Generative AI & Agent Engineering
Design, build and ship LLM applications that work against real documents and real tools, proven with an evaluation suite, not just a demo.
What you'll be able to do
Graduates can design, build, evaluate, secure and deploy an LLM application against a real document corpus and real tools, and can demonstrate that it works using an evaluation suite rather than a demo.
- Call LLM APIs with token, cost and latency control
- Write and version prompts as specifications
- Enforce schema-validated structured output
- Build a retrieval pipeline: chunking, embeddings, hybrid search, reranking, citations
- Build an agent that calls tools reliably and terminates correctly
- Construct a golden dataset and an LLM-as-judge evaluation harness with a CI regression gate
- Defend against prompt injection and data leakage
- Instrument traces, cost and latency
- Ship the application containerised, deployed and documented
Who it's for
- Working developers moving into AI product work
- CS/IT graduates who can already program in Python
- Data analysts with solid Python experience
Prerequisites
- Comfortable with Python: functions, classes, virtual environments, package installation
- Able to read JSON and consume an HTTP API
- Basic Git
Applicants complete a short Python take-home task before enrolment. Candidates without programming experience should take Python Backend Development first.
All students complete the two-session Engineering Onboarding module (Git, pull requests, code review, environment and secrets management, debugging method, AI-assisted development policy) before Module 1.
Tools and technologies
Target roles
Course curriculum
- Concepts
- Capabilities and limits of foundation models; tokenisation and context windows; temperature and sampling; non-determinism; latency and cost models; model selection; streaming; rate limits, retries and idempotency.
- Lab
- Streaming CLI chat client with conversation history and a token budget; per-request cost instrumentation; backoff handling under enforced rate limits; three-model comparison on cost, latency and output quality.
- Project
- Wrap the client as a validated FastAPI endpoint with structured logging.
- Concepts
- the prompt as a specification; message roles and instruction hierarchy; few-shot design; task decomposition and chaining; prompt versioning and diffing; prompt caching for cost control; JSON mode and schema-constrained generation; validate-and-repair loops.
- Lab
- convert 200 unstructured support tickets into validated typed records; build a repair loop; A/B two prompt versions on a fixed input set; measure cache savings on a long system prompt.
- Project
- Mini-project 1: Document Extraction Service. A deployed API that ingests messy PDFs and emails and returns schema-validated JSON, with confidence handling, a human-review path, prompts versioned in Git, and documented failure modes.
- Concepts
- embeddings and distance; embedding model selection; chunking strategies (fixed, recursive, semantic, structure-aware) and overlap; metadata design; keyword, semantic and hybrid retrieval; index types; when retrieval is the wrong solution.
- Lab
- ingest a mixed-format corpus; implement three chunking strategies and measure retrieval hit rate against 30 fixed questions; add metadata filtering; diagnose retrieval failure caused by bad chunk boundaries.
- Project
- reusable ingestion pipeline (parse, chunk, embed, upsert, verify) with idempotent re-ingestion and document versioning.
- Concepts
- RAG architecture and its failure taxonomy; query rewriting, expansion and decomposition; hybrid search with score fusion; reranking and its cost trade-off; context assembly; enforced citations; refusal when context is insufficient; incremental updates; multi-tenant isolation.
- Lab
- upgrade the pipeline to hybrid search with reranking and measure the change; implement inline source citations; verify the refusal path triggers; add per-tenant filtering and test for leakage.
- Project
- grounded knowledge assistant over a real corpus, with citations and a chat interface.
- Concepts
- defining quality for a specific task (correctness, faithfulness, relevance, tone, latency, cost); building a golden dataset from real questions; retrieval metrics against generation metrics; LLM-as-judge rubric design and calibration; judge bias; pairwise and pointwise scoring; regression testing prompts; error analysis and failure clustering; offline against online evaluation.
- Lab
- hand-label a 60-item golden set; write an LLM judge and iterate the rubric until it agrees with human labels; evaluate three configurations and select a winner on evidence; wire the suite into CI so a faithfulness regression fails the build.
- Project
- Mini-project 2: the assistant ships with a golden dataset, a before-and-after evaluation report and a CI regression gate.
- Concepts
- tool calling mechanics; tool design as API design (naming, descriptions, narrow parameters, idempotency, model-readable errors); reason-and-act loops; planning against reactive execution; short-term and long-term memory; termination and cost ceilings; human approval for destructive actions; failure recovery; when a fixed workflow beats an agent; MCP for standardised tool exposure.
- Lab
- add three tools (SQL query, retriever, external HTTP API) and make selection reliable; observe misuse from a poor tool description and correct it; add step and cost limits; add an approval gate before write actions; build a trajectory evaluation set scoring tool selection order.
- Project
- the assistant becomes agentic and is evaluated on tool-selection accuracy.
- Concepts
- the LLM threat model, including direct and indirect prompt injection; treating retrieved content as untrusted input; input and output filtering; PII detection and redaction; least-privilege tool credentials; rate limiting; cost engineering through caching, model routing and context trimming; latency budgets; tracing and span design; logging without exposing secrets.
- Lab
- red-team exercise in which students attack each other's applications, including injection planted inside a retrieved document, then implement and retest defences; add semantic caching and measure the cost reduction; implement small-model-first routing; locate the real latency bottleneck from traces.
- Project
- written security and cost review of the student's own application.
- Concepts
- production architecture for AI applications: async handling, queues for long jobs, streaming, session storage, secrets management, environment separation, provider fallback, versioning prompts and indexes together, rollback; when fine-tuning is the right decision and why retrieval and prompting usually come first.
- Lab
- containerise the application; build the pipeline (lint, test, evaluation gate, build, deploy); load test and fix the first bottleneck; simulate a provider outage and verify fallback.
Capstone project
A production-style AI assistant for a chosen domain with a real document corpus and a real data source, such as clinic operations, insurance policy support, or a bank product knowledge base with a transaction database.
Requirements
- Hybrid RAG with citations and refusal behaviour
- At least three tools, including one database query and one write action behind an approval gate
- Golden dataset of 50 or more items with an evaluation report showing measured improvement
- CI regression gate on faithfulness
- Prompt-injection defences with a documented red-team result
- Full tracing with per-request cost and latency
- Usable frontend
- Docker, CI/CD and a live URL
- README covering architecture, evaluation results, cost model, security notes and known limitations
Assessment
Capstone standard: runs from a clean clone, tests and evaluations pass in CI, deployed at a URL, README explains architecture and trade-offs, commit history shows incremental work, and the student can defend every design decision.
Out of scope
- Model pre-training
- Transformer internals beyond intuition
- Full fine-tuning runs
- GPU infrastructure
- Multi-agent research architectures
- Voice agents
Enquire about this course
Ask about the next cohort, schedule or prerequisites and our team will get back to you.



