Engineering

Generative AI & Agent Engineering

Design, build and ship LLM applications that work against real documents and real tools, proven with an evaluation suite, not just a demo.

8 Weeks3-4 Sessions/WeekIn PersonIntermediate
RAG PipelinesAI AgentsLLM Evaluation

What you'll be able to do

Graduates can design, build, evaluate, secure and deploy an LLM application against a real document corpus and real tools, and can demonstrate that it works using an evaluation suite rather than a demo.

  • Call LLM APIs with token, cost and latency control
  • Write and version prompts as specifications
  • Enforce schema-validated structured output
  • Build a retrieval pipeline: chunking, embeddings, hybrid search, reranking, citations
  • Build an agent that calls tools reliably and terminates correctly
  • Construct a golden dataset and an LLM-as-judge evaluation harness with a CI regression gate
  • Defend against prompt injection and data leakage
  • Instrument traces, cost and latency
  • Ship the application containerised, deployed and documented

Who it's for

  • Working developers moving into AI product work
  • CS/IT graduates who can already program in Python
  • Data analysts with solid Python experience

Prerequisites

  • Comfortable with Python: functions, classes, virtual environments, package installation
  • Able to read JSON and consume an HTTP API
  • Basic Git

Applicants complete a short Python take-home task before enrolment. Candidates without programming experience should take Python Backend Development first.

All students complete the two-session Engineering Onboarding module (Git, pull requests, code review, environment and secrets management, debugging method, AI-assisted development policy) before Module 1.

Tools and technologies

PythonFastAPIOpenAI and Anthropic SDKsPydanticPostgreSQL with pgvectorLangGraphRagaspromptfooLangfuseDockerGitHub ActionsNext.js or Streamlit

Target roles

AI EngineerLLM Application DeveloperAgent DeveloperAI Product EngineerAI-augmented Full-Stack Developer

Course curriculum

8 modules · 6-8 weeks

Concepts
Capabilities and limits of foundation models; tokenisation and context windows; temperature and sampling; non-determinism; latency and cost models; model selection; streaming; rate limits, retries and idempotency.
Lab
Streaming CLI chat client with conversation history and a token budget; per-request cost instrumentation; backoff handling under enforced rate limits; three-model comparison on cost, latency and output quality.
Project
Wrap the client as a validated FastAPI endpoint with structured logging.

Concepts
the prompt as a specification; message roles and instruction hierarchy; few-shot design; task decomposition and chaining; prompt versioning and diffing; prompt caching for cost control; JSON mode and schema-constrained generation; validate-and-repair loops.
Lab
convert 200 unstructured support tickets into validated typed records; build a repair loop; A/B two prompt versions on a fixed input set; measure cache savings on a long system prompt.
Project
Mini-project 1: Document Extraction Service. A deployed API that ingests messy PDFs and emails and returns schema-validated JSON, with confidence handling, a human-review path, prompts versioned in Git, and documented failure modes.

Concepts
embeddings and distance; embedding model selection; chunking strategies (fixed, recursive, semantic, structure-aware) and overlap; metadata design; keyword, semantic and hybrid retrieval; index types; when retrieval is the wrong solution.
Lab
ingest a mixed-format corpus; implement three chunking strategies and measure retrieval hit rate against 30 fixed questions; add metadata filtering; diagnose retrieval failure caused by bad chunk boundaries.
Project
reusable ingestion pipeline (parse, chunk, embed, upsert, verify) with idempotent re-ingestion and document versioning.

Concepts
RAG architecture and its failure taxonomy; query rewriting, expansion and decomposition; hybrid search with score fusion; reranking and its cost trade-off; context assembly; enforced citations; refusal when context is insufficient; incremental updates; multi-tenant isolation.
Lab
upgrade the pipeline to hybrid search with reranking and measure the change; implement inline source citations; verify the refusal path triggers; add per-tenant filtering and test for leakage.
Project
grounded knowledge assistant over a real corpus, with citations and a chat interface.

Concepts
defining quality for a specific task (correctness, faithfulness, relevance, tone, latency, cost); building a golden dataset from real questions; retrieval metrics against generation metrics; LLM-as-judge rubric design and calibration; judge bias; pairwise and pointwise scoring; regression testing prompts; error analysis and failure clustering; offline against online evaluation.
Lab
hand-label a 60-item golden set; write an LLM judge and iterate the rubric until it agrees with human labels; evaluate three configurations and select a winner on evidence; wire the suite into CI so a faithfulness regression fails the build.
Project
Mini-project 2: the assistant ships with a golden dataset, a before-and-after evaluation report and a CI regression gate.

Concepts
tool calling mechanics; tool design as API design (naming, descriptions, narrow parameters, idempotency, model-readable errors); reason-and-act loops; planning against reactive execution; short-term and long-term memory; termination and cost ceilings; human approval for destructive actions; failure recovery; when a fixed workflow beats an agent; MCP for standardised tool exposure.
Lab
add three tools (SQL query, retriever, external HTTP API) and make selection reliable; observe misuse from a poor tool description and correct it; add step and cost limits; add an approval gate before write actions; build a trajectory evaluation set scoring tool selection order.
Project
the assistant becomes agentic and is evaluated on tool-selection accuracy.

Concepts
the LLM threat model, including direct and indirect prompt injection; treating retrieved content as untrusted input; input and output filtering; PII detection and redaction; least-privilege tool credentials; rate limiting; cost engineering through caching, model routing and context trimming; latency budgets; tracing and span design; logging without exposing secrets.
Lab
red-team exercise in which students attack each other's applications, including injection planted inside a retrieved document, then implement and retest defences; add semantic caching and measure the cost reduction; implement small-model-first routing; locate the real latency bottleneck from traces.
Project
written security and cost review of the student's own application.

Concepts
production architecture for AI applications: async handling, queues for long jobs, streaming, session storage, secrets management, environment separation, provider fallback, versioning prompts and indexes together, rollback; when fine-tuning is the right decision and why retrieval and prompting usually come first.
Lab
containerise the application; build the pipeline (lint, test, evaluation gate, build, deploy); load test and fix the first bottleneck; simulate a provider outage and verify fallback.

Capstone project

A production-style AI assistant for a chosen domain with a real document corpus and a real data source, such as clinic operations, insurance policy support, or a bank product knowledge base with a transaction database.

Requirements

  • Hybrid RAG with citations and refusal behaviour
  • At least three tools, including one database query and one write action behind an approval gate
  • Golden dataset of 50 or more items with an evaluation report showing measured improvement
  • CI regression gate on faithfulness
  • Prompt-injection defences with a documented red-team result
  • Full tracing with per-request cost and latency
  • Usable frontend
  • Docker, CI/CD and a live URL
  • README covering architecture, evaluation results, cost model, security notes and known limitations

Assessment

30%Weekly labs and mini-projects
20%Code review participation
35%Capstone
15%Demo Day presentation and technical questioning

Capstone standard: runs from a clean clone, tests and evaluations pass in CI, deployed at a URL, README explains architecture and trade-offs, commit history shows incremental work, and the student can defend every design decision.

Out of scope

  • Model pre-training
  • Transformer internals beyond intuition
  • Full fine-tuning runs
  • GPU infrastructure
  • Multi-agent research architectures
  • Voice agents

Enquire about this course

Ask about the next cohort, schedule or prerequisites and our team will get back to you.

Generative AI & Agent Engineering
Keep learning

Related Courses.

AI & Machine Learning
EngineeringIntermediate

AI & Machine Learning

Take a business problem from raw data to a deployed, monitored model, with honest evaluation and real production practice throughout.

Feature EngineeringModel TrainingMLOps
8 Weeks4 Sessions/WeekIn Person
View course
Full-Stack Web Development
EngineeringBeginner-friendly

Full-Stack Web Development

Build, test, containerise and deploy a complete multi-user web application in TypeScript, and defend every layer of it under questioning.

TypeScriptDatabase DesignDocker
8 Weeks4 Sessions/WeekIn Person
View course
Software QA & Test Automation
EngineeringBeginner-friendly

Software QA & Test Automation

Take an unfamiliar web application and build a complete, maintainable quality strategy, from risk based test design to automated pipelines.

Risk-Based TestingTest AutomationAPI Testing
7 Weeks4 Sessions/WeekIn Person
View course
AWS Cloud Engineering
EngineeringIntermediate

AWS Cloud Engineering

Stand up a full production environment on AWS, from networking and compute to monitoring, and defend the monthly cost with confidence.

VPC NetworkingTerraformCost Optimisation
7 Weeks4 Sessions/WeekIn Person
View course
DevOps Engineering
EngineeringIntermediate

DevOps Engineering

Build and operate the delivery platform for a multi-service application, from containers and Kubernetes to incident response and postmortems.

KubernetesCI/CDIncident Response
8 Weeks4 Sessions/WeekIn Person
View course
Data Analytics & Analytics Engineering
EngineeringBeginner-friendly

Data Analytics & Analytics Engineering

Take raw, messy business data through the full analytics chain, from SQL and Python to dashboards and a recommendation you can defend.

SQLData ModellingDashboards
6 Weeks3 Sessions/WeekIn Person
View course