
A demo isn't proof: how we teach students to evaluate AI applications.
Obrix Academy
12 September, 2026
00 Comments
Building an AI assistant that gives a good answer in a demo is easy. Knowing whether it will keep giving good answers next week, on questions nobody has tried yet, is the hard part. That gap is the centre of our Generative AI & Agent Engineering course.
The course outcome is stated in those terms: graduates can show that an LLM application works using an evaluation suite, not just a demo.
Start with a golden dataset
Quality has to be defined before it can be measured. In the evaluation module students decide what good means for their specific task, then hand-label a golden set of 60 items drawn from real questions. They also learn to separate retrieval metrics from generation metrics, so they can tell whether a bad answer came from finding the wrong source or from writing the wrong thing.
“The assistant ships with a golden dataset, a before-and-after evaluation report and a CI regression gate.”
Teach a judge, then check the judge
Reading hundreds of answers by hand does not scale, so students write an LLM judge with a scoring rubric. The catch is that a judge has biases too. They iterate on the rubric until the judge agrees with their own human labels, then use it to compare three configurations and pick a winner on evidence rather than instinct.
Make regressions fail the build
The suite is wired into CI so that a drop in faithfulness fails the build, the same way a broken unit test would. A change to a prompt, a model or a chunking strategy has to prove it did not make things worse before it ships.
Attack your own application
Security gets the same treatment. In a red-team exercise, students attack each other's applications, including injection planted inside a retrieved document, and then implement and retest defences. The lesson is that anything retrieved from a document store is untrusted input.
What the capstone has to show
| Capstone requirement | What it proves |
|---|---|
| Golden dataset of 50 or more items with an evaluation report | The system improved, and by how much |
| CI regression gate on faithfulness | Changes cannot quietly make answers worse |
| Prompt-injection defences with a documented red-team result | The application has been attacked and held |
| Full tracing with per-request cost and latency | You know what each answer costs and how long it takes |
| Docker, CI/CD and a live URL | It runs outside the author's laptop |
Who it is for
The course is for working developers, CS and IT graduates, and data analysts who already program in Python. It runs for six to eight weeks, with three to four sessions a week, mostly hands-on lab work. If you are not there yet, Python Backend Development is the designated starting point, and AI Adoption for Professionals suits people who want to use AI tools in their job without building them.
Leave a comment