Blog Details

A demo isn't proof: how we teach students to evaluate AI applications.

A demo isn't proof: how we teach students to evaluate AI applications.

Obrix Academy
Authored by
Obrix Academy
Date Released
12 September, 2026
Comments
00 Comments

Building an AI assistant that gives a good answer in a demo is easy. Knowing whether it will keep giving good answers next week, on questions nobody has tried yet, is the hard part. That gap is the centre of our Generative AI & Agent Engineering course.

The course outcome is stated in those terms: graduates can show that an LLM application works using an evaluation suite, not just a demo.

Start with a golden dataset

Quality has to be defined before it can be measured. In the evaluation module students decide what good means for their specific task, then hand-label a golden set of 60 items drawn from real questions. They also learn to separate retrieval metrics from generation metrics, so they can tell whether a bad answer came from finding the wrong source or from writing the wrong thing.

“The assistant ships with a golden dataset, a before-and-after evaluation report and a CI regression gate.”

Obrix Academy
By Obrix Academy
Generative AI & Agent Engineering, Mini-project 2

Teach a judge, then check the judge

Reading hundreds of answers by hand does not scale, so students write an LLM judge with a scoring rubric. The catch is that a judge has biases too. They iterate on the rubric until the judge agrees with their own human labels, then use it to compare three configurations and pick a winner on evidence rather than instinct.

Make regressions fail the build

The suite is wired into CI so that a drop in faithfulness fails the build, the same way a broken unit test would. A change to a prompt, a model or a chunking strategy has to prove it did not make things worse before it ships.

Attack your own application

Security gets the same treatment. In a red-team exercise, students attack each other's applications, including injection planted inside a retrieved document, and then implement and retest defences. The lesson is that anything retrieved from a document store is untrusted input.

What the capstone has to show

Capstone requirementWhat it proves
Golden dataset of 50 or more items with an evaluation reportThe system improved, and by how much
CI regression gate on faithfulnessChanges cannot quietly make answers worse
Prompt-injection defences with a documented red-team resultThe application has been attacked and held
Full tracing with per-request cost and latencyYou know what each answer costs and how long it takes
Docker, CI/CD and a live URLIt runs outside the author's laptop

Who it is for

The course is for working developers, CS and IT graduates, and data analysts who already program in Python. It runs for six to eight weeks, with three to four sessions a week, mostly hands-on lab work. If you are not there yet, Python Backend Development is the designated starting point, and AI Adoption for Professionals suits people who want to use AI tools in their job without building them.

Leave a comment