A 15-part course
Most eval advice is prose and dashboards. This one builds the measurement stack by hand, one concept at a time, so you can tell a real 2% gain from noise: metrics and the confusion matrix, golden sets, annotator agreement, LLM-as-judge and its biases, confidence intervals, significance and power, pass@k, arenas, calibration, and a CI regression gate. Every part runs on a small labeled set you can trace end to end, with runnable code and an interactive figure.
Start with Part 1: What Is a Score?Core track · Parts 1–11
Frontier track · Parts 12–15
The production and agent territory you reach for once the offline gate is solid: the eval flywheel, online A/B and guardrail metrics, agent-trajectory evals, and a tiny end-to-end harness.