Shreya Shankar and Hamel Husain, instructors of the industry-standard AI evaluations course with 4,500+ students from OpenAI and Google, give live feedback on a working eval system and walk through a five-step process for building better evals using Claude Code.
The value is not in the conclusion. It is in watching two practitioners critique a real system in real time, exposing the gaps most builders miss when they think their evals are good enough. They also demo how to run evals using ChatGPT or Claude without a complex engineering setup.
The takeaway is practical and transferable: eval quality is the bottleneck most one-person AI businesses hit before they realize it. If you are shipping AI-assisted work at any scale, this is the methodology to benchmark against.
[WATCH ON YOUTUBE →]