Writing evals is a core skill for making AI products that actually work. Evals are our "definition of what good looks like". They are both harder than they seem to get right, and at the same time not rocket science at all - anyone can learn to write evals.
How do we know if our AI systems are working well? *The* key skill for UX researchers and product people.
Full access to all courses, $79/m after the trial. Cancel anytime.
Peter Van Dijck
Peter has been building AI products since 2023 and teaching teams how to build them since 2024. He has spent that time in the trenches with evals, observability, synthetic data and context engineering, and turns what actually works into practical, no-fluff lessons.
Every episode is taught by Peter himself, in plain language, for product managers, designers, researchers and strategists who need to understand how AI systems are really built, without needing to be an engineer.
Peter on LinkedInYou already measure quality with people; evals let you do it with models too.
You need to know whether an AI feature actually works before it ships.
You want a shared, repeatable way to talk about AI quality with your team.
“Really practical, this is the course I wish I had a year ago. It walks you through building evals from nothing and then, which I found even more useful, how to figure out why an existing eval isn't telling you anything. Turns out ours had been broken for months. If you are working on anything with an LLM in it and you can't honestly answer "did that last change make it better", do this course. I still have it open as a reference.”
“Before this, evals were something the ML team did and the rest of us just nodded. Now I can run them myself and the conversations with engineering about quality are completely different. Really glad I took this.”
“I came in a bit sceptical because "evals" sounded like an engineering topic. It isn't, or at least not only. The course made me realise how much of eval design is really UX research, and I already had most of the skills, I just didn't know how to apply them here. Lots of small things that are easy to miss but make a huge difference to the output.”
This is what we came for: some hands-on eval writing.
What are evals, why do we need them, and why isn't this just QA?
This is the fun part, hands-on writing evals together.
Evals can be tricky, and it's easy to make some very expensive (in terms of quality, end result and cost) mistakes.
One reason evals are tricky, is that it can be hard to define what Good looks like when working (as we are) in a team.
There are no evals without data sets. How do we create solid data sets? How many data points are enough? What about synthetic data?
Your subscription unlocks every course on model context experience.
Build a deeper understanding of AI. Why do models have a personality? What is context engineering?
Course details →Despite the "code" in its name, Claude Code is perhaps the most popular agentic AI system right now. Understanding and using it gives you a glimpse into what's coming the coming months and years in terms of agents. And it can be incredibly useful for non-coding tasks.
Course details →If AI is different, and AI projects are different, how do we plan projects for AI? What are the roles and tracks we should consider? What are some common gotchas?
Course details →A hands-on walkthrough of Claude Design — Anthropic's tool that creates real, code-based designs. Set up a design system, generate and refine a landing page, and see where designing-by-code shines: interactive, animated, production-quality design with a design-to-engineering handoff measured in minutes.
Course details →How do you build evaluations for agents? Model capabilities are evolving fast, user expectations are shifting, and both inputs and outputs are highly variable. This series walks through how to think about agent evals — from the kinds of agents you might be building, to identifying risk, defining quality, and combining qualitative research with metrics.
Course details →Content strategy is changing now that LLMs are reading, writing, and rewriting most of what we publish. This series is a practical walkthrough for content folks: setting up the right tools, structuring content as markdown, defining tone of voice and microcopy in ways an LLM can actually follow, and evaluating what comes out the other end.
Course details →