Evals 101

Writing evals is a core skill for making AI products that actually work. Evals are our "definition of what good looks like". They are both harder than they seem to get right, and at the same time not rocket science at all - anyone can learn to write evals.

What will we discuss?

  • What evals are (it's not rocket science)
  • How do write your first eval
  • The structure of a good evaluation
  • When to use or not use evals
  • How to set up evals in a project
  • Which tools to use, and when to use them
  • How to create useful data sets

Who is this for?

  • Anyone who is building things with AI, whether that is an internal tool, an agent, a product or a workflow
  • Anyone interested in learning how to do evals in a practical, no-nonsense way
14-day free trial
6 videos

Evals 101

$79/m

How do we know if our AI systems are working well? *The* key skill for UX researchers and product people.

  • All 6 videos in this series
  • Prompts, templates and resources from each episode
  • Watch at your own pace, on any device
  • Full access to every other course on model context experience
Start free trial

Full access to all courses, $79/m after the trial. Cancel anytime.

What you'll learn

  • What evals are, why every AI product needs them, and why they are not QA
  • How to write your first eval, together, step by step
  • How to define what "good" looks like as a team
  • How to build data sets that are worth evaluating against
  • Common eval mistakes and how to avoid them

Learn from Peter

Peter Van Dijck

Peter Van Dijck

Peter has been building AI products since 2023 and teaching teams how to build them since 2024. He has spent that time in the trenches with evals, observability, synthetic data and context engineering, and turns what actually works into practical, no-fluff lessons.

Every episode is taught by Peter himself, in plain language, for product managers, designers, researchers and strategists who need to understand how AI systems are really built, without needing to be an engineer.

Peter on LinkedIn

Who this is for

UX researchers

You already measure quality with people; evals let you do it with models too.

Product managers

You need to know whether an AI feature actually works before it ships.

Team leads

You want a shared, repeatable way to talk about AI quality with your team.

What's included

  • All 6 videos in this series
  • Prompts, templates and resources from each episode
  • Watch at your own pace, on any device
  • Full access to every other course on model context experience
  • A first eval written together that you can copy for your own product

What participants say

“Really practical, this is the course I wish I had a year ago. It walks you through building evals from nothing and then, which I found even more useful, how to figure out why an existing eval isn't telling you anything. Turns out ours had been broken for months. If you are working on anything with an LLM in it and you can't honestly answer "did that last change make it better", do this course. I still have it open as a reference.”
Product Manager
“Before this, evals were something the ML team did and the rest of us just nodded. Now I can run them myself and the conversations with engineering about quality are completely different. Really glad I took this.”
Principal Designer
“I came in a bit sceptical because "evals" sounded like an engineering topic. It isn't, or at least not only. The course made me realise how much of eval design is really UX research, and I already had most of the skills, I just didn't know how to apply them here. Lots of small things that are easy to miss but make a huge difference to the output.”
UX Researcher

The videos

1. Let's write some evals

Episode 1.1 · 3:37
Evals intro: set up your accounts

This is what we came for: some hands-on eval writing.

Members only
Episode 1.2 · 11:11
An introduction to evals

What are evals, why do we need them, and why isn't this just QA?

Members only
Episode 1.3 · 13:53
Let's write an eval together

This is the fun part, hands-on writing evals together.

Members only
Episode 1.4 · 7:12
Eval Tips and Common Mistakes

Evals can be tricky, and it's easy to make some very expensive (in terms of quality, end result and cost) mistakes.

Members only
Episode 1.5 · 15:15
How we define what Good looks like

One reason evals are tricky, is that it can be hard to define what Good looks like when working (as we are) in a team.

Members only
Episode 1.5 · 8:12
Creating Data Sets

There are no evals without data sets. How do we create solid data sets? How many data points are enough? What about synthetic data?

Members only

Also included

Your subscription unlocks every course on model context experience.

10 videos

AI 101

Build a deeper understanding of AI. Why do models have a personality? What is context engineering?

Course details →
8 videos

Claude Code for non-engineers

Despite the "code" in its name, Claude Code is perhaps the most popular agentic AI system right now. Understanding and using it gives you a glimpse into what's coming the coming months and years in terms of agents. And it can be incredibly useful for non-coding tasks.

Course details →
3 videos

Project Planning for AI

If AI is different, and AI projects are different, how do we plan projects for AI? What are the roles and tracks we should consider? What are some common gotchas?

Course details →
5 videos

Designing with Claude

A hands-on walkthrough of Claude Design — Anthropic's tool that creates real, code-based designs. Set up a design system, generate and refine a landing page, and see where designing-by-code shines: interactive, animated, production-quality design with a design-to-engineering handoff measured in minutes.

Course details →
8 videos

Evals for agents

How do you build evaluations for agents? Model capabilities are evolving fast, user expectations are shifting, and both inputs and outputs are highly variable. This series walks through how to think about agent evals — from the kinds of agents you might be building, to identifying risk, defining quality, and combining qualitative research with metrics.

Course details →
7 videos

Content Strategy for LLMs

Content strategy is changing now that LLMs are reading, writing, and rewriting most of what we publish. This series is a practical walkthrough for content folks: setting up the right tools, structuring content as markdown, defining tone of voice and microcopy in ways an LLM can actually follow, and evaluating what comes out the other end.

Course details →