Evals for agents

How do you build evaluations for agents? Model capabilities are evolving fast, user expectations are shifting, and both inputs and outputs are highly variable. This series walks through how to think about agent evals — from the kinds of agents you might be building, to identifying risk, defining quality, and combining qualitative research with metrics.

What will we discuss?

  • Why evaluating agents is different from evaluating prompts or models
  • The two big families: agents that help vs. agents that take action
  • How to identify the risks that actually matter
  • How to turn fuzzy "quality" into something testable
  • Where qualitative research fits in
  • Where metrics genuinely help (and where they mislead)
  • A practical playbook for agent evals that survive model upgrades

Who is this for?

  • Anyone building agents — copilots, assistants, action-takers — who needs to know if they're actually any good
  • PMs, researchers, and designers working alongside engineers on agent products
  • Anyone who's done basic evals and now needs to handle the variability and risk that agents introduce
14-day free trial
8 videos

Evals for agents

$79/m
  • All 8 videos in this series
  • Prompts, templates and resources from each episode
  • Watch at your own pace, on any device
  • Full access to every other course on model context experience
Start free trial

Full access to all courses, $79/m after the trial. Cancel anytime.

What you'll learn

  • How agent evals differ from evals for simpler AI features
  • The difference between agents that help and agents that take action
  • How to identify risk and define quality for an agent
  • How to combine qualitative research with metrics

Learn from Peter

Peter Van Dijck

Peter Van Dijck

Peter has been building AI products since 2023 and teaching teams how to build them since 2024. He has spent that time in the trenches with evals, observability, synthetic data and context engineering, and turns what actually works into practical, no-fluff lessons.

Every episode is taught by Peter himself, in plain language, for product managers, designers, researchers and strategists who need to understand how AI systems are really built, without needing to be an engineer.

Peter on LinkedIn

Who this is for

Product managers

You are shipping an agent and need to know whether it is safe and useful.

UX researchers

You want to bring qualitative rigor to evaluating agent behaviour.

AI leads

You need an evaluation approach that survives fast-changing models and expectations.

What's included

  • All 8 videos in this series
  • Prompts, templates and resources from each episode
  • Watch at your own pace, on any device
  • Full access to every other course on model context experience
  • A framework for evaluating both helper and action-taking agents

What participants say

“This is opening up all sorts of new neural pathways for me to see under the hood more of how the sausage is made! 🙏”
Senior UX Designer
“Very timely at my enterprise software company as evaluation of AI features scales.”
Principal Product Designer
“Everything I know about evals is from Peter's talk, which is why I'm back to find out more!”
UX Researcher

The videos

1. Evals for agents

Episode 1.1
Evaluating agents — introduction

Why evaluating agents is different from evaluating prompts or models. Set the stage for the series: the shifting ground (models, users, inputs/outputs) and what makes a good agent eval.

Coming Soon
Episode 1.2
Agents that help

Agents whose job is to assist a person — copilots, researchers, summarizers. What "good" looks like when the human stays in the loop, and how that shapes what you measure.

Coming Soon
Episode 1.3
Agents that take action

Agents that do things in the world — book, send, write, deploy. The eval bar is higher: correctness, reversibility, and trust become first-class concerns.

Coming Soon
Episode 1.4
Identifying risk

Where agents can go wrong and which failures actually matter. A practical way to map risks so your evals cover what's costly, not just what's easy to measure.

Coming Soon
Episode 1.5
Defining quality

What does "good" even mean for an agent? Turning fuzzy expectations into concrete, testable criteria that hold up across variable inputs and outputs.

Coming Soon
Episode 1.6
The role of qual research

Why you can't eval your way out of not understanding users. How qualitative research surfaces the failure modes and quality dimensions that metrics alone will miss.

Coming Soon
Episode 1.7
The role of metrics

Where metrics genuinely help, where they mislead, and how to build a metric set that complements — rather than replaces — human judgment.

Coming Soon
Episode 1.8
Wrapping it up

Pulling the threads together: a practical playbook for building agent evals that survive model upgrades, shifting user expectations, and the inherent variability of agent work.

Coming Soon

Also included

Your subscription unlocks every course on model context experience.

10 videos

AI 101

Build a deeper understanding of AI. Why do models have a personality? What is context engineering?

Course details →
6 videos

Evals 101

How do we know if our AI systems are working well? *The* key skill for UX researchers and product people.

Course details →
8 videos

Claude Code for non-engineers

Despite the "code" in its name, Claude Code is perhaps the most popular agentic AI system right now. Understanding and using it gives you a glimpse into what's coming the coming months and years in terms of agents. And it can be incredibly useful for non-coding tasks.

Course details →
4 videos

UX.md

Why should I make an UX.md file, and how do I know it's working?

Course details →
3 videos

Project Planning for AI

If AI is different, and AI projects are different, how do we plan projects for AI? What are the roles and tracks we should consider? What are some common gotchas?

Course details →
5 videos

Designing with Claude

A hands-on walkthrough of Claude Design — Anthropic's tool that creates real, code-based designs. Set up a design system, generate and refine a landing page, and see where designing-by-code shines: interactive, animated, production-quality design with a design-to-engineering handoff measured in minutes.

Course details →
7 videos

Content Strategy for LLMs

Content strategy is changing now that LLMs are reading, writing, and rewriting most of what we publish. This series is a practical walkthrough for content folks: setting up the right tools, structuring content as markdown, defining tone of voice and microcopy in ways an LLM can actually follow, and evaluating what comes out the other end.

Course details →