Get AI Agent Evals on Leanpub from $14.99
Pay what you choose in USD: $14.99 minimum, $39 suggested. Leanpub currently lists the book as 91% complete. Check the retailer for the current edition and price.
AI Agent Evals by Thomas De Vos
Judge the run by what the agent actually did.
A plausible answer can hide an incorrect tool call, an unnecessary question or a refund that should never have happened. Testing AI agents means checking the actions and changed state against a task contract, then keeping enough evidence to challenge the result.
This book develops a small support-system example into an evaluation workflow you can inspect. You write task contracts, protect datasets, test the graders themselves and decide which failures should block a release.
24 chapters, eight appendices and a 498-page Leanpub edition. Supplied labs use synthetic tasks and offline execution; a separate recorded live-model case study examines actual model behaviour in a synthetic refund lab.

Who this book is for
AI Agent Evals is for engineers and architects who need to decide whether an agent change is good enough to ship. It is also for evaluation owners who must explain a result to reviewers, choose where human judgement belongs and justify the cost of repeated experiments.
You will get more from the exercises if you are comfortable reading code and inspecting structured data. The book is a practical engineering guide, rather than an introduction to using a chatbot. Its support-system example gives the evaluation work a concrete setting without treating that example as evidence about your production system.
What you will learn to build
- Task contracts that define acceptable outcomes without prescribing one exact conversation.
- Datasets with task families, provenance, reviewed labels and protected holdouts, so tuning does not quietly contaminate the comparison.
- Evaluation harnesses that isolate trial state, preserve crashes and timeouts, and keep live execution separate from offline replay.
- Code graders with their own tests, human review rubrics and model judges that you calibrate before trusting their scores.
- Comparisons that account for nondeterminism, incomplete evidence, latency and cost instead of hiding them behind one average.
- CI release gates that can block a change, require review or call for another experiment, with a traceable reason for the decision.
A useful evaluation should change an engineering decision. A reassuring dashboard that cannot explain a failed run is a poor substitute.
Inside the book
Foundations: chapters 1 to 4
Start by watching an agent fail. Turn the expected behaviour into a testable task, build a dataset that represents the work and create labels whose meaning survives review. These chapters separate golden cases from holdouts and make dataset permissions explicit before tuning begins.
Evaluation machinery: chapters 5 to 8
Build the harness, prepare it for real models and make every run reproducible and inspectable. Provider boundaries, isolated state, retry behaviour and retained failures matter here. Then write and test code graders against the evidence the task actually requires.
Human and model judgement: chapters 9 to 11
Design human review that produces useful disagreement rather than an unexplained vote. Build a model judge, calibrate it and challenge its failure modes. The evaluator is part of the system under test.
Measuring under uncertainty: chapters 12 and 13
Measure nondeterministic agents and compare changes with statistical care. A small improvement in an aggregate score does not settle whether a change helped the tasks that matter, or whether the experiment has enough evidence to support the claim.
Agent-specific evaluation: chapters 14 to 18
Evaluate retrieval and evidence use, tool trajectories, long-running state, simulated conversations, adversarial behaviour and coding agents. These chapters ask whether the agent respected the task and its constraints, including cases where the correct action is to stop or do nothing.
Production and operations: chapters 19 to 24
Gate releases on complete evidence, instrument traces and learn from production without confusing observation with proof. Work through drift, regressions, rollout and rollback, failure diagnosis and the operating cost of an evaluation programme.
The appendices cover metrics and denominators, exercise checks, reusable templates, defensive evaluation engineering, statistics, framework and benchmark selection, delivery records and governance questions.
A live run with an uncomfortable result
The recorded Astra case study contains ten tasks: three passed and seven failed. It examines questions, actions and refund ledgers rather than accepting a convincing final answer as proof of success. Correct refund outcomes can coexist with unnecessary questions that violate the task contract.
This is a small experiment in a synthetic refund lab. It is not a model ranking, a production benchmark or evidence that another deployment will behave the same way. The teaching value is in the retained evidence and the distinction between an apparently correct outcome and a fully satisfied task.
How to use the exercises
Work through the cumulative support-system example and inspect the supplied evidence before adapting it. Keep offline checks, replay of an existing run and a new live-model execution separate in your own reports. A passing protocol test does not establish model quality.
The AI Agent Evals companion code contains the reader exercises. Follow its setup instructions and the book’s edition-matching guidance when you run the labs, and retain failed results before attempting repairs.
Read more before you buy
For a shorter introduction, read why Claude Code evals should start with bad runs. The AI coding agents guide puts evaluation alongside permissions and review. For the security boundary around an agent, see Securing Enterprise AI Agents.
AI Agent Evals concentrates on how to establish and challenge evaluation evidence. Compare the other books if your immediate problem is production Claude Code, agent security or legacy modernization.
Get the book
View AI Agent Evals on Leanpub
The minimum price is $14.99 USD; the suggested price is $39 USD. Those are pay-what-you-choose options, not a discount claim. The Leanpub page has the full table of contents and current publication status.