How to evaluate an AI agent
Evaluating an agent requires measuring the whole task, traces, tools, cost, and recovery from failure. A convincing final answer is not enough.
Links containing ?t= open the video at a specific second.
Video summary
The ideas to retain
The unit of evaluation is a task
A useful task should specify:
Four dimensions worth measuring
1. Outcome
Did the task finish correctly? This is the most visible dimension, but not the only one. The metric should represent partial results and distinguish “could not complete” from “completed…


