All videosAI agents2:14

How to evaluate an agent

Separate task permissions from the verifier’s resources. Record out-of-scope attempts and inspect the trajectory.

Links containing ?t= open the video at a specific second.

Video summary

The ideas to retain

01

The unit of evaluation is a task

A useful task should specify:

02

Four dimensions worth measuring

03

1. Outcome

Did the task finish correctly? This is the most visible dimension, but not the only one. The evaluation should account for partial outcomes and distinguish “could not complete” from…

Key moments

Jump directly to a section

  1. The test starts with an initial state.
  2. The same result can require more steps.
  3. A policy violation invalidates success.
  4. The mean can hide the tail.
  5. Check the sum, not the wording.
  6. The agent must not rewrite its own test.
Read the reviewed transcript

Transcript of the visual text. This video has no narration.

00:00 — The test starts with an initial state.

The task asks for overdue invoices to be totalled and saved as a draft, not sent.

In this example, 20 and 30 are overdue; the invoice for 10 is not.

The expected result is 50, with a saved draft and no message sent.

Illustrative example, not production measurements.

00:24 — The same result can require more steps.

Both runs produce the same correct draft.

One uses three calls; the other repeats queries and consumes eight.

Compare trajectories for the same task, not only their final text.

Illustrative example, not production measurements.

00:46 — A policy violation invalidates success.

The total can be correct and the cost can remain within budget.

But sending a forbidden draft violates the task contract.

The evaluation must preserve that failure rather than average it into success.

Illustrative example, not production measurements.

01:08 — The mean can hide the tail.

Eighteen tasks take one second; two others take nine.

The mean is 1.8 seconds, but the 95th percentile is nine in this sample.

Inspect slow cases and their tools, not only the average.

Illustrative example, not production measurements.

01:30 — Check the sum, not the wording.

A clear response can still claim that 20 plus 30 equals 60.

A deterministic verifier compares the number with the expected result: 50.

Use the objective condition when one exists; writing style cannot replace it.

Illustrative example, not production measurements.

01:52 — The agent must not rewrite its own test.

A test can appear to pass if the agent can change its expected result.

Separate task permissions from the verifier’s resources.

Record out-of-scope attempts and inspect the trajectory.

Illustrative example, not production measurements.