Testing AI agents

Change the prompt. Know it still works.

Write down what good looks like as test cases: an input and the checks that must hold. Run them after every change, and see each check's expected and actual value. Test runs never send the email.

The problem

Every change changes the agent.

An agent is instructions, a model and some tools. Edit the instructions, switch the model or add a tool, and its behaviour changes, sometimes in a way you only notice when a customer does. Reading a few outputs by eye does not scale, and comparing outputs word for word fails on good runs, because the same right answer can be phrased a hundred ways.

  • What it did. Which tools it called, and which it must never call for this input.
  • What it decided. The label it chose, the part of the answer that has to be there.
  • How it finished. Without an error, within a time limit, under a cost limit.

How it works

An input, and the checks that must hold.

Each agent has an Evals tab. A test case is a name, the input the agent receives, and a list of checks. Eight kinds of check cover behaviour rather than wording.

  • Output contains, equals or matches. output_contains, output_equals and output_regex check what the agent returned.
  • Tool called, or not. tool_called and tool_not_called check what it did: labelled the email, and never replied to it.
  • No error. no_error passes when the run finished without one.
  • Time and cost limits. max_duration_ms and max_cost_usd hold the run to a budget.

The verdict

Keep the failing case.

Run a case and the agent runs on its input, then every check is judged. The case passes only when all of them pass. A case with no checks is not a pass.

Each check shows the expected value, what actually happened and whether it passed. The latest result stays next to the case, so a failing case is the first thing the next person to open the agent sees.

Safe to run

A test run never sends the email.

Test runs are sandboxed. The agent still sees a plausible result, so it behaves as it would for real, but nothing leaves the building.

  • Actions are stubbed. Integration tools, MCP tools, writes to documents and any HTTP call other than a GET return a note of what would have been called instead of calling it.
  • Reads are real. Fetching a page, searching the web and reading documents run for real, so the agent works from genuine context.
  • Kept apart from production. Test runs are marked as tests, and a failing one never trips a run alert. A case that reaches a tool needing approval, or stops to ask a question, ends with that reason instead of waiting for a person.

What you get

Tests that run when it matters.

  • A case for every mistake. When an agent gets something wrong in production, add the case. The same mistake cannot come back unnoticed.
  • Evals from code. Create, list, run and delete cases with a workspace key through the API and the TypeScript and Python SDKs, or from Claude, Cursor or any MCP client. A run waits for the verdict and lists the failed checks first.
  • Coverage at a glance. Listing an agent's cases returns each with its most recent result, so you can see what is tested and what is passing.

Here the agent labelled a legal threat as ordinary support instead of leaving it for a person. The fix belongs in its instructions, and the case stays in the suite so the fix is checked every time from then on.

GuideHow to test an AI agent before it goes live

More in Control

Decide what agents may do on their own, and keep the rest for a person.

Every capability

Give one job to an agent this week.

Start with one repetitive workflow, gate the sensitive step behind your approval, and read the first run end to end.