AI agent monitoring
Agents that review your agents. Worst first, with the fix.
Built-in intelligence agents read your fleet's runs for failure trends, slower runs, odd patterns, weak prompts and wasted tokens. What they find lands in one inbox, ranked by severity, each finding with what to change.
The problem
An alert says something broke. Not why.
Finding out why means reading the failed runs side by side until the pattern shows: the same source returning 403, the same step getting slower every day, the same instructions sent on every call. It is the step people skip, and it is work an agent can do.
The analysts
A team of reviewers in every workspace.
Every workspace comes with a set of intelligence agents, each with one job. They read your runs and your agents' setup, and each writes what it found as a finding.
- Stability Watchdog. Failure-rate trends, clusters of the same error and silent degradation over the last seven days.
- Performance Auditor. An agent's latency this week against the week before, and the regressions.
- Pattern Analyst. Recurring sequences of events, unusual flows and failure clusters across recent runs.
- Anomaly Detective. Silent failures, out-of-band errors and rare tool calls.
- Output Quality Monitor. Samples recent completed runs and scores the output from 1 to 5 for task completion, accuracy, conciseness and tone.
- Prompt Strength Classifier. Rates every agent's instructions as Weak, Good, Great or Perfect, with specific improvements.
- Token Optimizer. Finds the token waste in recent runs and ranks the prompt fixes by estimated saving.
- Fix Advisor. Reads a failing agent's history and setup and prescribes the change, down to the integration to reconnect.
How they run
On a schedule, without being asked.
A built-in pipeline, Fleet Sweep, runs automatically on a schedule. It picks the active agents that ran in the last seven days, up to the ten busiest, and has Stability Watchdog, Performance Auditor and Pattern Analyst review each one. Fix Advisor goes last, so it prescribes after the diagnosis.
You can also run a check yourself whenever you want one: the Intelligence page lists the analysts with when each last ran, and every one of them is on the Built-in tab of the Agents page. They run on the model provider your workspace already has connected.
The inbox
Everything they found, on one screen.
Findings are grouped by the agent they are about, and the groups are ranked by their worst severity: critical, then warn, then info. Filter to one severity to see only what needs a decision.
The inbox stays readable as the analysts keep running: a finding reported again folds into the one before it, and a check that found nothing is tucked out of the way. Open a finding for its detail and its recommendation.
From finding to fix
What is failing, why, and what to change.
For a failing agent, Fix Advisor reads its failure history and its setup and writes the fix: which step fails, the cause, and exactly what to change and where.
Token Optimizer does the same for spend. Whole pages fetched to find three lines, context sent on every step, calls that could be batched: each comes with the change and its expected saving.
What you get
A review that happens whether or not anyone asks.
- Findings with a recommendation. Not a chart to interpret: what the analyst saw, and what to do about it.
- Escalations that alert. A rule can fire when an analyst raises its finding about an agent to warn or to critical, in the app, by email or on a webhook.
- Only your agents. The sweep reviews the agents your team built, never the built-in ones, and skips agents that have not run all week.
In the last 30 days about two dozen of our agents did real work in our own workspace. On top of that, the built-in intelligence agents watch the fleet: failure trends, slow runs, odd tool-call patterns, weak prompts, wasted tokens. Their findings land in one inbox.
Go deeper
How it works, in detail
More in Observe
See what every agent did, what it cost, and when it breaks.
Give one job to an agent this week.
Start with one repetitive workflow, gate the sensitive step behind your approval, and read the first run end to end.