AI agent observability

What did the agent do? Open the run.

Every run is recorded in full, step by step: each model call, each tool call with its arguments and result, the tokens, the cost and the decisions people made. Hosted agents need no setup; agents you run yourself report through the SDK or OpenTelemetry.

The problem

Agents fail politely.

Traditional software fails with an error. Agents often take a wrong turn halfway through a run, finish anyway, and report something reasonable. If all you kept is the final answer, you cannot say what happened, why, or what it cost.

  • What did it do? Which tools it called, with which arguments, and what came back.
  • Why did it do it? What the model said between steps, where it explains its reasoning.
  • What did it cost? Tokens in and out on every model call, per run and per agent.

The record

Every step, in order, on one page.

Each run keeps what started it, its input, every model call and every tool call paired with its result, every error, and its outcome, with timings and tokens throughout.

The run page reads it several ways: a transcript of the agent's reasoning and actions, the raw event stream filtered to tools, model calls or errors, and a plain-language summary written from the events when you ask for one.

Traces

See where the time and the tokens went.

A trace is the run as a tree of timed steps: the agent, each model call and each tool call, with a duration and a token count on every node. A slow dependency or a step that reads far more than it needs stands out at a glance.

Runs also stream live as they happen, so you can watch an agent work instead of waiting for it to finish.

Any agent

Hosted, or already running somewhere else.

The record is the same wherever the agent runs.

  • Hosted agents: nothing to set up. Agents that run on AgentOS are recorded and traced automatically, every model call and tool call included.
  • Your own code: a few SDK lines. Wrap a run with the TypeScript or Python SDK and emit steps, model calls and tool calls, or post events over HTTP from any language.
  • Already on OpenTelemetry: change the endpoint. Point your OTLP exporter at AgentOS. Each trace becomes a run, and GenAI spans keep their models, tokens and tool calls.

Cost

Cost per run, per agent.

A monthly bill tells you that you spent money. Cost per run tells you which agent spent it, and whether that changed yesterday. Usage breaks spend down by agent and by model, estimated from the tokens recorded on every run, and monthly budgets can be set for the workspace and for single agents.

We tried an agent that searched X every morning for conversations worth joining. It worked, and it was the most expensive agent we had: one search run cost around $0.60, against about $0.02 for publishing a post.

A monthly bill would have hidden that for weeks. Cost per run, per agent, made it obvious right away.

Case studyHow we run AgentOS on AgentOS

What you get

Answers without a meeting.

  • Decisions made by people. Who approved or rejected a step, when, and their note; the question an agent asked and the answer it got.
  • Events you define. Add your own event types with a level (debug, info, warn, error) next to the built-in ones.
  • Alerts on the same data. Failures, error rates, slow runs, agents that go quiet and budgets, sent in the app, by email or to a webhook.

One Tuesday the agent published Wednesday's post, and on Wednesday it published it again. Its time tool returned the date without the weekday, and the model worked out the weekday itself and got it wrong. The output of both runs looked fine; the run's steps showed the mistake.

Case studyEvery post on our LinkedIn page is drafted by an agent and approved by a person

Give one job to an agent this week.

Start with one repetitive workflow, gate the sensitive step behind your approval, and read the first run end to end.