Releases
What's new in AgentOS
What changed in the AgentOS platform and its TypeScript and Python SDKs, newest first. RSS
Upload many files at once, and file them into folders
- Resources takes many files at a time, from the Upload button or by dragging them in, and shows progress as they go.
- Drop files onto a folder to upload them straight into it, or drag a file onto a folder to move it there.
- LinkedIn posts published by agents keep their links and hashtags intact.
A public site for the product
- New pages at /product walk through every capability, backed by a blog of guides and case studies and a Company section with About and Security pages.
- Press / or Ctrl+K anywhere on the site to search docs, guides, product pages and release notes from one place.
- The MCP server has its own docs page, with setup steps for Claude Code, Cursor and Claude Desktop.
- Agents can post images to LinkedIn straight from workspace Resources, and HTML email previews now match the look of the plain-text ones in both themes.
- Alert rules created through the API or an agent now fire reliably, including their webhooks.
- Web requests made by agents now recheck every redirect they follow, not just the first address, closing a gap that could reach internal network addresses.
A new look across the product
- The product shares the website's visual language: clear labels, mono figures and quieter surfaces.
- Approvals read as one request: "Outbound Outreach wants to send 3 emails", the first call previewed, then Approve all, Review one by one or Reject all.
- The run page reads like a record. The transcript is a log with tool calls and their input and output inline, and an agent's markdown renders, tables included.
- Agent pages load faster: model lists and integration schemas are cached instead of fetched on every visit.
Every action shows its progress
- Pausing or activating an agent shows its progress and no longer freezes the page.
- Every button, switch and control that saves something shows that it is working, and you can keep navigating while it does.
- When something goes wrong, the message no longer exposes internal details. It ends with a short reference you can quote to support.
Previews for social posts
- Twitter and LinkedIn posts waiting for approval show the text, visibility, any attached link and the character count against the network's limit.
- Email previews look closer to how a mail client shows them.
- The approvals list stays in place while a decision is in flight, and its buttons show their progress.
See the action before you approve it
- Approvals preview what a call will do: an email with its recipients, subject and body, a file with its contents. Arguments a preview does not show are still listed beneath it.
- Email bodies render in a sandbox, so opening an approval never loads a tracking pixel.
- Every run gets a Summary tab: generate a short account of what happened, plus the list of actions it took, rejected ones marked.
- A paused turn shows all of its calls together, with Approve all and Reject all. Reject the exception on its row, then approve the rest.
Accurate 7-day health cards
- The 7-day run and success cards on Runs are counted live, so they never show stale figures.
- Reasoning models get enough output budget to finish: their maximum completion tokens never drop below 16,384, so long replies are no longer cut off.
Approvals are decided call by call
- When an agent pauses several gated calls in one turn, approving one of them runs only that call. Each call needs its own decision.
- Select or clear every call of a turn at once from its chip.
- A decision stays on screen, with a loading indicator, until it is recorded.
- Conversations, evals and folder changes respect each member's folder access again.
Security fixes for the control plane
- A member's personal key can only reach the agents and folders that member has access to, including when moving an agent between folders.
- Workspace keys can no longer invite or remove members, and control-plane changes are recorded in the audit log like dashboard changes.
- A rate-limited control-plane call answers
429withRetry-After, so clients can back off. - TypeScript SDK 0.4.1 keeps the API key out of logged or serialised runs, tasks and spans. Upgrade, and rotate any key that may have been logged.
The control plane over HTTP
- Every workspace operation the MCP server offers is now also an HTTP endpoint:
POST /api/v1/rpc/<method>, with the same keys, workspace resolution and role checks. - 56 operations across 18 areas: agents, runs, approvals, evals, schedules, resources, alerts, reports, tools and more.
- The TypeScript and Python SDKs expose them as typed methods from version 0.4.0.
- A new control plane reference documents every method.
- Every workspace operation the MCP server offers is now also an HTTP endpoint:
Security
- The API key no longer appears when a run, task or span is serialised with
JSON.stringifyor printed withutil.inspect. In 0.3.0 and 0.4.0, logging or returning one of these objects exposed the key. Upgrade, and rotate any key that may have been logged. - Redirects are no longer followed. A redirect fails with an
AgentOSErrorcodedREDIRECT; setbaseUrlto the final address. - A
baseUrlthat uses plainhttp://for anything other than localhost logs a warning, since the key would travel unencrypted.
- The API key no longer appears when a run, task or span is serialised with
Added
- Control plane. Every workspace operation as a typed method:
aos.agents,aos.runs,aos.approvals,aos.evals,aos.schedules,aos.resources,aos.alerts,aos.reports,aos.tools,aos.folders,aos.conversations,aos.members,aos.apiKeys,aos.notifications,aos.integrations,aos.marketplace,aos.fleetandaos.workspaces(56 methods). Generated from the server's schemas; see the reference. workspaceoption andAGENTOS_WORKSPACE, for control-plane calls with a personal key.
Changed
- The README is rewritten as a complete guide.
- Control plane. Every workspace operation as a typed method:
Breaking
invoke()returns{ runId, status, output, durationMs, sessionId, approval, inputRequest }.statusis"completed","awaiting_approval"or"awaiting_input"; a paused run used to look like a success with an empty output.outputisnullwhile paused, and typed with a generic (invoke<T>()).emitLLMResponse()takes token usage in camelCase:{ inputTokens, outputTokens }.- The LangChain handler records a span tree instead of tasks and events.
- Node 20 or later.
Added
- Configuration from the environment:
AGENTOS_API_KEY,AGENTOS_AGENT_ID,AGENTOS_BASE_URL,AGENTOS_DISABLED.agentIdis optional and can be passed per call, for workspace keys. invokeStream(),reply(), and the full invoke API:sessionId,startSession,history,externalUserId,parentRunId,test.withRun(), which completes or fails a run around a function.run.cancel(),run.complete({ usage })for token counts, andspan.fail().client.flush()andclient.shutdown(). In Node, queued events are also sent automatically before exit.- Retries with backoff for transient failures; a 429's
Retry-Afteris honoured. Invokes are only retried on 429, since the agent may already be running. AgentOSTransientError,AgentOSConfigError,statuson every error, andretryAfteronAgentOSRateLimitError.- Options:
maxQueueSize,maxRetries,timeoutMs,onError,fetch. startRun({ parentRunId }).
Fixed
- A failed background flush could crash the Node process with an unhandled rejection.
- Events were lost silently when a send failed. Transient failures now stay queued and retry; rejected batches are dropped and counted in
run.droppedEvents. - The event queue is bounded. When it fills, debug and info events are dropped before warnings and errors.
- The LangChain handler reported every model as
"llm"and every tool as"tool", and made a network call on each chain step. - The package declares
"type": "module"and separate type declarations forimportandrequire. Its ES module build was being loaded as CommonJS, which can breakimporton older Node versions.
Added
- Control plane. Every workspace operation as a typed method with keyword arguments:
aos.agents,aos.runs,aos.approvals,aos.evals,aos.schedules,aos.resources,aos.alerts,aos.reports,aos.tools,aos.folders,aos.conversations,aos.members,aos.api_keys,aos.notifications,aos.integrations,aos.marketplace,aos.fleetandaos.workspaces(56 methods). Generated from the server's schemas; see the reference. workspaceargument andAGENTOS_WORKSPACE, for control-plane calls with a personal key.
Changed
- The README is rewritten as a complete guide.
- Control plane. Every workspace operation as a typed method with keyword arguments:
Breaking
invoke()returns anInvokeResult(run_id,status,output,duration_ms,session_id,approval_id,input_request) instead of a dict.statusis"completed","awaiting_approval"or"awaiting_input"; a paused run used to look like a success with an empty output.emit_llm_response()takes token counts asusage=(wastoken_usage=).- The LangChain handler records a span tree instead of tasks and events.
- Licence is MIT (was Apache-2.0), matching the TypeScript SDK.
Added
- Configuration from the environment:
AGENTOS_API_KEY,AGENTOS_AGENT_ID,AGENTOS_BASE_URL,AGENTOS_DISABLED.agent_idis optional and can be passed per call, for workspace keys. invoke_stream(),reply(), and the full invoke API:session_id,start_session,history,external_user_id,parent_run_id,test.Runis a context manager: completed when the block ends, failed with the traceback if it raises.run.cancel(),run.complete(usage=...)for token counts, andspan.fail().client.flush()andclient.shutdown(). Queued events are also sent automatically at interpreter exit.- Retries with backoff for transient failures; a 429's
Retry-Afteris honoured. Invokes are only retried on 429, since the agent may already be running. AgentOSTransientError,AgentOSConfigError,statuson every error, andretry_afteronAgentOSRateLimitError.- Options:
max_queue_size,max_retries,on_error,http_client. start_run(parent_run_id=...).invoke_stream()yields typed events (InvokeStreamEvent) and skips event types it doesn't know.- Ships
py.typed, so type-checkers use the SDK's type hints.
Fixed
invoke()timed out after 10 s, while hosted runs can take 300 s.emit()could send a batch synchronously on the caller's thread. Sending now always happens in the background.- Events were lost silently when a send failed. Transient failures now stay queued and retry; rejected batches are dropped and counted in
run.dropped_events. - The event queue is bounded. When it fills, debug and info events are dropped before warnings and errors.
- Payloads with values JSON can't encode (datetimes, SDK objects) are sent as text instead of failing.
- A span with no parent or no kind sent those fields as
null, which the server rejects, so the whole batch of events was lost. They are now left out. - The LangChain handler reported every model as
"llm"and every tool as"tool", and made a network call on each chain step.
SDK 0.3.0 support, and run cancellation
- The platform side of SDK 0.3.0 for TypeScript and Python: typed
invoke()results, streaming, sessions and retries. See the SDK entries for the full list. - New
POST /api/ingest/v1/run/cancelcancels a run and any human-input request it is waiting on. Calling it twice is safe. - An agent-scoped key can only complete, fail, add tasks to or pause its own agent's runs, never another agent's in the same workspace.
- The SDK documentation pages have their own descriptions and working tracing examples.
- The platform side of SDK 0.3.0 for TypeScript and Python: typed
Resources show when they last changed
- Every resource records when it was last modified, whichever way it was written: dashboard, API, MCP or an agent.
- The resource viewer shows that date, opens as a modal on desktop and no longer overflows its box.
- Help, alert and prompt-variable copy now describes the product as it is, and the user guide has been rewritten.
The Operator can read what it supervises
- The Operator now reads a run's full event timeline, so asking why a run failed gets the actual tool calls and errors, not just the error message.
- It also reads fleet health, the real list of built-in and integration tools, alerts, the approval queue and members. Fleet health is available to MCP clients as the
fleet_healthtool. - Evals run to a verdict in one call: pass, fail, or error with the reason. The Operator can run them to check that a change helped.
- Modals stand out from the page behind them, and Run agent, Run pipeline and the folder picker open as modals.
- Prompt variables and parent run ids on
invoke().
- Prompt variables and parent run ids on
- Prompt variables and parent run ids on
invoke().
- Prompt variables and parent run ids on