Docs contents
Scripts
Scripts are not generally available yet. They are on for workspaces that have them turned on.
A script is JavaScript that AgentOS runs in a sandbox, with no model in the loop. It calls the tools you enable on it, the same tools an agent uses, and spends no tokens. Use a script for work whose steps are known in advance: fetch, transform, post. Use an agent when a model has to decide what to do.
A script run is a run like any other: it has a timeline, it can wait for approvals, it can be scheduled, and an agent can call it as a tool.
Writing a script
A new script starts from a template: blank, a weekly run report (runs as is), a scheduled post, a webhook relay with approval, or a fan-out to an agent. Each uses built-in tools only, so it runs in any workspace.
A script is a module whose default export takes the run's input and the enabled tools:
export default async function run(input, tools) {
log("posting for", input.day);
const res = await tools.http_request({ method: "GET", url: `https://example.com/posts/${input.day}` });
return { status: res.status, ok: res.ok };
}
inputis the JSON object the run was started with. Describe it with an input schema (JSON Schema): it is also the parameters an agent sees when the script is its tool.tools.<name>(args)calls an enabled tool and resolves with its result. Only the tools you enable are present;ask_humanand control-plane tools are never available to a script.log(...)writes a line to the run's timeline.- What the function returns is the run's output.
There are no imports, no network access and no timers. Everything a script does outside itself goes through a tool, so the same authentication, secrets and approval rules apply as for an agent. A few globals of browsers and Node are absent too: console (use log), structuredClone, TextEncoder, queueMicrotask and Intl. Date reads the workspace's time zone.
Saving checks that the code compiles, with the same interpreter that runs it: a syntax error refuses the save and names its line. Nothing runs during the check.
Under the code, the tool reference lists the saved tools by the names the script calls, with their parameters and a call to copy. A tool that waits for an approval is marked.
Running a script
- Dashboard: open the script and press Run. A script with an input schema gets a form. A run opens its page, which shows its steps as they happen.
- Test run tries it without side effects: tools that write return what they would have sent, reads run for real, and scripts and agents it calls run as tests too. A call that needs an approval still waits, decided on the run's page. Test runs are left out of metrics and the runs list.
- API:
POST /api/v1/rpc/scripts_runwith a workspace key. - SDK:
aos.scripts.run({ scriptId, input })in TypeScript,aos.scripts.run(script_id=..., input=...)in Python. - Schedule: the script's Schedule tab, or the Schedules page. A scheduled run's input also carries
_schedule:due_at, the time it was scheduled for, andlast_run_at, when the schedule last started a run (nullthe first time; a run that could not start does not count). A script that handles only what is new since then needs no state of its own.last_run_atdoes not say that run finished or succeeded, so a script that must never skip data should keep its own cursor._scheduleis reserved: a configured input value of that name is replaced.
curl -X POST https://agentos-ai.dev/api/v1/rpc/scripts_run \
-H "Authorization: Bearer aos_ws_..." \
-H "Content-Type: application/json" \
-d '{ "script_id": "<scriptId>", "input": { "day": "mon" } }'
To run a script at most once for an event, pass idempotency_key (up to 128 characters): a repeated call with the same key answers with the run the first one started (replayed: true) instead of starting another, so a retry or a webhook delivered twice does nothing new. Keys are per script, and the first call's input is the one that ran. The answer is that run as it stands now: still in progress, it says running with no output yet. A paused or archived script still answers a key it has already run.
The answer is { run_id, status, output, duration_ms } with status: "completed". A run that is waiting answers awaiting_approval with approval_ids, or awaiting_child with child_run_ids. A script that throws answers status: "failed" with an error message, which names the line. A paused or archived script cannot be run. A workspace runs a limited number of scripts at once (2 on Solo, 5 on Operations, 10 on Evolution, 25 on Enterprise); a new run past that is refused with RATE_LIMITED until one ends. Pausing a script also turns its schedule off.
A script's run reads as a script everywhere it is observed. Its events are script.run.started, script.run.completed, script.run.failed, script.run.cancelled, script.run.awaiting_approval, script.run.awaiting_child and script.run.resumed, where an agent's run has agent.run.*, next to script.started, script.log and script.error. On the run page its trace starts at a Script run span that shows 0→0 tok, and a call that started a child run shows that child as an agent or a script. A run_failed or run_duration_exceeded alert names it with runnable_kind: "script".
Approvals
Switch a tool from Auto-run to Approval required on the script and every call to it waits for a person. The run stops at that call and resumes once it is decided:
- Approved: the call executes once, and the script continues from where it stopped.
- Rejected: the call does not execute and throws a
ToolRejectedErrorin the script, which the script can catch:
try {
await tools.write_resource({ filename: "weekly-report.md", content: input.report });
} catch (e) {
if (e instanceof ToolRejectedError) return { written: false };
throw e;
}
Batches
Gated calls the script makes together wait together. The run keeps going after a gated call until it cannot go further without one of them, then pauses once, with one approval per waiting call, as an agent asks for the calls of one turn. Each approval is decided on its own; the run resumes when every one of them is decided, and each call gets its own outcome: an approved call executes, a rejected one throws its ToolRejectedError.
// One pause, one approval per lead. Each entry records its own outcome.
const outcomes = await Promise.allSettled(
leads.map((lead) => tools.composio_GMAIL_SEND_EMAIL({ recipient_email: lead.email, subject, body })),
);
- Parallel (
Promise.all,Promise.allSettled, or calls started before any is awaited): one pause for all of them. - Sequential (each call awaited before the next starts): one pause per call, since the next call is not known until the previous one is decided.
- Calls that need no approval, made while gated calls wait, run as usual.
- A batch holds at most 20 gated calls. A script that starts more at once fails; send them in groups of 20 or fewer.
- The outcomes reach the script in the order the calls were made, whichever finished first.
- If any approval of a batch expires, the run ends, as for a single approval.
A script resumes by running again from the top: calls it already made return their recorded results without running again. Keep a script's steps the same from one run to the next. Date, now(), Math.random() and random() are recorded too, so they return the same values on a resume.
Code saved while a run is waiting does not change that run: it finishes the version it started with.
Determinism
A resume replays the script from the top, so a few things behave differently from a single straight run:
- The run's
inputand every tool result reach the script with object keys in sorted order, on the first run and on every replay, so iterating over an object gives the same order each time. Promise.raceandPromise.anysettle in the order the calls were made on a replay, not the order they finished. Do not rely on which call finishes first before a pause.- Calls that do not need approval may run again if a resume is retried, so they run at least once, not exactly once. Make them safe to repeat.
- An approved call runs with the workspace's tool configuration as it is when the run resumes, not as it was when the call was made.
Scripts as tools
Enable a script on an agent or on another script and it becomes a tool named script_<slug>, where the slug comes from the script's name when it is created and never changes. The call starts a child run of the script and resolves with { run_id, status, output }.
If the child run waits (for an approval, say), the caller waits too, and continues with the child's output once the child has ended, within about a minute. A child agent whose every gated action a reviewer rejected ends cancelled, but its caller gets a normal result, { run_id, status: "declined", output }, with what the child returned: the pipeline continues. A calling agent is also told that a person declined the actions. A child that fails, is cancelled, or whose approval expires throws a ChildRunError in a calling script, and reaches a calling agent as an error result.
The same holds for tools.invoke_agent:
const reply = await tools.invoke_agent({ agent_name: "Lead Reply", input: { lead } });
if (reply.status === "declined") return { sent: false, draft: reply.output };
Secrets
Script code is stored and versioned in plain text and every member can read it. Never put a key or a password in a script: store it in a tool (an integration, or a custom tool's secret) and call that tool. The editor warns when the code looks like it holds a key.
Limits
| Limit | Per run |
|---|---|
| Sandbox time (time waiting on tools excluded) | 30 s |
| Total time, tools included | 240 s |
| Memory | 64 MB |
| Tool calls | 200 |
| Log lines | 500, each up to 8 KB, 256 KB in all |
| Arguments of one call / one result | 64 KB each |
| Output | 1 MB |
| Tool arguments and results, all calls together | 2 MB |
| Gated calls waiting for approval at once (one batch) | 20 |
A log line longer than 8 KB is cut off. Passing any other limit fails the run, and its timeline names the limit.
An integration (Composio) call has a lower cap of its own: its arguments may be at most 32 KB. A call between 32 KB and 64 KB passes the script's check and then fails with "Composio tool input too large", so keep a large email body or post under 32 KB.
The number of scripts a workspace can have depends on its plan. Archived scripts do not count, and scripts never use an agent slot.