Case study · 27 September 2026
How we run AgentOS on AgentOS
We build a platform for running AI agents at work, so we run our own company on it. In the last 30 days about two dozen of our agents did real work in our own workspace: answering support email, finding and writing to prospects, welcoming new signups, drafting and publishing our social posts, and turning product feedback into GitHub issues.
This is what that fleet looks like, and the four things it taught us that we did not expect.
The fleet
We group agents by the job they do. Most jobs are a small pipeline: one agent hands work to the next, and a person steps in where it matters.
| Job | Agents | How it stays safe |
|---|---|---|
| Support | Email Triage labels the inbox, Support Reply answers routine questions | Sensitive or ambiguous threads are left as drafts for a person |
| Outreach | Prospect Finder, Outbound Outreach, Outreach Reply Classifier, Lead Reply Agent | Every first outreach email waits for approval |
| Onboarding | New Signup Monitor hands each new account to Welcome & Onboarding | Every email is built from one fixed template |
| Social | Social Content drafts the week, Daily LinkedIn Post and Daily X Post publish | Every scheduled post waits for approval |
| Product | Linear Feedback Reader promotes feedback, GitHub Issue Creator files it | Issues are filed for the team to triage, not acted on |
On top of that, the built-in intelligence agents watch the fleet: failure trends, slow runs, odd tool-call patterns, weak prompts, wasted tokens. Their findings land in one inbox.
1. Approval gates work. Approval fatigue is the real risk.
Every scheduled post on our LinkedIn page and every first outreach email waits for a person to approve it. That part works exactly as intended: none of them goes out until someone has read it.
What we underestimated is the other side. An approval that nobody answers expires after 24 hours, and the run ends without posting. Between 18 August and 25 September, 17 of the 29 weekday LinkedIn posts expired that way. The agent did its job; the humans did not show up for theirs.
The lesson is not to remove the gate. It is to ask for fewer, better decisions: gate only what leaves the building, batch related calls into one decision, and post less often. We are moving to three posts a week.
2. Read the run, not just the output
One Tuesday our LinkedIn agent published Wednesday's post, and the next day it published it again. The output looked fine both times: "post prepared and submitted for review".
The run told the real story. The time tool returned the date without the weekday, and the model worked out the weekday itself and got it wrong. We only saw that because every tool call and its result is on the record, in order. The fix we are making is to key each day's post by its date instead of asking the model to reason about weekdays.
3. Watch cost per run, not per month
We tried an agent that searched X every morning for conversations worth joining. It worked, and it was the most expensive agent we had: one search run cost around $0.60, against about $0.02 for publishing a post. Reading is priced roughly thirty times higher than writing on that API.
A monthly bill would have hidden that for weeks. Cost per run, per agent, made it obvious right away. The agent is paused until reading is worth that price to us.
4. Agents are honest about bad inputs, if you look
Our prospecting agent searches the web for companies that match our ideal customer, verifies a public contact email, and adds them to a sheet. The leads it finds are mostly agencies, which is not who we want.
No amount of prompt work fixes a weak source. The runs made that clear quickly, because the tool results show exactly what came back from each search. Our first real traction came from founder-led replies, not from automated outbound.
What we would tell a team starting out
- Start with one job that already has a clear owner, like the support inbox.
- Gate the actions that leave the building: sending, posting, paying, deleting. Leave the rest to run.
- Make someone responsible for the approval queue, or it will expire on you.
- Look at runs every week, not just outputs. The interesting failures do not throw errors.
- Track cost per run from day one.
Everything in this post happened in our own workspace, on the same product you can use.
See your own agents like this
AgentOS records every run, pauses risky actions for approval, and tells you when an agent breaks.