← Blog

Case study · 27 September 2026

How we run AgentOS on AgentOS

We build a platform for running AI agents at work, so we run our own company on it. In the last 30 days about two dozen of our agents did real work in our own workspace: answering support email, finding and writing to prospects, welcoming new signups, drafting and publishing our social posts, and turning product feedback into GitHub issues.

This is what that fleet looks like, and the four things it taught us that we did not expect.

The fleet

We group agents by the job they do. Most jobs are a small pipeline: one agent hands work to the next, and a person steps in where it matters.

JobAgentsHow it stays safe
SupportEmail Triage labels the inbox, Support Reply answers routine questionsSensitive or ambiguous threads are left as drafts for a person
OutreachProspect Finder, Outbound Outreach, Outreach Reply Classifier, Lead Reply AgentEvery first outreach email waits for approval
OnboardingNew Signup Monitor hands each new account to Welcome & OnboardingEvery email is built from one fixed template
SocialSocial Content drafts the week, Daily LinkedIn Post and Daily X Post publishEvery scheduled post waits for approval
ProductLinear Feedback Reader promotes feedback, GitHub Issue Creator files itIssues are filed for the team to triage, not acted on

On top of that, the built-in intelligence agents watch the fleet: failure trends, slow runs, odd tool-call patterns, weak prompts, wasted tokens. Their findings land in one inbox.

1. Approval gates work. Approval fatigue is the real risk.

Every scheduled post on our LinkedIn page and every first outreach email waits for a person to approve it. That part works exactly as intended: none of them goes out until someone has read it.

What we underestimated is the other side. An approval that nobody answers expires after 24 hours, and the run ends without posting. Between 18 August and 25 September, 17 of the 29 weekday LinkedIn posts expired that way. The agent did its job; the humans did not show up for theirs.

The lesson is not to remove the gate. It is to ask for fewer, better decisions: gate only what leaves the building, batch related calls into one decision, and post less often. We are moving to three posts a week.

2. Read the run, not just the output

One Tuesday our LinkedIn agent published Wednesday's post, and the next day it published it again. The output looked fine both times: "post prepared and submitted for review".

The run told the real story. The time tool returned the date without the weekday, and the model worked out the weekday itself and got it wrong. We only saw that because every tool call and its result is on the record, in order. The fix we are making is to key each day's post by its date instead of asking the model to reason about weekdays.

3. Watch cost per run, not per month

We tried an agent that searched X every morning for conversations worth joining. It worked, and it was the most expensive agent we had: one search run cost around $0.60, against about $0.02 for publishing a post. Reading is priced roughly thirty times higher than writing on that API.

A monthly bill would have hidden that for weeks. Cost per run, per agent, made it obvious right away. The agent is paused until reading is worth that price to us.

4. Agents are honest about bad inputs, if you look

Our prospecting agent searches the web for companies that match our ideal customer, verifies a public contact email, and adds them to a sheet. The leads it finds are mostly agencies, which is not who we want.

No amount of prompt work fixes a weak source. The runs made that clear quickly, because the tool results show exactly what came back from each search. Our first real traction came from founder-led replies, not from automated outbound.

What we would tell a team starting out

  • Start with one job that already has a clear owner, like the support inbox.
  • Gate the actions that leave the building: sending, posting, paying, deleting. Leave the rest to run.
  • Make someone responsible for the approval queue, or it will expire on you.
  • Look at runs every week, not just outputs. The interesting failures do not throw errors.
  • Track cost per run from day one.

Everything in this post happened in our own workspace, on the same product you can use.

See your own agents like this

AgentOS records every run, pauses risky actions for approval, and tells you when an agent breaks.