← Blog

Case study · 9 August 2026

What we learned from letting agents do our cold outreach

Outbound was one of the first jobs we gave our own agents. It looked like a perfect fit: repetitive, rule-based, and easy to measure. Some of it was. This is an honest account of what worked, what didn't, and why.

The pipeline

  1. Prospect Finder searches the web for companies that match our ideal customer, reads their site to find a public contact email, skips anyone already on our list, and adds the rest.
  2. Outbound Outreach writes a short first email to each new prospect and sends it. Every send waits for a person to approve it.
  3. Outreach Reply Classifier reads the replies and sorts them by intent: interested, needs a follow-up, and so on.
  4. Lead Reply Agent follows up with the ones who were interested.

What worked: the writing, with a gate

The outreach agent writes well, because we told it exactly what a good first email is: short, in the prospect's language, with no guesses about their business. It asks rather than assumes, and never claims to know what a company's problems are.

An outreach email waiting for approval, with the recipient, subject and full body shown, and Approve and Reject buttons.
Every outreach email waits for a person · Demo workspace

The approval gate earned its place. A person reads every email before it goes out, and that keeps volume honest: when every send needs a person, you send fewer and better emails.

What didn't work: finding the right people

The weak link was the first step. Web search is good at finding companies and bad at finding the right ones. Most of what our prospecting agent found were agencies, not the companies we wanted to reach.

No amount of prompt work fixed it, because the problem wasn't the prompt. It was the source. We could see that quickly, because every search and every result was on the record: the agent was doing exactly what we asked, with inputs that couldn't produce what we wanted.

The downstream agents felt it too. Almost no replies came back, so the reply classifier had very little to classify.

What we would do differently

  • Let a person pick the prospects. Our first real conversations came from founder-led outreach, not from automated prospecting. Agents are good at the writing and the bookkeeping; deciding who is worth writing to needed a person.
  • Fix the source before automating more. A proper contact-data source would change the equation. Without one, automating the search just produces more of the wrong list, faster.
  • Keep the gate. Every email still waits for approval.

The general lesson

An agent can only be as good as its inputs. When an agent's results are poor, read its runs before you rewrite its prompt: often it is doing its job perfectly on material that was never going to work.

See your own agents like this

AgentOS records every run, pauses risky actions for approval, and tells you when an agent breaks.