← Blog

Guide · 16 August 2026

Human-in-the-loop without the bottleneck

"Human-in-the-loop" usually means one thing: a person approves what the agent is about to do. That is one of two ways to bring a person in, and using only that one is how approval queues end up full of decisions nobody wants to make.

Two different moments

Approving an action. The agent has decided what to do and needs sign-off before it does it: send this email, post this text, issue this refund. The person sees the exact action and says yes or no. This is an approval gate.

Answering a question. The agent can't decide on its own and asks. "This customer mentions a chargeback. Do you want to take it, or should I send the standard billing reply?" The person answers, and the agent carries on with that answer.

The first protects you from a wrong action. The second stops the agent from guessing in the first place.

When the agent should ask

An agent should ask when the right answer depends on something it can't know from the input:

  • Judgment calls. A complaint that might need a manager, or might not.
  • Policy exceptions. A refund outside the usual terms.
  • Missing information. Which account, which date, which of two customers with the same name.
  • High stakes with no clear rule. Anything involving money, legal language or a threat to leave.

Write those situations into the agent's instructions, together with what to do when there is a clear rule. The goal is an agent that asks rarely, and always about something that deserves a person.

Make the answer easy

A good question comes with its context and with choices. "What should I do?" costs the person a minute of reading; two buttons cost them a second.

A support pipeline's timeline: email triage and the reply agent completed, then a question from the reply agent about a customer mentioning a chargeback, with two choices.
The reply agent asks about a chargeback, with two choices · Demo workspace

Keep the run moving

While the agent waits for an answer, the run pauses. When the person answers, the run resumes with the answer as if the agent had known it all along. Work that doesn't depend on the answer can be done before the question is asked, so the person only holds up what truly needs them.

How it works in AgentOS

  • Give a hosted agent the built-in ask_human tool. It takes a question, optional context, and up to eight choices.
  • When the agent asks, the run pauses and waits for input. The question shows on the run, and on the pipeline's timeline when the agent is part of one.
  • Answer in the dashboard, and the run resumes with the answer.
  • For actions rather than questions, gate the tool instead: see approval gates.

See your own agents like this

AgentOS records every run, pauses risky actions for approval, and tells you when an agent breaks.