The Delegation Boundary
How to decide what an AI agent may do alone, when a person steps in, and how that boundary moves as evidence changes. Part 3 of 7.

AI Native Operating Model · Part 3 of 7
Most organisations already have some kind of delegation boundary between people and AI agents, even if nobody has formally designed it.
It usually emerges from a series of practical decisions. An approval is added because a new agent has not yet earned much trust. A permission is opened up to make a pilot possible. Another review step is introduced after an incident. Important actions gradually accumulate extra checks, while routine actions inherit controls that may no longer be necessary.
Each decision can be reasonable on its own. The difficulty is that, taken together, they often become an operating model by accident. People end up reviewing work that carries very little risk, while more consequential decisions can move through processes that were never explicitly designed around their impact.
That is why delegation needs to be treated as something deliberate. The goal is not to automate everything that can technically be automated. It is to give agents enough autonomy to be useful, while keeping the right controls around the work where mistakes are difficult to reverse or require genuine judgment.
Authority is set action by action rather than granted to an agent as a whole, and the boundary moves in both directions as evidence changes.
What the delegation boundary needs to define
Describing an agent as "autonomous" is too broad to be useful because autonomy depends entirely on what the agent is being allowed to do.
An onboarding agent, for example, might be allowed to analyse approved support records without asking anyone. It might also be allowed to build a prototype inside an isolated test environment. Releasing that change could still require approval, while changing the underlying eligibility rules might remain entirely with the product owner.
The same agent can therefore operate with several different levels of authority at once. The useful question is not whether the agent is autonomous, but what authority it has for each class of action.
There are five things that need to be clear.
The first is what the agent may do without asking. This should be explicitly defined and is usually narrower than the set of actions the agent is technically capable of performing.
The second is when the agent has to stop and escalate. These conditions need to be concrete enough for the system to recognise them. "Use good judgment" is not a useful control, whereas a change affecting a payment path, an exception above an agreed threshold, or work outside a validated scope can be detected and handled consistently.
The third is who receives the escalation. The role needs to be clear, and someone must actually hold that decision right. An approval process that does not have a clear owner simply creates another queue.
The fourth is what information should accompany the request. The person making the decision needs enough context to make it properly. Depending on the situation, that might include test results, expected impact, affected systems, alternatives considered and a rollback plan. An agent saying only that the work is complete and asking for approval pushes the decision to a person without giving them the evidence they need.
The fifth is what the system should do while it waits. In some cases it should pause completely. In others it may be able to continue unrelated work, prepare alternatives or fall back to a safe default. This behaviour should be agreed beforehand rather than discovered during the first incident.
Only part of this is really about the agent itself. Ownership, evidence, escalation and waiting behaviour are questions about how the organisation operates around the agent, which is why delegation belongs in the operating model as much as it belongs in the technical platform.
Where possible, the limits should also be enforced by the environment. A restricted account, isolated environment or deployment permission is a stronger control than a sentence in a prompt telling an agent not to do something.
What human intervention tells you about the system
Whenever a person has to step in, it is worth understanding why they were needed. Human involvement is not automatically a sign that the system has failed. In some cases, the escalation is exactly what the system was designed to do.
There are usually three broad reasons for an intervention.
The first is that the decision genuinely required judgment. Imagine an agent proposes an onboarding change that improves the customer experience but increases the workload for support. There may be no technical rule that settles whether the trade-off is worthwhile. Someone with the appropriate authority needs to decide.
In that case, the intervention should remain. The improvement is not to remove the person, but to make sure the decision reaches the right person with enough evidence to make it well.
The second reason is that the surrounding system was missing something. The agent may not have been able to find an agreed policy, so someone had to supply it manually. The available tests may not have been good enough to determine whether the result was correct, forcing somebody to review the output line by line. Low-risk and high-risk actions may have been bundled behind the same permission because the system could not distinguish between them.
Those interventions expose weaknesses in the delivery system. Better context, stronger tests, narrower permissions or more reliable rollback can often remove the need for the same intervention in future.
The third reason is that the organisation has deliberately chosen to keep a decision with a person. Some actions are difficult to reverse, have consequences that extend far beyond the immediate change, or simply involve decisions the organisation does not want to delegate. Those boundaries are entirely reasonable as long as they are explicit.
An intervention can reveal more than one of these at the same time. A specialist might make a genuine judgment while also discovering that the agent was missing important information. Fixing that missing information improves the system, but it does not make the judgment itself unnecessary.
The behaviour of blocked agents can reveal similar problems. If an agent cannot turn a ticket into tests because the ticket is ambiguous, the weakness is in planning rather than implementation. If it repeatedly exhausts its attempts because a task is too large, the problem may be the way the work was scoped. Limits on attempts are useful partly because hitting them produces information about the system.
The important thing is to record why the intervention happened, what the person contributed and whether anything should change before the same kind of work appears again. That information is much more useful than simply measuring how often humans were involved.
Turning intervention rates into a performance target would create the wrong incentive. A lower number of escalations is only useful if the system has actually become safer and more capable, rather than simply becoming less likely to ask.
Why supervision should follow consequence rather than size
One of the easiest mistakes to make is using the size of a change as a proxy for how much supervision it needs.
A one-line change to an authentication rule can carry far more risk than hundreds of lines of generated test fixtures. Similarly, a small irreversible data migration may deserve more scrutiny than a much larger configuration change that can be rolled back immediately.
The level of supervision should therefore follow the possible consequence of a wrong decision and the organisation's ability to recover from it.
This is an established principle in reliability engineering. Google's SRE practice makes a similar point about on-call pages: every page should be actionable, and every response should require intelligence.1 The same idea applies to approval gates. Human attention should be used where a real decision is required, rather than spent confirming routine outcomes that could be checked automatically.
Classifying work by consequence also allows low-risk changes to move quickly without lowering the bar for the work that actually matters. Without that distinction, an organisation can end up applying the controls required for its most dangerous changes to almost everything else.
The classification itself should be as deterministic as possible. Rules based on the systems being touched, the type of data involved, the environment or the potential impact are preferable to relying entirely on a model's judgment. A model can provide an additional signal, but asking one model to decide whether another model's work is risky does not create a genuinely independent control.
This leads to a broader point about the role of people in the process. If someone repeatedly sits at a gate only to confirm that a change which is almost always acceptable is, once again, acceptable, then the organisation probably has a validation problem rather than a governance solution. The better response is to automate the check properly or make an explicit decision about the risk.
Why another agent is not always an independent check
Adding a second agent can look like an easy way to create redundancy. One agent produces the work, another reviews it, and the assumption is that the second will catch errors made by the first.
The weakness in that approach is that both agents may fail for the same reasons.
Software engineering has explored this problem before. Knight and Leveson asked students at two universities to independently build 27 versions of the same program. The versions failed together much more often than independence would predict because the difficult parts of the specification were difficult for multiple teams in similar ways.2
AI agents can share even more of the same assumptions. They may use the same model family, similar prompts, the same tools, the same source material and similar context. A second agent can still be useful, particularly if it receives fresh context and did not participate in the original generation process, but it should not automatically be treated as an independent form of verification.
For consequential work, a stronger approach is to combine different kinds of evidence. That can include deterministic tests, a different validation method, properties that can be asserted independently, observations from a limited rollout, or review by a person with information the agents do not have.
Acceptance tests written before implementation are a good example. If the building agent cannot edit those tests and does not decide whether it has passed them, the pipeline becomes an independent check against a predefined expectation.
That also keeps the evidence honest. A green pipeline tells you that the behaviour covered by the tests worked. It does not establish that every behaviour outside those tests is safe.
Autonomy should expand and contract
Greater autonomy should follow evidence rather than precede it. Early versions of a system will often involve more human involvement because the organisation is still learning where the agent performs reliably, which failures matter and which technical controls are strong enough to replace manual checks.
As tests improve, rollbacks become reliable and teams gather evidence about particular classes of work, some of those approvals can be removed.
One team we worked with approached this at repository level. They allowed autonomy over documentation first, then tests, then implementation. Broader authority was introduced only after the previous level had proved reliable, and each expansion required explicit approval.
The objective was not to reach a fully autonomous "dark factory". It was to increase autonomy selectively where the evidence justified it.
Consider a routine interface change that currently requires someone to approve the release. If the team introduces representative tests, constrains the rollout and demonstrates that rollback works reliably, future changes in that class may no longer need the same manual gate.
That evidence applies to that class of work under those conditions. It does not provide a general argument for giving the same agent authority over identity policies, payment logic or other unrelated areas.
The same logic also needs to work in reverse. A model can change, an agent can begin operating outside the area it was validated on, volume can increase beyond what existing monitoring can reasonably cover, or a near miss can reveal that earlier validation was weaker than expected.
In those situations, the sensible response may be to reduce autonomy until confidence has been rebuilt.
For that reason, the conditions that reduce autonomy should be defined in advance. A consequential model change, expansion of scope, an exception rate crossing an agreed threshold, a near miss or a regulatory change are all examples of events that could trigger additional supervision.
Agreeing those conditions while the system is operating normally is much easier than debating them in the middle of an incident.
Waiting is part of the design
A delegation model is incomplete if it explains whom the agent should ask but says nothing about what happens while the agent waits for an answer.
The appropriate behaviour depends on the type of work. The agent may be able to continue with unrelated tasks, prepare alternatives, fall back to a safe state or stop completely. That behaviour should be part of the design of the workflow.
The expected response time also matters. If meaningful harm can occur faster than a person can reasonably understand the situation and respond, a human approval step alone is not a sufficient control. The system needs stronger containment or a different division of authority.
After release, it is useful to make a similar distinction between safety and judgment.
An automated system can determine that a rollout is failing, that an important metric has deteriorated or that the outcome cannot yet be established. If a predefined safety condition is breached, the system can roll back a release or disable a feature without waiting for someone to approve the response.
What happens after that may still require judgment. Deciding whether to retry, redesign the change, investigate further or abandon the idea is a different kind of decision.
The practical relationship between speed and safety is therefore fairly straightforward. Better validation, observability and reversibility make it possible to give agents more authority without relying on blind trust.
What to try
A useful first step is to look at the interventions already happening rather than trying to design the entire model from scratch.
For four weeks, record why a person was needed whenever an agent escalates. Keep the categories simple: genuine judgment, something missing from the system, or a limit that should remain. Patterns usually appear quickly, and a relatively small number of missing capabilities or unclear rules often account for a large share of the interruptions.
At the same time, define the conditions that would cause autonomy to be reduced. Writing those conditions down before they are needed makes it much easier to respond consistently when the system changes.
It is also worth choosing one existing approval gate and reclassifying the work that passes through it based on consequence rather than size. Look at how much of the work genuinely requires human judgment and how much could be handled through stronger automated checks.
Finally, review places where one agent is used to check another. The useful question is not simply whether there are two agents involved, but whether the second check introduces genuinely different evidence.
The aim is not to minimise human involvement. It is to use human attention where it contributes something the system cannot reliably provide on its own, while improving the system around everything else.
The next post looks at how these rights, decisions and outcomes can be made visible across the organisation.
Sources
- Betsy Beyer and colleagues (eds), Site Reliability Engineering, O'Reilly, 2016, Chapter 6, "Monitoring Distributed Systems".
- John Knight and Nancy Leveson, "An Experimental Evaluation of the Assumption of Independence in Multiversion Programming", IEEE Transactions on Software Engineering, 1986.
Background: Pynadath, Scerri and Tambe, "Towards Adjustable Autonomy for the Real World", Journal of Artificial Intelligence Research, 2002, discusses transfers of control, including the cost of waiting for a person, as design choices.
Previously in the series: The Hybrid Team.
Next in the series: The AI Native Control Room.
Where did your constraint actually move? We run value-stream assessments that find the real bottleneck in hybrid delivery, then help design the operating model around it. info@re-cinq.com
Your next read

Towards an AI Native Operating Model
An AI native operating model is the system by which an organisation turns intent into outcomes using people and machines together. Part 1 of 7.
Pini Reznik9 mins readSubscribe to Our Bi-Weekly AI Native Newsletter
Lessons from our client work, engineering deep-dives, and the AI Native research worth reading.
Bi-weekly. No spam, unsubscribe anytime.






