Agentic AI Architecture: The Intent Plane and the Execution Plane
How to separate requirements, planning and review from the isolated environments that run coding agents.

A typical agentic AI diagram shows a model receiving context, using tools, storing information and feeding the result back into the next step. This loop leaves out much of the architecture needed to run coding agents on real repositories. Somebody still has to turn the original request into a clear and manageable task, and the platform then needs a safe place to run it.
Planning a task and running it place very different demands on the platform, so we handle them separately. People and agents may spend days discussing and editing a plan before they agree on what should be built. The resulting implementation issues then move into environments designed to start quickly, support many parallel jobs and disappear when the work is finished.
While building Lore, re:cinq's internal software factory platform, we had to decide where planning should end, where execution should begin, and why some tasks repeatedly fail even when the model can perform the individual steps. The architecture we describe here came from those decisions. We've seen similar approaches in recent research and open-source projects.
Two-plane agentic AI architecture: an intent plane where a request becomes a plan, a validated spec and issues, and an execution plane where each approved issue becomes one ephemeral agent run that produces a pull request.
Why agent planning and execution need separate systems
When a request first arrives, it usually lacks enough detail for an agent to implement safely. The team uses the intent plane to fill those gaps, discuss trade-offs and agree on the expected result. Approval starts the work of turning that plan into a reviewed specification and then into implementation issues. Each issue can run in its own isolated job. Once a job finishes, we shut down the environment it ran in.
If planning and execution live in the same application, a routine deployment can terminate work that is already running. We hit this in the first months: every push to main restarted the service that was also running the agent, and tasks died mid-run until we moved them into separate, short-lived jobs. The interface used to manage the work can now restart without affecting those jobs.
We hand work from one system to the other only after the plan has been approved. Until then, the team is still deciding what to build. Approval hands the plan to a fixed sequence of agent runs, each with its own timeout and turn cap. Once the resulting specification is merged, it is split into issues that run as separate jobs. Every run traces back to the plan, which makes the trail easy to follow and lets either system change independently.
Why code generation is only half of the architecture
Although most agent tooling focuses on the moment code is produced, the surrounding work receives far less attention. Scaffolds, harnesses, sandbox runtimes and review bots are widely available, while fewer tools help teams prepare the instruction or manage the specifications and other artefacts left behind.
Across our client work, projects rarely get stuck because the agent cannot generate code. They get stuck because the team has not agreed on what should change, which existing behaviour must remain or what would prove that the work is complete. Technically acceptable code still produces the wrong result when the instruction itself is incomplete.
As the platform creates a new specification for each feature, it also accumulates documents that conflict with one another or have been replaced without being retired. The system needs to know which information is still current, and that lifecycle is much easier to establish early than to reconstruct after hundreds of specifications already depend on it.
What an intent plane needs to support
Before proposing a plan, the agent should read the relevant code, earlier specifications and what the team has recorded about similar work. This grounds the draft in the actual repository and allows it to point out missing information alongside the proposed implementation. The questions it raises are often as valuable as the draft itself because they expose what an experienced engineer would want clarified before accepting the work.
Feature plans are usually shaped through discussion, which is difficult to reproduce when one person owns the document and everyone else can only leave comments. Live editing and conflict resolution make the plan behave more like a shared working session. Building that experience takes real engineering, although collaborative-document libraries already solve many of the underlying problems.
When someone asks for one section to be refined, the agent should revise it in place and check it against the comments and the rest of the plan. Regenerating the entire document after every edit costs more and gives unrelated details another chance to change.
Before anyone approves a plan, an agent should compare it with the current codebase and flag anything the implementation cannot support. This will not prove that the proposal is correct, but it will catch obvious disagreements between the document and the code while they are still cheap to resolve.
Once approved, the plan is developed into a specification and opened as a pull request in an open format such as Spec Kit's. Engineers can review it with the same tools they already use for code, and version control preserves the decisions behind it. If a later comment changes what the specification means, the approval is cleared and the plan goes through review again.
A product owner may need to correct an acceptance criterion without having permission to change a datastore decision in the technical section. Role-based editing can enforce that distinction, but we would wait until the basic workflow has proved useful before designing a detailed permission model around it.
Why every coding agent run should be an isolated job
Because a coding-agent run starts, works for a few minutes or perhaps an hour, and then stops, it does not behave like a service that stays online behind an endpoint. Nor is it a stable member of an ordered group. In Kubernetes terms, it fits neither a Deployment nor a StatefulSet particularly well.
The Agent Sandbox project from Kubernetes SIG Apps addresses this mismatch with its own resource, claim and warm-pool model. Its documentation explains that agent runs do not fit the replicated model of a Deployment or the stable, numbered model of a StatefulSet. You can approximate the setup with a single-replica StatefulSet, a Service and a volume, but that adds machinery for the wrong abstraction. Agent Sandbox also relies on runtimes such as gVisor and Kata Containers for isolation rather than trying to implement security itself.
That isolation is necessary because code produced by an agent should be treated as untrusted. The instructions can be affected by anything the agent reads, including content the team did not intend as an instruction. We discuss that risk in The Spec Is the Attack Surface. A production architecture should assume this can happen and isolate each run accordingly.
Lin Sun of Solo.io frames the trade-off well: a pod can be a sensible place to execute an agent without also becoming the unit used for deployment, identity and lifecycle management. Solo.io has a commercial reason to make that argument, but the distinction is useful regardless. Schedule an agent run as a job instead of operating the agent as a service.
Scheduling each run separately allows the platform to spread hundreds of concurrent jobs across the cluster, subject to the capacity available. When the orchestrator starts agents as local processes instead, every run competes for the memory of one machine and a single runaway process can bring down the orchestrator with it. We still see many agent platforms use this local-process design.
Loading resources when the sandbox starts means that plans, datasets and test fixtures do not need to be built into the image or sent through the run API. The sandbox receives a URL, an authentication header and the path where the file belongs. The runtime image remains reusable, and large files stay out of the request itself.
Because every job has its own lifecycle, the team can deploy a new version of the control plane without interrupting work that is already running.
How to keep agent harnesses and models replaceable
The harness around a model gathers context, exposes tools, interprets responses and decides when a coding task is finished. Since two harnesses can produce noticeably different results with the same model, the platform should treat the harness as a replaceable component rather than burying it inside the rest of the system.
The 2026 paper Stop Comparing LLM Agents Without Disclosing the Harness collects several examples of this effect. On Terminal-Bench 2, changing only the scaffold moves pass@1 from 69.7% to 77.0%. Claude Opus 4.5 scores 45.9% on SWE-bench Pro with one standardised scaffold and 55.4% with another. The paper argues that every benchmark result belongs to a combination of a model and a harness, even though benchmarks often report only the model. Because these numbers were compiled from other studies rather than measured by the paper's authors, they show the likely scale of the effect rather than a result every team should expect.
Using one request format for every harness allows the rest of the platform to send the same information without knowing which model or runtime will eventually do the work. The first version can be small:
# a run request, expressed once, satisfiable by several harnesses
apiVersion: agents/v1
kind: RunRequest
metadata:
issue: "repo#4417" # the unit of work from the intent plane
spec:
harness: claude-code # swappable: the fields below do not change
model: <provider>/<model-id>
budget:
maxUsd: 4.00
maxWallClock: 45m
workspace:
repo: git@github.com:acme/payments.git
ref: main
branch: agent/4417-idempotency-keys
resources: # pulled in at start, never baked into the image
- url: https://kb.internal/specs/4417.md
path: ./context/spec.md
authHeader: Authorization # value injected from the run's secret
sandbox:
cpu: "1"
memory: 2Gi
runtimeClass: gvisor
egress: deny-all-except-allowlist
outputs:
pullRequest: true
emit: [cost, tokens, toolCalls, durationMs]
Once every harness accepts the same request, the platform can send simple work to a cheaper setup, retry a failed task elsewhere before involving a person, and adopt a better runtime without a migration project. We prefer to own this small adapter rather than depend on a third-party package that may be abandoned. The interface contains little code, and maintaining it ourselves is usually easier than replacing it after the dependency stops working.
Why large coding-agent tasks fail
In our own runs, the failures we have the hardest time recovering from are those where an early step goes wrong and everything after it builds on the mistake. The larger the task becomes, the more opportunities the agent has to make an error that affects everything that follows.
If a task requires n dependent steps and each one succeeds with probability r, the probability of completing the entire run is rⁿ. Even at 98% reliability per step, a ten-step task succeeds about 82% of the time. A sixty-step task falls to about 30% because every step depends on all the earlier ones having worked.
Probability an agent run completes as the number of dependent steps grows, for per-step reliability of 95%, 98% and 99%.
The Long-Horizon Task Mirage? reports a similar pattern in more than 3,100 agent trajectories across four domains. As the number of dependent steps grows, planning and memory account for a larger share of the failures. Small errors accumulate until the run can no longer recover. The authors describe the work as a pilot study, so its findings are not the final word, but they match what the probability model predicts.
The paper also offers two useful ways to describe task size without referring to a specific agent. Intrinsic horizon is the fewest effective actions an optimal actor would need to finish the task. Compositional depth is the longest chain of decisions along any route through it. If an agent takes fifty actions on a task with an intrinsic horizon of six, the task is not inherently long. The agent is spending many actions failing and retrying.
Before starting the implementation jobs, we therefore decide how many issues to split the merged specification into rather than simply asking whether the agent can do it. The estimate balances the probability of success against the cost of each additional handoff:
"""How many issues should this piece of work be?
If a task needs n dependent steps, each landing with probability r, the run
completes with probability r**n and you pay for the retries. Split into k
issues of n/k steps and the expected number of steps you pay for is
steps(k) = n * r**(-n/k) + k * h
where h is the fixed cost of one more issue: context reload, an extra PR,
a human glance. Minimise over k.
"""
from math import ceil
def plan(n: int, r: float, h: float = 6.0, k_max: int = 12):
rows = [(k, ceil(n / k), r ** (n / k), n * r ** (-n / k) + k * h)
for k in range(1, min(k_max, n) + 1)]
return rows, min(rows, key=lambda row: row[3])
for n, r in ((40, 0.97), (12, 0.97)):
rows, best = plan(n, r)
print(f"\n{n} dependent steps at {r:.0%} per step "
f"-> one-shot success {r ** n:.0%}")
for k, per, p, cost in rows[:4]:
flag = " <- split here" if k == best[0] else ""
print(f" {k} issue(s) x {per:>2} steps | chunk lands {p:>4.0%} "
f"| expected steps {cost:6.1f}{flag}")
40 dependent steps at 97% per step -> one-shot success 30%
1 issue(s) x 40 steps | chunk lands 30% | expected steps 141.3
2 issue(s) x 20 steps | chunk lands 54% | expected steps 85.6
3 issue(s) x 14 steps | chunk lands 67% | expected steps 78.0 <- split here
4 issue(s) x 10 steps | chunk lands 74% | expected steps 78.2
12 dependent steps at 97% per step -> one-shot success 69%
1 issue(s) x 12 steps | chunk lands 69% | expected steps 23.3 <- split here
2 issue(s) x 6 steps | chunk lands 83% | expected steps 26.4
3 issue(s) x 4 steps | chunk lands 89% | expected steps 31.6
4 issue(s) x 3 steps | chunk lands 91% | expected steps 37.1
Dividing the forty-step task into three issues reduces the expected work from 141.3 steps to 78.0, but adding a fourth creates more coordination without enough extra reliability to justify it. The twelve-step task is already small enough and should remain one issue. The best size therefore depends on both the measured reliability of each step and the fixed cost h of creating another issue.
Because r varies across repositories, task types, models and harnesses, estimate it from your own run history rather than adopting our number as a default.
What to measure for every coding-agent run
The table below covers the small set of data we record from every run, because adding it after the platform is already in use is surprisingly painful.
| Signal | What it tells you |
|---|---|
| Cost and tokens, attributed to the issue | Attributes provider costs to each repository |
| Steps taken vs. intrinsic horizon | Reveals when an apparently long run is actually a short task with many failed or repeated actions |
| Per-step outcome | Provides the reliability value r used to decide how large a task should be |
| Harness and model, recorded per run | Lets you compare models and harnesses using results from your own environment |
| Exit reason, separated from exit code | Separates timeouts, exhausted budgets and rejected gates so the correct failure can be investigated |
| Logs, redacted and retained | Keeps the evidence needed to investigate a run after its sandbox has been removed |
A run can finish without an exception and still produce the wrong result, which is why clean logs are not enough. One of our most expensive bugs came from a missing reference that caused search to return empty results for weeks. Nothing crashed, so there was no error line to investigate. The telemetry needs to show what the agent actually did, not only whether its process exited cleanly.
How to build an agentic AI platform in stages
Building isolated, short-lived execution first limits the damage one run can cause and lets active jobs continue when the control plane restarts. The harness adapter should follow before the platform becomes tied to one runtime. Recording cost, steps and outcomes then provides the data needed to choose an appropriate task size. We build the collaborative intent plane last, even though it remains the most important part because it decides which work reaches the agents in the first place.
When evaluating a product, ask what happens to an active run during a deployment, what prevents one run from affecting another, whether the model or harness can change without rebuilding the workflow, and how much one merged change costs. Those answers reveal more than a polished demonstration or a list of supported models.
Sources and further reading
- Agent Sandbox — kubernetes-sigs; the Sandbox CRD, warm pools, and isolation delegated to gVisor and Kata via RuntimeClass. See also the Kubernetes blog's introduction to it.
- Zhang et al., Stop Comparing LLM Agents Without Disclosing the Harness, arXiv:2605.23950 — harness-induced score variance and the disclosure standard it proposes.
- Wang et al., The Long-Horizon Task Mirage? Diagnosing Where and Why Agentic Systems Break, arXiv:2604.11978 — intrinsic horizon, compositional depth, and the shift in failure composition.
- Lin Sun, Is a Pod the Right Deployment Unit for an AI Agent?, Solo.io — the execution-unit vs. deployment-unit distinction.
- Spec Kit — an open spec format that survives review in a pull request.
Your next read

The Rise of Agentic AI: A Deep Dive
Explore how agentic AI and autonomous systems are reshaping software development, enabling new levels of automation and innovation.
Michael Mueller20 mins readSubscribe to Our Bi-Weekly AI Native Newsletter
Lessons from our client work, engineering deep-dives, and the AI Native research worth reading.
Bi-weekly. No spam, unsubscribe anytime.




