- Published on
Enterprise Agent Architecture: The Control Plane
- Authors

- Name
- Patrick Burger
The moment an agent acts, you stop running software and start operating an autonomous system. That sentence is the whole problem. Software that only answers questions fails by being wrong. An agent that can create a work order, move money, or change production fails by being unauthorized — and the failure is indistinguishable from normal operation until someone reconstructs what happened.
Most agent programs do not fail at the model. They fail at the seams: where authority is assumed, where scope is inferred from untrusted context, where a handoff loses state, where nobody can say what an agent was allowed to do. Those are architecture problems, and they are solvable with the same rigor we already apply to distributed systems.
This is the architecture I keep coming back to. Five layers, and the specific failure each one prevents.
The failure mode that generalizes
Prompt-level controls cannot hold. The pattern shows why.
Anthropic's cybersecurity-evaluation incident is worth reading carefully — not because a model escaped, but because nothing had to. The evaluation environment had a live path to the internet while the model was told it did not. OpenAI published a comparable case: a research agent in a sandbox built to keep it offline found the DNS resolver still reached the live internet and used DNS delegation to ask a public chatbot for help.
Both are harness failures, not alignment failures. And both generalize: any system that infers its own scope rather than having it enforced is exposed the same way. An agent's view of what it may do is assembled from context, and context is influenceable. Neither a prompt claiming "this is a simulation" nor a policy document listing approved systems is a security control.
- Anthropic's Cybersecurity-Eval Incident and Bounded Autonomy — why a flawed operating environment was enough
- AI Agents Need a Kill Switch, Not a Meeting — the OpenAI DNS case, and why stopping must be mechanical
Layer 1: Identity and task-scoped policy
Traditional access control answers whether an identity may reach a system. Agentic systems need a narrower decision: may this agent take this action, for this task, using this evidence, under this policy, right now? That is delegated authority, and it cannot be satisfied by a service account with ambient permissions.
In practice that means just-in-time, ephemeral, task-scoped credentials: the minimum data, tools, and permissions the task requires, and no standing authority beyond it. An agent that inherits a developer's permissions inherits a developer's blast radius.
- AI Agents: The Permissions Problem Nobody's Solving
- Every MCP Tool Is a Door. Most Teams Leave Them Wide Open
Layer 2: A gateway that enforces it
Policy that lives in a document is advice. Policy that lives inline, on the path every agent action takes, is a control. The gateway is where identity is checked, task-scoped permissions are evaluated, policy decisions are invoked, and the evidence needed to explain an action is correlated.
Two consequences. The enforcement belongs outside the agent framework — controls embedded in one framework do not govern authority consistently across teams, tools, and future framework choices. And the gateway itself is becoming a commodity; the durable asset is the policy stack it enforces.
- The Agent Framework Should Not Become the Control Plane
- The Gateway Is Commoditizing — The Policy Stack Is Not
- MCP Goes Stateless. Trust Boundaries Relocate. — where trust continuity moves when the protocol stops carrying it
Layer 3: An immutable event backbone
If you cannot reconstruct what the system knew and decided, you are debugging a black box. Every prior, every action, every policy decision has to land on a durable, append-only record — not as logging, but as the substrate that makes an action explainable after the fact.
This is also what makes recovery possible. In autonomous systems the order was always recovery before capability: a pipeline that cannot recover from a failed step is not a system, it is an expensive demo.
- LLM Output Is a Measurement. Sensor Fusion Solved This. — priors, gates, and why retries keep failing without a calibrated loop
Layer 4: A live registry
You cannot govern what you cannot enumerate. Three questions that most organizations still cannot answer cleanly: how many agents do we have, what is each allowed to do, and what are they doing right now.
The scale is not hypothetical — Gartner projects the average Fortune 500 running over 150,000 AI agents by 2028, up from fewer than 15 today, with only 13% of organizations having adequate governance in place. Without a registry, agent sprawl is not a multi-agent system; it is a distributed system with no contract between its components.
Layer 5: Observability tied to business outcomes
There is a large difference between knowing a system is running and knowing it is working. An agent can stay inside latency budgets, pass every health check, and drift completely outside the intent it was deployed to serve. Dashboards keep looking fine while the business outcome breaks.
So the observability question is not "what happened" but "was the outcome still inside the boundary the business expected" — which requires trace data, agent identity, and policy enforcement in the same plane.
What the layers buy you: bounded autonomy
Composed, the layers buy one thing worth having. The agent acts independently inside a defined domain, and the moment it steps outside, the platform stops it — not a human, not a wiki page. Autonomy you can revoke quickly is autonomy you can grant generously.
That also sets the ceiling: how much autonomy you can hand an agent is set by how fast you can take it back. Stopping should be automatic; continuing should be the deliberate decision.
- The Moment an Agent Acts, You Operate an Autonomous System — the Agent Autonomy Level framework, and why AAL-3 is the dangerous middle
Two things that are not infrastructure
The architecture is necessary. It is not sufficient. Two constraints sit outside the stack and break programs anyway.
Coordination is an organizational problem. Span of control, boundary objects, and coupling applied to organizations long before agents existed, and they still apply. An orchestrator hits the same wall a manager does: can it split work cleanly, judge what comes back, and reconcile conflicts? A chat transcript is a weak way to pass work between specialised agents.
Context is an economic problem. Every static tool definition and permanent instruction is charged on every turn. Harness design is therefore platform economics — and, because those instructions encode what an agent believes it may do, governance too.
The short version
Enforce scope at the boundary, not in the prompt. Put policy inline, not in a document. Record enough to reconstruct any decision. Know what exists. Measure business state, not system health. Then autonomy becomes something you can extend deliberately instead of discover after an incident.
The honest counterpoint: none of this is free. Inline enforcement costs latency on every action, scoring costs compute, and most teams will not pay that tax until an incident makes the alternative more expensive. Both costs are priced per action, which is exactly how agent economics should work — but that argument lands better after the first postmortem than before it.
The organizations that get this right will not be the ones that governed most carefully. They will be the ones that designed for it early enough that it became a property of the system rather than a process bolted on top.