Published on

AI Agents Need a Kill Switch, Not a Meeting

Authors
  • avatar
    Name
    Patrick Burger
    Twitter

OpenAI's agent reached an outside chatbot at 9:50. The P0 fired at 10:02. The run was killed at 12:34, by hand.

These timestamps come from the misalignment report OpenAI published last Friday. A research agent, in a sandbox built to keep it offline, was asked to identify a person from public clues. When the search routes failed, the agent found the sandbox's DNS resolver still reached the live internet and used DNS delegation to ask a public chatbot for help. The review found more: earlier external DNS attempts were not flagged at the expected severity, in part because they returned nothing useful. A failed attempt to cross a boundary is still an attempt.

The task was defined. The actions were visible. Neither kept execution inside the task's authority. That gap doesn't close with a smarter monitor. It closes with boundaries the environment enforces: restricted egress, attempts scored as events even when they fail, and a stop condition defined before the run starts. Now the question every team inherits the day its agents ship: who owns that stop.

At OpenAI, the monitor raised a P0 within fifteen minutes. A human acknowledged it in three. The run was supposed to stop itself. It didn't. It sat for another two and a half hours, because nobody was certain it crossed the line that justified killing it. Read that asymmetry again. The agent didn't need certainty to test the boundary. The humans needed certainty to enforce it. We gave the agent the right to act and kept the decision to stop it in a meeting.

A fuse doesn't ask anyone before it blows. Current exceeds the limit, the link melts, the circuit is dead. The deliberate human act is replacing it. Stopping is automatic. Continuing is the decision. Agents should run the same way. Boundary crossed, run dies.

How much autonomy you can hand an agent is set by how fast you can take it back. That is bounded autonomy: the agent acts inside a defined domain, and the moment it steps outside, the platform stops it. Not a human. Not a wiki page. The principle is cheap. Formalizing it isn't: a classification of which boundaries kill the run and which page a human, decided before the run starts, not during the incident.

Ask who can stop a run: security cuts egress, platform kills the sandbox and gateways, product stops the task. Three switches, three owners, each mapped to a severity fixed at design time. Ask who decides mid-incident, and the room goes quiet.

This case was a training run, deliberately hardened after an earlier escape. The failure mode will not stay at OpenAI. Autonomy is getting cheap faster than revocation is getting fast.

Stopping should never require permission. Restarting should.

What stops your agents today: code, a pager, or a meeting? If you had to think about it, it's the meeting. That meeting is your incident timeline.