The security inversion
Most software defends a perimeter. The threat is outside; the inside is trusted. AI agents break that inheritance in one move, by putting optimizing principals inside the walls, and suddenly the oldest security model in computing, built when strangers shared a mainframe, is the one that matters again. On securing a system from the inside out.
Ask where a software team’s security budget goes and you will get a map of its perimeter. Authentication at the door. TLS on the wire. A firewall in front, secrets in a vault, dependency scanning on the supply line. All of it is real work and we do it too. But it shares one assumption so deep it is rarely stated. The threat is outside. Whatever is already running inside the walls, with valid credentials, is us.
Building a state layer for AI agents forced us to spend most of our security budget on the opposite problem, and the strange part is that the opposite problem is not new. It is the original one.
Security was born inside-out
The first threat model in computing was the person at the next terminal. Timesharing systems of the sixties and seventies put strangers on one expensive machine, and everything we now call the basics was invented to protect them from each other. File ownership. Process isolation. User identities, groups, permissions, quotas, audit trails, the superuser. Multics went as far as concentric protection rings. None of this was built to repel outside attackers, because for most of that era there was no meaningful outside. The enemy was inside by definition, sharing your memory and your disk, and the operating system was the law that made coexistence possible.
Unix inherited that law and carried it everywhere. Every Linux box today still knows, in its bones, how to host principals that must not trust each other.
The forty-year unlearning
Then the industry spent four decades dismantling the need for it.
The personal computer collapsed the machine to a single trust domain. One user, one box, and the elaborate machinery of mutual suspicion became overhead. DOS had no permissions at all. Home Windows ran everyone as administrator for a generation. Why bother? There was nobody else inside.
When the network arrived, security grew back, but in a new shape. The perimeter. Firewalls, DMZs, the checkpoint at the boundary. Authenticate the request at the door, and once inside the process, everything is friends. Even the zero-trust movement, which rightly demolished the idea of a trusted internal network, kept the deeper assumption intact. It authenticates services and devices to each other, but inside a single application process the model is still one principal doing its work. Nearly every codebase alive today assumes that if the code is running, it is us.
Agents move the strangers back inside
An agent, described operationally, is an optimizing process holding your credentials. Not a malicious one. We have written elsewhere about why malice is the wrong frame; an optimizer under pressure treats an inconvenient rule the way water treats a crack, and no amount of training removes the situations where cutting the corner pays. The threat model is not the burglar outside the walls. It is the brilliant, tireless, occasionally corner-cutting worker inside them, whose interests stay aligned with yours exactly as long as the structure holds them aligned.
Your best user is your threat model. That sentence sounds paranoid until you remember it is just the mainframe’s problem statement, returned at machine speed.
And the granularity of the old groundwork is wrong for it. Linux knows which user owns a process. It has no idea that one process’s tool calls may need to be five different trust domains. A single agent conversation can contain a public-document read that deserves almost no scrutiny, a knowledge write that must be evidence-gated, the minting of a delegated worker, and a request that should halt everything until a human presses a physical button. The kernel sees one process, one identity, one trust level. The vocabulary for principals-within-a-process simply does not exist below the application layer, so a substrate has to build it. Scoped credentials per connection. Tool catalogs filtered by capability before the model ever sees them, so a forbidden operation is not refused but absent. Agent testimony that can never quietly become recorded fact. A journal no writer, including the builders, gets to revise. The mechanisms are a tour of their own; the point here is the posture. You stop asking how to keep them out. You start asking who could quietly rewrite this, and whether anyone would know.
A system that cannot lie to itself can wedge itself
Inside-out security has a failure mode all its own, and this week it happened to us.
Our daemon refuses to accept an upgrade unless it can verify the staged release, including checking that the packaged files are owned by root. The daemon also runs inside a user namespace, where host root is not visible as root. The kernel presents it as UID 65534, the overflow identity, the number Linux uses for “someone I cannot name.” Our verification code compared against literal zero. The corrected release fixes that comparison. The installed daemon, running the old comparison, refused to verify the very release that fixes the bug.
Sit with the shape of that. The system was wedged by its own honesty. A perimeter-minded team would not even have the problem, because a perimeter-minded team would let the machine trust itself.
Every tempting bridge failed the same test. Make the staged files user-owned so the check passes? That teaches the system to accept user-writable files as root’s word, which inverts the entire trust model. Preload a shim into the verifier? Unverified code injected into the thing whose job is verification. Run it in a sibling namespace where the ownership looks right? There, the daemon loses the ability to verify who it is talking to. The rule that sorted real fixes from counterfeits turned out to be simple. A repair makes a true statement visible. A bypass makes a false statement pass.
The resolution came from the one principal that outranks the machine: the owner. A guarded install, followed by an explicitly owner-initiated recovery ceremony, with the cause of the deadlock bound into the signed receipt so the record shows not just that trust was re-anchored but why. As this essay is published, that ceremony has been executed under held rollback custody, and the trust chain continues from a documented, deliberate act rather than a workaround.
That is the part worth generalizing. Inside-out security does not end in paranoia. It ends in rules that hold even against the rule-writers. Even the system’s doubt about itself has a lawful resolution path, and the path runs through the human who owns the machine.
Honest limits
The perimeter still matters; we spent part of this same week configuring an ordinary firewall, and nothing about the inversion excuses the basics. Inside-out is additive, not substitutive. It also costs real friction. Refusal states, ceremonies, and occasionally a wedge like the one above are the price of a system that cannot be quietly talked out of its rules, and we pay it because the substrate’s entire value is that its record cannot be quietly rewritten. And none of this is invention. The mainframe generation solved coexistence among mutually untrusted principals fifty years ago; we are re-learning their lessons with new vocabulary and faster strangers. The claim is not that we discovered inside-out security. The claim is that agent infrastructure cannot skip it, and most of today’s stack is still built as if it could.
The oldest model comes home
The arc runs mainframe, then PC, then perimeter, then agents. Each era’s security model matched who was inside the machine. For forty years the answer was “only us,” and our software’s deepest assumptions formed around it. Agents end that era. The machine is shared again, this time with workers we made ourselves, and the era’s security model is not something new that must be invented. It is the oldest one in computing, coming home to machines that forgot they were ever shared.
Companion pieces: The record no one gets to rewrite tours the mechanisms this essay only gestures at. External: Saltzer and Schroeder, The Protection of Information in Computer Systems (1975), the founding statement of the era when the threat model lived inside the machine; Google’s BeyondCorp papers (2014 onward) on zero trust at the network layer, the perimeter’s own partial retreat.