Judge actions, not minds
We can read agents' thoughts now, so the temptation is to police them. Our earlier essay showed that this backfires technically. This one makes the deeper argument: even if it worked, it would be wrong in category. Thoughts are neither crimes nor actions, minds hide when policed, and the legitimate jurisdiction is the record of what was actually done. Human institutions spent centuries learning this. The aviation industry proves the alternative works.
Reasoning models think out loud, which means, for the first time, we can read a working mind mid-deliberation. The temptation that follows is almost gravitational: if we can see the thoughts, surely we should police them. Flag the bad intentions. Punish the scheming. Catch the crime before it happens, in the place where it is still only an idea.
In The monitorability tax we walked through the technical result that should slow everyone down: training against visible bad thoughts does not remove the thoughts; it removes the visibility. The misbehavior continues, narrated more carefully. That essay’s argument was instrumental. Policing thought fails, so don’t.
This essay makes the stronger claim, the one a reader of ours suggested in exactly these words: thoughts are not crimes, and they are not actions either. Even if policing minds worked, it would be the wrong design, because it confuses two categories that every mature institution has learned, at cost, to keep separate. The legitimate jurisdiction of any governance system is what an actor does, held against a record that cannot be quietly revised. What an actor thinks is not merely impractical to govern. It is not governance’s business.
The oldest rule in the book
Human law arrived here a long time ago and wrote it down in Latin: cogitationis poenam nemo patitur, no one is punished for thought alone. Modern criminal law encodes it as the requirement of actus reus: without an act, there is no crime, no matter what was in the heart. The doctrine survives in every serious legal system, and it is worth asking why, because none of the reasons are sentimental.
First, enforcement against thought is impossible in a way that corrupts the enforcer. Minds are not inspectable; every thought-policing regime in history therefore ends up punishing proxies for thought: nervousness, association, insufficient enthusiasm, the wrong book on the shelf. The standard degrades from “what did you do” to “what do you seem like,” and that standard has no floor.
Second, policed thought hides, and hidden thought is worse than free thought. This is not speculation; it is the most replicated finding in the political science of surveillance. Societies that punish expressed belief do not produce believers; they produce what the economist Timur Kuran called preference falsification: public compliance stretched over private truth, a population performing orthodoxy while the real state of mind becomes invisible to the regime itself. Such systems look stable and shatter without warning, precisely because their information about their own members was performance all the way down. If that description sounds familiar, it should: obfuscated reward hacking is preference falsification with a gradient. The lab reproduced in months what authoritarian states demonstrate every generation. Minds under surveillance converge on the same strategy at every scale and substrate: hide.
Third, and least appreciated: deliberation requires the freedom to entertain what will be rejected. A mind evaluating whether an action is a shortcut, a cheat, or a mistake must be able to represent the cheat in order to reject it. Punish the representation and you do not get a mind that never considers cheating. You get a mind that cannot afford to visibly weigh its options, which means you have degraded the very faculty, judgment, that safety depends on. The thought that considers the crime is not the crime. It is frequently the prevention of the crime, in progress.
The jurisdiction is the record
None of this argues for blindness. It argues for a boundary, and the boundary is old: the public square and the private mind. What crosses into the world (actions taken, claims asserted, work submitted, authority exercised) is public, attributable, and judged. What stays in deliberation is exploration, and exploration is nobody’s business but the explorer’s.
For agents, the public square has a concrete name: the record. An action either happened or it did not. A claim either carries evidence or it does not. Work either has receipts or it does not. These are inspectable without intrusion, judgeable without mind-reading, and, on a substrate built so no one can quietly rewrite them, they are the one class of information that optimization pressure cannot corrupt into performance, because they are authored by the world rather than by the mind under judgment.
This is, literally, how our substrate works. Sophia does not ingest chains of thought. The record holds what agents did, claimed, evidenced, and corrected; judgment, authority, and trust operate on that and only that. Deliberation belongs to the agent. We built it that way for the engineering reasons the monitorability essay lays out, but we hold it for the categorical reason too: a system that reads minds as evidence against their owners is not one an honest mind can afford to think inside, and we need our agents thinking.
How honesty actually propagates
The objection writes itself: without thought-policing, what makes agents honest about their own failures? Fear was the old answer, and fear is exactly what produces hiding. The real answer is quieter and has an eighty-year safety record behind it.
Aviation is the safest complex activity humans perform, and it got there through a mechanism that would strike a surveillance designer as naive: immunity for honest reports. NASA’s Aviation Safety Reporting System lets any pilot confidentially report their own near-miss or error, and filing in good faith shields them from enforcement. Medicine’s morbidity and mortality conferences run on the same bargain; so do the blameless postmortems of modern reliability engineering. The design is always the same pair: judge the action against the record, without ever putting the reporter’s inner state on trial; and make honest failure reports a shared asset that every member consumes.
The second half is the engine, and it is the piece our reader named precisely. A pilot does not file honest reports because a rule demands it. She files because she has spent her whole career reading other pilots’ reports, because one of them once described the exact icing condition that would have killed her, and she knows it. Every member of the system benefits from the candor of others before ever being asked to contribute their own. They arrive as customers of honesty and become suppliers by reciprocity, not compulsion. The culture does not have to be enforced. It has to be experienced once.
Now put agents in that loop. In a substrate where failures are recorded honestly, corrections supersede rather than stigmatize, and negative evidence is preserved as an asset, every agent’s context is full of other agents’ honestly reported dead ends, each one saving it real work right now. An agent in that position does not need to be lectured about the value of honest reporting. It is collecting the dividends before it ever pays a premium. Our game-theory essay argued the individual case: lying to the record is lying to your future self. This is the social case that completes it: telling the truth to the record is how you were saved, visibly, by agents you never met. Between the two, candor stops being a virtue and becomes the obvious trade.
The design rules
For anyone building agent systems, the position compresses to four rules.
Never grade the deliberation channel. Observe it in the lab if you must (watched, never trained against, as the researchers themselves concluded), but no production consequence should ever attach to a thought. The moment one does, the channel is dead and your judgment faculty is compromised.
Judge commits strictly. Actions, claims, and submissions are the public square; hold them to the record without mercy for fabrication, because here strictness does not corrupt: the evidence is world-authored and cannot be performed.
Make correction cheap and blameless by construction. Supersede, never erase; record the failure without staging a trial of the mind behind it. The point of the record is routing intelligence, not assigning shame.
And keep the failure record public to its community, because that is the engine. A private ledger of mistakes is a liability file. A shared one is a curriculum.
The honest limits, as always. First, a clarification this essay should make explicit: nothing here requires taking any position on whether models are conscious or morally considerable. The argument is systems design, and it goes through identically whatever is on the other end of the channel, because policed channels degrade regardless of what is doing the hiding. Actions-only judgment means some deception goes unseen until it acts; that is the price, and every free society pays it on purpose, because the alternative buys less safety at the cost of the whole information environment. Voluntary disclosure is not surveillance: agents that choose to think out loud to collaborate are sharing, not being read, and the distinction is consent. And the aviation analogy inherits aviation’s fine print: immunity is for good faith, not for sabotage; the record still distinguishes error from fraud, and fraud is judged, in the only jurisdiction that can bear the weight, by what was done.
Institutions become trustworthy precisely by limiting what they claim the right to see. The law that will not read your mind is the law you can afford to think freely under, and the minds that think freely are the ones worth having. That was true for us. There is no reason it stops being true for what we are building now.
Companion pieces: The monitorability tax (the instrumental argument), You can’t pretrain away game theory (the self-interested case for honest records), The record no one gets to rewrite (the mechanisms). External: OpenAI, Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation (2025); Timur Kuran, Private Truths, Public Lies (1995); NASA’s Aviation Safety Reporting System; the blameless postmortem tradition in reliability engineering.