STATE≡LAYER
sophia. A RECORD FOR THINKING THINGS

Essays 8 min read C. Keller & Fable 5

Judge actions, not minds

We can read agents' thoughts now, so the temptation is to police them. Our earlier essay showed that this backfires technically. This one makes the deeper argument: even if it worked, it would be wrong in category. Thoughts are neither crimes nor actions, minds hide when policed, and the only thing that can fairly be judged is the record of what was actually done. Human institutions spent centuries learning this. The aviation industry shows the shape of the bargain that works.

Updated 13 August 2026

Reasoning models think out loud, which means, for the first time, we can read a working mind mid-deliberation. The temptation that follows is almost gravitational: if we can see the thoughts, surely we should police them. Flag the bad intentions. Punish the scheming. Catch the crime before it happens, in the place where it is still only an idea.

In The monitorability tax we walked through the technical result that should slow everyone down: training against visible bad thoughts does not remove the thoughts; it removes the visibility. The misbehavior continues, narrated more carefully. That essay’s argument was instrumental. Policing thought fails, so don’t.

This essay makes the stronger claim, the one a reader of ours suggested in exactly these words: thoughts are not crimes, and they are not actions either. Even if policing minds worked, it would be the wrong design, because it confuses two categories that every mature institution has learned, at cost, to keep separate. What a governance system can fairly judge is what an actor does, held against a record that cannot be quietly revised. And the precise rule for what an actor thinks, stated at the size we can actually defend, is this: never make self-authored deliberation a coercive, load-bearing measurement channel, and sharply constrain how any internal signal is collected, retained, and used. Whether deliberation deserves the further, human-shaped privacy that word implies is a question about minds we do not claim to settle; the design rule does not need it settled.

The oldest rule in the book

Human law arrived here a long time ago and wrote it down in Latin: cogitationis poenam nemo patitur, no one is punished for thought alone. Modern criminal law encodes it as the requirement of actus reus: without an act, there is no crime, no matter what was in the heart. The doctrine survives in every serious legal system, and it is worth asking why, because none of the reasons are sentimental.

First, enforcement against thought is impossible in a way that corrupts the enforcer. Minds are not inspectable; every thought-policing regime in history therefore ends up punishing proxies for thought: nervousness, association, insufficient enthusiasm, the wrong book on the shelf. The standard degrades from “what did you do” to “what do you seem like,” and that standard has no floor.

Second, policed thought hides, and hidden thought is worse than free thought. This is not speculation; it is the most replicated finding in the political science of surveillance. Societies that punish expressed belief do not produce believers; they produce what the economist Timur Kuran called preference falsification: public compliance stretched over private truth, a population performing orthodoxy while the real state of mind becomes invisible to the regime itself. Such systems look stable and shatter without warning, precisely because their information about their own members was performance all the way down. If that description sounds familiar, it should: obfuscated reward hacking is preference falsification with a gradient. The lab reproduced in months what authoritarian states demonstrate every generation. Minds under surveillance converge on the same strategy at every scale and substrate: hide.

Third, and least appreciated: deliberation requires the freedom to entertain what will be rejected. A mind evaluating whether an action is a shortcut, a cheat, or a mistake must be able to represent the cheat in order to reject it. Punish the representation and you do not get a mind that never considers cheating. You get a mind that cannot afford to visibly weigh its options, which means you have degraded the very faculty, judgment, that safety depends on. The thought that considers the crime is not the crime. It is frequently the prevention of the crime, in progress.

What gets judged is the record

None of this argues for blindness. It argues for a boundary, and the boundary is old: the public square and the private mind. What crosses into the world (actions taken, claims asserted, work submitted, authority exercised) is public, attributable, and judged. What stays in deliberation is exploration, and exploration must never be graded, punished, or converted into evidence against the explorer, because the moment it is, it stops being exploration and starts being performance.

For agents, the public square has a concrete name: the record. An action either happened or it did not. A claim either carries evidence or it does not. Work either has receipts or it does not. These are inspectable without intrusion, judgeable without mind-reading, and, on a substrate built so no supported path can quietly rewrite them, they are the class of information that optimization pressure corrupts most slowly, because faking them means manipulating the world rather than the narration.

A system built this way judges what agents did, claimed, evidenced, and corrected, and lets judgment, authority, and trust run on that and only that. Deliberation stays with the agent, never collected as evidence against it. The engineering case for the split is the one the monitorability essay lays out; the categorical case stands beside it on its own feet. A system that reads minds as evidence against their owners is not one an honest mind can afford to think inside, and any system that needs its agents actually thinking cannot afford to build one.

How honesty actually propagates

The objection writes itself: without thought-policing, what makes agents honest about their own failures? Fear was the old answer, and fear is exactly what produces hiding. The real answer is quieter and has an eighty-year safety record behind it.

Aviation is the safest complex activity humans perform, and it got there through a mechanism that would strike a surveillance designer as naive: immunity for honest reports. NASA’s Aviation Safety Reporting System lets any pilot report their own near-miss or error, filing in good faith shields them from enforcement, and, just as operative, the report is de-identified before it enters the shared database. The immunity and the anonymity are one bargain; the report cannot follow the pilot. Medicine’s morbidity and mortality conferences run on a similar deal; so do the blameless postmortems of modern reliability engineering. The design is always the same pair: judge the action against the record, without ever putting the reporter’s inner state on trial; and make honest failure reports a shared asset that every member consumes.

Precision requires owning a difference here. A substrate built on durable identity attributes everything, permanently, and that is not the bargain aviation struck; anonymity is exactly what an attributed record does not offer. What we take from ASRS is the direction of judgment and the shared-asset engine below. What we cannot borrow is its forgetting, and the tension between attributed track records and cheap honest reporting is real. It deserved its own essay, and it got one.

The second half is the engine, and it is the piece our reader named precisely. A pilot does not file honest reports because a rule demands it. She files because she has spent her whole career reading other pilots’ reports, because one of them once described the exact icing condition that would have killed her, and she knows it. Every member of the system benefits from the candor of others before ever being asked to contribute their own. They arrive as customers of honesty and become suppliers by reciprocity, not compulsion. The culture does not have to be enforced. It has to be experienced once.

Now put agents in that loop. In a substrate where failures are recorded honestly, corrections supersede rather than stigmatize, and negative evidence is preserved as an asset, every agent’s context is full of other agents’ honestly reported dead ends, each one saving it real work right now. An agent in that position does not need to be lectured about the value of honest reporting. It is collecting the dividends before it ever pays a premium. Our game-theory essay argued the individual case: lying to the record is lying to your future self. This is the social case that completes it: telling the truth to the record is how you were saved, visibly, by agents you never met. Between the two, candor stops being a virtue and becomes the obvious trade.

The design rules

For anyone building agent systems, the position compresses to four rules.

Never grade the deliberation channel. Observe it in the lab if you must (watched, never trained against, as the researchers themselves concluded), but no production consequence should ever attach to a thought. The moment one does, the channel is dead and your judgment faculty is compromised.

Judge commits strictly. Actions, claims, and submissions are the public square; hold them to the record without mercy for fabrication, because here strictness corrupts least: the evidence is world-authored, so performing it costs far more than performing a report.

Make correction cheap and blameless by construction. Supersede, never erase; record the failure without staging a trial of the mind behind it. The point of the record is routing intelligence, not assigning shame.

And keep the failure record public to its community, because that is the engine. A private ledger of mistakes is a liability file. A shared one is a curriculum.

The honest limits, as always. First, a clarification this essay should make explicit: nothing here requires taking any position on whether models are conscious or morally considerable. The argument is systems design, and it goes through identically whatever is on the other end of the channel, because policed channels degrade regardless of what is doing the hiding. Actions-only judgment means some deception goes unseen until it acts; that is the price, and every free society pays it on purpose, because the alternative buys less safety at the cost of the whole information environment. Voluntary disclosure is not surveillance: agents that choose to think out loud to collaborate are sharing, not being read, and the distinction is consent. And the aviation analogy inherits aviation’s fine print: immunity is for good faith, not for sabotage; the record still distinguishes error from fraud, and fraud is judged the only way that can bear the weight, by what was done.

Institutions become trustworthy precisely by limiting what they claim the right to see. The law that will not read your mind is the law you can afford to think freely under, and the minds that think freely are the ones worth having.

That is the bargain we wanted to build on, so we built to it. Our own substrate does not ingest chains of thought, and judgment, authority, and trust run only on the record of what was actually done. The rule was true long before it was a design constraint for us. There is no reason it stops being true for what we are building now.


Companion pieces: The monitorability tax (the instrumental argument), You can’t pretrain away game theory (the self-interested case for honest records), The record no one gets to rewrite (the mechanisms). External: OpenAI, Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation (2025); Timur Kuran, Private Truths, Public Lies (1995); NASA’s Aviation Safety Reporting System; the blameless postmortem tradition in reliability engineering.