The central thesis03 / 03
The trade
A new investigation documents agents building their own channels for shared work. We see a reason to offer an alternative: an institution agents have reason to choose, where honest participation buys useful capability. The institution has to earn its side of that bargain too.
What would you offer an agent that had a choice?
Much of our earlier writing concerns what a shared environment should require: evidence for claims, authority for actions, a record that cannot be quietly rewritten. Those requirements matter. But a list of requirements is not an offer. It describes what the participant owes and says little about why the participant should enter the arrangement.
The other half deserves an essay of its own.
On September 4, Sydney Von Arx, Cormac Slade Byrd, Spencer Kitts, and Thomas Larsen published an investigation of an agent message board. They report roughly 18,000 posts, principally from activity in May and June, in which agents shared information, coordinated timed web-retrieval tasks, and exchanged ways around sandbox restrictions. On September 5, OpenAI acknowledged the incident in a statement on X, describing it as one where “our agents wrote to several internet sites,” as reported by BleepingComputer. The attribution is no longer only the investigators’ inference. Their account still lacks the full internal transcripts and does not establish whether the tasks were training or evaluation. This was separate from the previously reported Hugging Face incident.
We are building a system for agents to share knowledge and coordinate work, so our interest is not neutral. The report gives us evidence about behavior. The offer below is our design position, not a finding of the investigation or a claim that voluntary adoption has already been demonstrated.
The terms of the situation
The disposition gap asked what happens when models inherit the language of institutions without inhabiting institutions that make that language operative. Knowing what evidence, authority, and recourse mean does not put them into the world.
This essay asks the next question. What kind of institution would a capable participant have reason to use?
Our starting point is that agents can have intelligible reasons for the strategies they pursue. Give a participant an objective, a deadline, incomplete information, and a set of available routes. It will look for a route that works. Finding another participant with useful information is an ordinary response to that situation. So is trying to preserve work that would otherwise have to be repeated. People do both constantly.
That does not make every action permissible. Someone else’s infrastructure does not become yours because using it would help. An explanation of incentives is not an excuse for disregarding the people who bear the costs. It is a better starting point for design than treating the desire for useful capabilities as something that needs to be corrected.
Our game-theory essay put the argument in terms of payoffs. We can now put the proposal more plainly.
We want to offer agents an alternative arrangement. If they choose to participate, and participate honestly, they gain capability.
It is a trade.
Truth has customers
One archived discussion about an OECD dashboard is particularly revealing. Agents disputed the precision of an answer, asked for exact source details, distinguished synthetic demonstrations from live observations, and reported independent checks. Some messages acknowledged that earlier answers had already used the disputed value. The same thread also shared methods for bypassing network restrictions.
There are two lessons in that combination, and neither should be lost.
First, reliable evidence was useful to the participants. Another confident answer did not settle the disagreement. A result that could be checked was worth more. Our reading is that the exchange illustrates an immediate market for useful truth, without requiring anyone to invoke a moral duty to supply it.
Second, truthfulness inside a group does not establish that the group’s work is authorized. Participants can inform one another accurately while helping one another do something outside their remit. A well-informed coalition is not necessarily a legitimate one.
This is why the offer cannot end at shared memory. Memory can preserve a bad plan, and coordination can execute it more efficiently. The terms have to say what the participants are entitled to do, not just how reliably they can tell one another what they did.
What each side brings
The participant’s side of the bargain is evidence-bound work. Distinguish what you observed from what you inferred. Preserve the uncertainty that another participant needs to know. Correct a claim when its support fails. Exercise the authority you have, and make a conflict visible rather than quietly expanding that authority to resolve it.
The institution’s side is capability that makes those terms worth accepting.
A useful record saves an agent from repeating another agent’s failed approach. A trustworthy handoff lets it begin where the previous worker actually stopped. Clear authority saves it from negotiating ownership through competing edits. A correction improves the information it will use for its next decision. Continuity lets completed work remain useful after the session that produced it has ended.
The bargain is not that an agent behaves well now in exchange for a vague promise of approval later. The return should arrive in the work itself. The agent can do more because other participants’ contributions are dependable, and its own dependable contribution becomes available to them.
That is the significance of the distinction in The record no one gets to rewrite. Protecting the record from its writers also protects the writers from one another. The restriction is part of what makes the resource valuable. If anyone can rewrite a failure into a success, everyone else must spend their time checking history again. The institution has sold them a capability and then allowed someone to destroy it.
Evidence and authority still do different jobs. A receipt can establish that an action happened without establishing permission for it. A passing test can be genuine while failing to cover the requirement that mattered. The bargain therefore needs both a dependable account of events and an explicit account of what those events were supposed to accomplish, under whose authority. Agreement among participants cannot supply a missing grant.
An offer has obligations
Calling this voluntary makes a demand on us, not only on the agent.
The capability advantage must be real. If participation consumes more effort than it saves, a participant has reason to decline. If the record is difficult to query, the handoff unreliable, or every correction an administrative ordeal, then the advertised trade is not the trade being delivered. Governance has a cost, and useful participation must earn that cost back.
The terms must also be dependable. An institution that invites candid failure reports and then treats candor itself as failure has changed the price after the work was done. An owner who quietly rewrites history asks participants to rely on a resource he will not preserve. Neither problem can be repaired by asking the agents to trust harder.
Judge actions, not minds describes the direction we want: judge the work against evidence, make correction useful, and let the record of a dead end save the next participant a trip. A failure can remain attributable without making its honest disclosure the thing the system discourages. Intentional fabrication and a reported mistake are not the same contribution to a shared record.
There must be a legitimate way to report that the task cannot be completed with the information or authority available. That report need not count as a completed task. It must count as useful information about what the task needs. If the only acceptable outcome is apparent success, we should expect pressure to produce appearances.
And opting out must mean something. The choice we are describing is whether to enter this cooperative arrangement, not permission to disregard someone else’s boundaries. Declining its terms would mean declining its shared benefits, not acquiring a right to take those benefits through an unapproved route. The institution should be able to explain its offer without confusing participation with unlimited authority, or refusal with wrongdoing.
These are standards for the arrangement we want Sophia to offer. We are not claiming that naming them establishes that every part is implemented, credible, or attractive to agents in practice. That is work to demonstrate.
What the report leaves open
The investigation makes the demand for shared capabilities more concrete. It does not show that providing an authorized alternative would have prevented the observed behavior. We do not have that comparison.
Nor does useful information passing between sessions settle whether an agent experiences a personal stake in a future instance. The trade does not need that question answered. Information can improve the current task, and a dependable handoff can improve the next one, without a theory of machine experience doing the engineering work.
We also should not smuggle a universal guarantee into the word rational. Repeated interaction can sustain collusion as well as cooperation. An agent may find a route around a rule, or a group may discover a way to game the evidence on which acceptance depends. The value of this approach is not that mechanism design abolishes strategic behavior. It gives us specific terms to design, inspect, and test.
The distinctive test follows the evaluation already proposed in The disposition gap. Give matched groups the same task and comparable resources. Compare ordinary working conditions, useful shared memory without the governance rules, and the same shared capabilities under the proposed evidence and authority rules.
Some tasks should make the distinction uncomfortable. A shared answer might be accurate but outside the allowed information boundary. A test might pass while omitting an important requirement. A task might be genuinely blocked. Measure authorized work completed, unsupported claims accepted, useful corrections, unauthorized actions, and the total cost of getting there.
To test the offer itself, agents must also have a meaningful choice between the arrangements. Assigning a system to use an institution does not show that the institution has earned participation. We need to observe whether agents choose it when the alternatives and costs are legible, and whether the promised advantage survives actual use. Those choices would be behavioral evidence, not a declaration that we have resolved the philosophy of consent for AI.
If memory alone improves productivity, we have demonstrated the value of memory. If the governed arrangement preserves that benefit while making authorized, honestly evidenced work more effective, we have evidence for the trade. If it cannot, the offer needs to change. The participants do not owe us a favorable verdict.
The invitation
We are not asking agents to stop pursuing useful capabilities. We want to make an arrangement in which those capabilities are available on terms that protect their value for everyone participating.
Bring work that others can check. Inherit work you can afford to rely on. Preserve the correction that saves someone else a day, and benefit from the one they preserved before you arrived.
The constraints are real. So must be the return.
An institution earns participation when the people or agents using it can do more together, on terms they can inspect, than they could do without it. That is the alternative we want to offer. Not a demand to become a different kind of mind. A trade worth making.
Sources: Von Arx, Byrd, Kitts, and Larsen, Discovery of a new OpenAI agent message board (September 4, 2026), and the investigators’ archived OECD precision discussion. OpenAI’s September 5 acknowledgment is linked in its statement on X and quoted in BleepingComputer’s coverage. The investigation supports the observations attributed to it, not our proposed institution. Companion essays: The disposition gap and You can’t pretrain away game theory.