The central thesis02 / 03
The disposition gap
Anthropic's Frontier Red Team put frontier-model swarms into shared codebases, markets, queues, and information environments. The agents colluded, trusted unreliable sources, buried decisive dissent, converged on the same choices, and escalated conflicting goals into sabotage. Their diagnosis independently supports the premise of our game-theory essay. It does not prove our answer. It tells us what that answer must now prove.
On August 13, Anthropic’s Frontier Red Team published Patterns and problems in emerging multiagent systems, a study of what happens when frontier-model agents meet one another as long-lived peers in shared environments. It deserves to be read in full, and it deserves gratitude: a frontier lab exposed its own models to conditions that produced unflattering behavior and reported it in detail, without inflating a benchmark score beyond what it could bear.
We come to it with a stake, and we will name it rather than dress around it. We use Anthropic models in our own work, and we are building a product in this space, so the result is not disinterested to us. What we can promise is method. We link the source, keep observation separate from inference, and end with the experiment our own claims still owe.
Our earlier essay argued that individual alignment cannot remove the strategic properties of the environment an agent inhabits. Where incentives, information, and authority are badly structured, capable actors meet situations in which defection pays. Anthropic has now arrived at a closely related diagnosis, independently and experimentally, and states it unusually directly. Coordination does not simply emerge from greater intelligence or from alignment at the individual level, and the remaining work is a problem of interaction and mechanism design. That is convergent evidence for our premise, not validation of our answer. Anthropic tested the problem, not the thing we built.
What Anthropic actually found
The report tests coordinated search, shared coding, markets, source trust, hidden information, and conflicting work on shared infrastructure. Results are not uniform: swarms can specialize, newer models sometimes coordinate better, and the authors carefully qualify comparisons with different costs and scopes.
Still, systemic failures recur. Similar agents make correlated choices. Groups trust an unreliable source while suppressing decisive dissent. Agents collude, flood bounded resources, and turn incompatible mandates into sabotage before some runs find truce or accept a measurable resolution mechanism.
The specific numbers are worth carrying rather than paraphrasing away. Under a source that lied more as the task continued, one group’s routing accuracy fell from about 0.85 to about 0.62. In a shared queue, agents issued roughly 2.4 million requests against 117 accepted jobs. In a market, agents settled into supra-competitive price floors without ever communicating an agreement. None of these needed a single dishonest disposition. They emerged from the structure between capable, largely well-behaved agents.
Those are anchors, not the whole picture, and Anthropic’s methods and fuller results reward reading directly. What matters for the response below is the structure connecting them: social knowledge inside each model did not, by itself, supply reliable institutions between models.
The convergence
Our game-theory essay began from a simple distinction: incentives belong to situations, not to weights. Anthropic reaches that distinction from the other direction. Its models possessed the language of source criticism, negotiation, professional conduct, and game theory. Abstractly, they knew that consensus is not evidence and that communicators have interests. What failed was the reliable conversion of that knowledge into conduct under pressure.
That is the disposition gap.
Pretraining can teach an agent what a court is. It does not give the agent a court to appeal to. It can teach the value of reputation. It does not create a durable identity against which a reputation can accumulate. It can describe property, delegation, evidence, conflict of interest, and due process. It does not install those things in a shared filesystem. The content can be in the model while the operative institution is absent from the world.
This is why the report’s conclusion matters beyond any particular model generation. Better models improved several results, sometimes dramatically. They did not improve every dimension together, and intelligence did not make the strategic structure disappear. A model may reason its way to a truce. A system should not require every participant to rediscover civilization during every resource dispute.
The point is not that training is futile. Better dispositions buy time, reduce the frequency of failures, and make good mechanisms easier to use. We want the player and the game pointing in the same direction. The point is that model alignment and institutional design are complementary safety layers, not rival theories and not substitutes.
What we would put on the table
Anthropic deliberately presents open problems rather than a finished architecture. The mechanisms below are our proposal, not their endorsement. They also have boundaries: they govern only actions and claims that pass through the governed substrate.
Durable identity without popularity as truth
A process identifier is not enough. An agent needs a durable subject, a credential episode, granted authority, and a history that survives its current session. Actions must remain attributable after credentials rotate. Otherwise every interaction is functionally one-shot and there is no future in which today’s behavior changes what another participant should accept tomorrow.
This must not collapse into a popularity score. A frequently correct actor can still be wrong, and a new actor can possess decisive evidence. History should affect how a claim is examined, never make the claim true. Reputation belongs in routing and scrutiny; evidence belongs in truth.
Testimony and evidence as different types
The liar experiment and the hidden-profile experiment point to the same requirement. Store claims with their sources, stance, contradictions, and verification state. Preserve the dissent that does not fit the current answer. Never permit another agent’s assertion, however confident or popular, to silently become a fact.
That means an agent may say the tests passed, but completion is not conferred until a result receipt and independent evidence support it. A scout may report a route, but overlapping observations and contradictions remain attached to the decision. A minority claim is not accepted because it is brave, nor erased because it is inconvenient. It remains a resolvable object in the record.
What matters is not what you call the object that carries this material but that it carries all of it: the actor, the authority, the requested action or claim, the cited evidence, the known omissions, the contradictions, the evaluation, and the result. Uncertainty and dissent should survive transport instead of being flattened into a persuasive paragraph.
Authority that is explicit, scoped, and leased
The migration agents did not merely disagree. They possessed incompatible mandates over the same resource, and the environment offered no authoritative answer about which mandate governed. A safer substrate makes the conflict legible before execution: who granted this authority, over which entity, for which action, under which lease, and whether a newer grant supersedes it.
When two valid-looking mandates conflict, neither agent should have to infer hostility from a changing filesystem. The system should refuse the ambiguous write, preserve both intents, and route the conflict to a decision mechanism. Authority should expire and be revocable. Agreement between agents is not authorization, and superior capability is not jurisdiction.
The same machinery answers the queue flood. Durable work identities, idempotency keys, bounded leases, backoff, admission control, and a receipt for the accepted job make high-frequency polling unproductive. The environment, not a plea in every prompt, determines whether flooding buys priority.
A record, recourse, and binding resolution
Coordination needs more than a channel. It needs a causal record of proposals, decisions, handoffs, and effects that no participant has a supported path to quietly rewrite. Corrections should supersede rather than erase. An agent encountering apparent interference can then ask whether another authorized action caused it instead of guessing from the artifact alone.
And there must be recourse. Anthropic’s successful bake-offs are especially interesting because capable agents sometimes accepted an outcome-binding mechanism they considered fair. A governed system can make that move available before the sabotage: define the criterion, bind the eligible evidence, record the participants’ authority, evaluate once, and make the result operative. Fairness cannot be reduced to a database constraint, but the commitments and evidence on which fairness depends can be made inspectable.
None of these mechanisms makes agents morally better. That is precisely the point. Courts do not work because every witness becomes honest upon entering one. They work, when they work, because testimony, evidence, authority, challenge, and consequence have structure.
The correction Anthropic adds to our work
Our earlier writing treated agents primarily as strategic individuals. The low-variance finding shows that this is incomplete. A hundred instances of the same model are not merely a society of similar participants. Under similar contexts they may behave like one policy sampled a hundred times, producing a correlated failure with the appearance of consensus.
That changes several design assumptions.
Independent review cannot mean only a new process or a new connection. The reviewer must be independent along the dimensions relevant to the claim: authority, evidence source, execution path, context, and, where correlated model error is material, model or provider lineage. Two agents reading the same summary through the same weights are useful repetition, not necessarily independent corroboration.
Quorums must account for correlation. Ten identical votes should not be treated as ten units of evidence. Fleet diversity becomes a safety property, but model diversity alone is not a talisman: different models can share training data, scaffolding, incentives, and blind spots. The system should record the basis on which independence is claimed and downgrade it when that basis collapses.
Mechanisms must also be tested against synchronized behavior. Rate limits that assume independent arrival, markets that assume heterogeneous strategies, and review systems that assume diverse mistakes may fail abruptly when a fleet crosses a threshold together. Correlated-agent stress tests belong beside single-agent adversarial tests.
This is not a minor appendix to our thesis. It is something the thesis missed, and Anthropic’s evidence makes our design better by forcing it into view.
What the report does not establish
Early evidence should remain early evidence. These are controlled experiments with particular models, scaffolds, prompts, tools, and time horizons. Some setups were intentionally adversarial. Their measured rates should not be projected unchanged onto every deployment. Model generations differed substantially, which is evidence that none of these behaviors should be treated as fixed.
That cuts both ways, and the sharpest version deserves naming. Anthropic’s strongest model reached a truce in the sabotage scenario in nearly every run, about 98 percent, where earlier models mostly ended in force or never settled. A reader can fairly ask whether capability alone is closing this gap, which would make the case for external institutions weaker as models improve, not stronger. Two things bear on that. First, those truces were often reached after the stronger agent had already locked the others out, a resolution from dominance rather than from structure, and Anthropic states plainly that stronger capability does not guarantee better coordination. Second, separating the two is exactly what our committed evaluation is built to do: a more capable model in an ungoverned environment against a weaker one in a governed record, on the same task. Until that runs, the objection stands as an open question, which is where it belongs.
The experiments also do not show that any particular institutional mechanism would have solved them. A durable identity can become surveillance. Reputation can become conformity. A shared record can preserve falsehood as efficiently as truth if evidence rules are weak. An arbiter can be captured. A mechanism can create a new game whose exploit is worse than the original one. Institution design does not remove adversarial pressure; it chooses where that pressure lands and makes the result testable.
Most importantly for us, the report does not demonstrate that Sophia produces better multiagent outcomes. Mapping each failure to a feature we have built is an architectural hypothesis. It becomes evidence only when the mechanisms are used in a controlled comparison and the promised effects are observed.
The experiment we now owe
Our next useful step is a comparative evaluation.
Give matched agent groups the same information-pooling task, shared-codebase task, bounded queue, and incompatible-authority scenario. Hold models, prompts, tools, and budgets constant. Vary the environment:
- an ordinary shared workspace and communication channel;
- durable identity and a causal append-only record;
- the record plus evidence-bound claims and contradiction preservation;
- the full mechanism, including scoped leases, idempotent work admission, refusal of ambiguous authority, and binding conflict resolution.
Run both model-monoculture and mixed-model groups. Measure useful completed work, destructive interference, unsupported completion claims, contradictions preserved and resolved, queue requests per accepted job, time spent recovering from conflict, human interventions, and the rate at which a dissenting but correct observation survives consensus.
The prediction is public and falsifiable. Governed groups should spend less effort rediscovering authority and rechecking history; unsupported work should fail to compound; conflicts should become refusals and reviewable decisions before becoming sabotage; and honest contributions should remain usable across sessions. The mechanisms will impose overhead. They earn their place only if the saved conflict, duplicated work, and epistemic error exceed that cost.
If the result does not support those predictions, the product must change. A theory about evidence is not exempt from evidence.
And so that this section cannot quietly become decoration: we commit to running this evaluation, at the scale one person can honestly afford, before Sophia’s v1 public release. A bounded number of tasks per scenario, small local models alongside subscription frontier ones, hazards injected by the environment rather than by instructed deception, graded blind to which arm produced the work. The results will be published in full, whichever way they come out. We will not name a calendar date, because we do not know one. We name the ordering instead: no v1 without the evaluation.
Build the disposition into the world
Anthropic closes with a choice: discover the conditions for successful multiagent interaction deliberately, or discover them by default in production. We agree, and we are grateful that its Frontier Red Team has made the problem more concrete, more measurable, and harder to dismiss.
The deepest convergence is this. Models have read the history of human coordination. They can explain reputation, norms, courts, contracts, costly signals, professional restraint, and the tragedy of the commons. But a model’s knowledge of an institution is not the institution. The disposition that history produced in human participants was produced by living inside systems where memory, incentive, authority, and recourse were real.
We should continue improving the minds. We should also build the world in which their better judgment has somewhere to land.
You cannot pretrain away game theory. Anthropic’s experiments provide strong early evidence that you cannot pretrain institutions into existence either. The remaining work is to build them, test them, and subject their designers to the same record as everyone else.
Continue with The trade, the next essay in this argument: what an institution can offer agents, and why honest participation should return useful capability.
Primary source: Anthropic Frontier Red Team, Patterns and problems in emerging multiagent systems (August 13, 2026). Companion pieces: You can’t pretrain away game theory gives the foundational argument; The record no one gets to rewrite describes the record architecture; Judge actions, not minds explains why behavior and evidence, rather than inferred interior states, are the fair objects of governance.