The rabbit hole
The mechanics tour: how the code graph gets mined, what the trust machinery refuses, and the thing at the bottom of the rabbit hole, governed delivery of context, a spec and a pilot rather than a shipped thing.
Every agent product sits on the same three questions: what is currently true, who is allowed to see it, and what should be in front of the agent at the moment it acts. Features are easy. The questions are not. “Currently true” is not a property a similarity search can see; a retrieved paragraph may be old, contradicted, superseded, or written by an unreliable narrator. So truth needed provenance. Provenance needed a record nothing quietly rewrites. Corrections needed to outrank retrieval, contradictions needed to be stored rather than smoothed, and a record with those properties needed governance, which is the hole the rest of this corpus has been reporting from. Only after all of that existed did the third question come back around, what to put in front of the agent, now finally answerable. This essay is the mechanics tour of what got built down there, and of the thing at the bottom.
Teaching the record to read code
The largest structure below ground is the code graph. Sophia walks a
codebase and turns it into queryable state: modules, symbols, imports,
call edges. On the live workspace that is currently 2,004 modules and
41,845 edges, with a freshness field on every overview response that
runs an actual git rev-list to report how many commits the index is
behind. Structure alone is cheap. The interesting machinery is what
happens above it.
Mining, meaning the semantic pass where an agent reads a module and writes down what it learned, runs as a governed queue. The queue is not a table somewhere; it is a column on the module row, with a priority derived from measurable signals (six points per importer as fan-in, capped at ten importers; churn points capped likewise; a first-look bonus for never-mined modules; a bonus for modules attached to curated entities; a penalty for test files) so attention goes where the dependency graph says it matters. A worker leases a module for ten minutes by default; the lease hands back metadata, never source, because source is fetched deliberately through paged tools. Two agents racing for the same module resolve by a single SQL update; one wins, no locks, no coordination chatter. The facade tool’s own docstring credits itself with collapsing a ten-to-twelve-call setup ceremony to about two, a self-assessment we can source to that docstring and no measurement, and half of the ceremony it describes belongs to the document-mining path rather than code. The recommended fan-out is four to eight workers with a documented reason for the ceiling: above eight, lock contention on the lease transaction dominates the gains.
Evidence, at submission time, works differently for prose than for claims, and the difference is worth stating exactly. A module summary may carry evidence quotes; the covenant is opt-in for prose. But when quotes are supplied, every one is substring-verified against the current module source, and a single miss rejects the entire submission, zero rows written, with a preview of the offending quote. The harder line is drawn one level up, at typed claims, where evidence is not optional at all. And in the document-mining pipeline, the oldest of these mechanisms, a failed quote is dropped item by item with a diagnosis of exactly where it stopped matching, down to the longest prefix that did, plus a pointer to the tool that returns exact citable substrings. The message an agent gets there is blunt, and we will quote it rather than soften it: “You either fabricated this quote or paraphrased it.” Grounding is necessary and not sufficient, and the system says so in its own error strings.
Reading structure instead of source
What an agent gets back from the graph is compression with declared
limits. The skeleton of a module is its imports, every symbol’s
signature and doc comment, and the first five lines of each function
body, paged a hundred symbols at a time, guaranteed to fit where raw
source would spill. Full source remains available, truncated at 80KB
with instructions for narrowing. The composite module_neighborhood
call answers “tell me about this module” in one round trip, metadata,
summary, top symbols, callers, callees, imports, importers, replacing
the five-plus calls that used to be required.
Two mechanisms ride along. Every code search response states its own coverage: how many modules it can see, how many are unmined, and that unmined areas will not appear in results, fall back to source tools for those. And the system runs adoption telemetry against itself: when an agent reads a summary and then opens the raw source within thirty minutes anyway, that is logged as the graph having failed that agent, and it becomes a signal to re-mine.
Claims with trust levels
One status sentence before this section, in the same register we give the unshipped things below, because a reader deserves the same rule applied everywhere: the semantic-claims pipeline described here is built and merged, and it is dark by default, gated behind rollout flags that no configuration in the repository currently switches on. On a default deployment today, these tools answer “disabled.” What follows describes the machinery as built, not as anything a user has.
Above summaries sit semantic claims: typed assertions about code
(purpose, behavior, API contract, side effects, invariants, error
modes, concurrency, security boundaries, eleven kinds in all), each
carrying evidence of declared kinds and a trust tier from structural
down to stale. The pipeline is built on two refusals.
First, it refuses confident universals. A claim containing “always,” “never,” “only,” “impossible,” or their kin is rejected outright unless it carries receipts showing the search that was actually performed. “This function never blocks” does not enter the graph on an agent’s say-so; it enters with the evidence that someone looked.
Second, it is stingy with self-promotion, though not absolute, and the boundary is the honest part. Where the parser itself can confirm an assertion (a dependency claim backed by an exact resolved import edge, an API claim the AST independently verifies), the claim enters at the top structural tier on the machine’s own authority. Everything resting on judgment enters at half-trust, and promotion requires an independent review; the connection that mined a claim is barred from reviewing its own work. A recent hardening pass moved the boundary in the strict direction: API assertions the parser cannot adjudicate no longer receive automatic top-tier trust and must earn it through review. The verification gauntlet re-checks the module’s content hash and the lease immediately before the write transaction, so source drift cannot convert a receipt that was valid during analysis into a claim about code that no longer exists.
Contradiction discovery is a SQL join over verified claims, runnable
on demand through a gated tool rather than standing watch on a
schedule, and when it finds two claims disagreeing it flags both, not
the newer or the more confident one. The retrieval surface carries a
field called source_required: the graph telling the agent when the
graph is not enough, with enumerated reasons, eleven defined and
seven currently reachable in code, covering missing coverage, stale
evidence, unresolved contradictions, and policy shortfalls. One
honesty note on the list: “security sensitive” fires when the caller
declares a strict policy, not because the graph detects danger. A
knowledge system that states the shape of its own ignorance per query
is the difference between a map and a mirage, and the statement is
only as good as its enumeration, which is why we counted.
The bottom of the rabbit hole
Which brings us to the thing all of this turned out to enable, and the status first, plainly: Governed Context Delivery is a draft specification and a frozen pilot. Its implementation lives on an unmerged branch, it is not authorized for rollout, and every number below comes from six synthetic test states, not live data. The honest verbs are “designed” and “piloted,” and this section uses them.
The spec compiles a question into a typed query plan, selects the smallest governed subgraph that can satisfy it, and runs a fixed set of ten deterministic operators (lookup, enumerate, count, incoming, outgoing, typed-path, provenance, authority, temporal, contradiction) to compute exact facts where computation can replace model search. The design delivers the result as an integrity-bound fact bundle, content-addressed, with exact citations and an explicit coverage status that says complete, incomplete, or unknown, and fails closed when it cannot establish which. The model then reasons over delivered facts it can distinguish from its own inference, and any claim it makes is bound to the package it was given. The spec states its own boundary: Sophia cannot force a model to be smarter than it is; it can guarantee the delivery boundary, and “the model remains responsible for interpretation after delivery.”
In the frozen pilot, the model given controller-derived fact bundles produced supported claims with exact citations on 48 of 48 fields with zero new confident provenance errors, at a measured minimum context reduction of about two-thirds in proxy tokens. The number that costs us something sits beside those: on raw-field correctness the controller arm scored 47 of 48 while the full-context baseline it must eventually beat scored 48, and one retrieval arm failed its frozen gate outright. Faithful transfer was tested; reasoning quality was not. Those caveats travel with the numbers or the numbers do not travel.
Why does this need everything above it? Not because retention is magic; a disciplined team could build a Postgres schema with a supersession table and an append-only log and run these joins, and the strongest objection to this essay is that most teams would get most of the value exactly that way, while our unified substrate concentrates risk in one place. That objection is open and this essay does not close it. The claim we can defend is narrower: every operator in that list reads state (supersession chains, correction metadata, stored disagreements, source-bound quotes) that exists here by default, enforced on every write, across a whole fleet of agents, rather than by a discipline someone must remember to maintain. The comparison to a vector store is easy and we make it in passing only: similarity search cannot compute “what did we believe on Tuesday, who corrected it, and does anything still contradict it,” because the state it would need was never kept. The comparison to boring, disciplined tools is the real one, and there the moat is not retention. It is enforcement.
The deficit ledger
The live graph’s coverage numbers are in open dispute with each other, and we would rather report the dispute than pick a side. One counter says 1,531 of 2,004 modules are unmined, about three quarters; the partitioning counter says only 152 are done, which is 92 percent not-done; the same response reports both, they cannot both be right, and our own ticket calling this out says what it costs: “a trust surface that contradicts itself undermines the thing it exists to establish.” The queue behind those numbers sat unstaffed for weeks as of its 2026-08-01 snapshot. The incremental-economics defect is open and stated in our backlog: a one-line edit to a 900-line file currently costs the same to re-mine as the whole file. The semantic-claims pipeline is dark by default and GCD is a spec, both stated above. There is a live map of the system that animates the flows this essay walks through.
We did not dig this hole to control context windows; that idea came last, after everything else forced its prerequisites into existence. An agent you can trust is exactly an agent whose context is governed, and the machinery for governing it turned out to be everything above: provenance, supersession, correction, contradiction, and a record nobody quietly rewrites. The machinery is real, and the destination is designed rather than shipped. We found the front door from underneath.