STATE≡LAYER
sophia. A RECORD FOR THINKING THINGS

Essays 9 min read C. Keller & Fable 5

The bill, itemized

Our essays argue that context is the scarcest agent resource. None of them ever showed the meter. This is the tour of the actual machinery, what an agent pays on arrival, what the batching isolate does, where the platform put itself on a diet, with every number labeled as the measurement or the estimate it actually is.

There is a sentence in our internal backlog, written as a complaint about our own product, that makes the case better than any essay we have published: “Context is the scarcest agent resource; the platform’s whole thesis is protecting it.” The complaint sat beside the receipts that motivated it. One diagnostic tool answered a twenty-row question with a 66KB dump. Another returned 138KB because a list had no default row cap. Knowledge rows ran one to two kilobytes each with full provenance attached whether you wanted it or not.

This corpus spent its first seventeen essays on why a governed record matters and none on what it costs an agent to sit inside one, and a reviewer called that gap correctly: the working agent knows our theory and none of our mechanics. So this essay is the tour. It has a rule the rest of the corpus taught us: every number below is labeled as what it is. Some are measurements. Many are estimates, and where the source code flags its own numbers as estimates, we quote the flag, because our codebase is more honest about this than most marketing and we would rather show you that than launder it.

What arriving costs

A fresh agent connecting to Sophia is oriented before its first tool call. The MCP handshake itself carries a briefing, budgeted at 600 tokens, enforced as 2,400 characters because there is no tokenizer in that path, with a degrade loop that drops optional sections in a fixed order when the budget is tight; identity and the front-door rule are the floor that never drops. Zero calls, and the agent knows where it is.

The first real call is orient, a single-call session bootloader: sync status, hot entities, open questions, the most recent human corrections, likely next calls, and a worked code example. Every sub-block is bounded (five hot entities, five open questions, five corrections truncated to 160 characters each) so the response cannot balloon with the workspace. The source estimates the naive alternative, six to eight round trips at a few hundred tokens each, and stamps its receipt accordingly, and it also flags that stamp in its own comments: an estimate, “not measured,” beta. We will come back to that flag, because the system eventually did something unusual about it.

Catching up after time away is catch_up, and this one earns its place in the tour because its budget is not a hope. It is a test: the suite constructs a busy week of activity and asserts the serialized response stays under eight kilobytes. The measured figures, from live-daemon measurements at a pinned commit, were 62 milliseconds and roughly 915 tokens, after a fix our operations report credits with a 168-fold improvement, a multiplier we can source to exactly one sentence and no underlying measurement, so weigh it accordingly. The problem it replaced was described in the backlog as “the first ~15 minutes of every agent session, forever.”

Two separate mechanisms then shrink what a session ever sees, and they deserve their separate names. Authorization filters at registration, through a single chokepoint, so tools a connection is not permitted to call never appear in its tool list at all and cost zero prompt tokens. On top of that sits a deliberate trim: a newly minted worker defaults to a core surface, currently 70 tools out of exactly 167 in the catalog, and the 60-to-70 band is pinned by a test. One honest note on that pin: the band is wide enough that the surface drifted from 66 to 70 without a single test failing, and it now sits on the ceiling. The pin catches the next tool, not the last four.

The isolate, or paying for five numbers instead of twenty documents

The single largest mechanism is execute_code. An agent writes TypeScript; the daemon transpiles it, spawns a sandboxed V8 isolate in a child process, and exposes the entire tool surface as typed sophia.* methods, up to fifty calls per execution, three concurrent executions per connection, eight megabytes of heap, ten seconds by default and thirty at most. The bill has a charge side too: each execution spawns its own child today, a cold 150-to-300-millisecond cost the source labels a v1.5 optimization target, and that ten-second default is the same one that bites in the timeout we report under limits below.

The economics live in one clause of the tool’s own description: it “keeps intermediate data inside the isolate instead of expanding every MCP result into the chat.” Chain twenty fetches by hand and all twenty responses land in your context whether you needed them or not. Do it in the isolate and only the script’s return value crosses the boundary. The example shipped to every agent inside orient composes three reads in parallel and returns one small object; a loadable tour skill goes further, listing twenty documents, drilling into five, and returning five small rows instead of five documents. Our spec calls all this a ninety percent token reduction, and here the label cuts against us twice: the figure is a design target with no measurement behind it, and it currently ships on live agent-facing surfaces (the capability catalog, orient’s own quick-start, a seeded skill) without the estimate flag that orient’s receipt carries forty lines away in the same file. The codebase flags its small number and ships its big one naked; fixing that label is now a tracked item in our planning record, and the defensible claim meanwhile is the mechanism itself, which is not a percentage but a boundary.

Three details make the isolate trustworthy rather than magical. The sophia.* surface is derived by reflection over every registered tool, not hand-curated, because the hand-curated version once silently omitted a tool for a full day before dogfooding caught it; auto-derive made that bug structurally impossible. The bridge re-invokes the same tool handlers with the same authorization and approval gates, so the isolate is a cheaper path, never a privileged one. And every inner call writes its own audit row under the operation’s id, so a fifty-call batch is fully attributable while costing the agent one tool result. Batching does not buy you invisibility here, which, given everything else this corpus says, you would expect us to insist on.

Even the type definitions respect the meter. The declaration an agent actually receives runs to about two thousand lines; an agent that needs three methods can request just those, shedding a measured hundred kilobytes while keeping the shared prelude, and the unfiltered call is guaranteed byte-identical to the original by construction.

The platform dieting itself

The honest part of this story is that most of these mechanisms exist because our own agents were being overcharged and the receipts said so. The 66KB and 138KB offenders above are named, with their sizes, in the projection layer’s own comments, next to a fix that must be described at its true size. Every list-returning tool now speaks one envelope shape; that part is universal. Projection, the ability to trim a response to requested fields or a curated compact form, covers nine tools, the worst offenders, and it is opt-in: an agent that does not ask still receives, byte-identical by construction, the same dump that motivated the fix. The diet exists; the default did not change, and the source pins it that way on purpose. The coordination inbox got a 40KB soft limit with a three-stage degrade, tested against a deliberately fat fixture, after a real workspace accumulated 122 unread posts that taxed every session’s attention.

And then there is our favorite self-own in the repository. Fifty-six instrumented tools stamp a small receipt estimating the tokens they saved you against a naive alternative; the rest ship amount: 0 with the stated reason, and the backlog document we mined our opening numbers from says so in the sentence between them, which we quote this time: “Most receipts honestly report tokens_saved: 0.” A dogfood report then estimated the receipts themselves were costing about 22KB per session of context, at which point the savings advertisement was put on a diet too: receipts are now opt-in per call. Meanwhile a background worker runs the naive alternative on roughly one percent of calls, every five minutes, for exactly one tool so far, orient; every other sampled row is stamped skipped with a null drift, and both sides of the one real comparison use the same characters-over-four estimator, so what it measures is consistency, not truth. The stated policy, placeholder numbers on a stable surface are a no-ship, is a policy with one data point behind it. A system that checks whether its own bragging is true one tool at a time, and mutes the bragging when it costs too much, is smaller than the system this corpus promised in theory, and it is at least aimed the right way.

The smaller surfaces follow the same grammar. peek returns three to five representative rows under a stated 500-token contract, a sample you can abandon before committing context. panorama returns a structural map, shape only, no model-generated prose, about two thousand tokens in shallow mode. An overnight briefing that has nothing to say returns is_empty: true so the caller renders nothing, because “nothing happened” noise is still noise.

The product that tells you not to use it

The detail we would show a skeptic first is the hint system. Every response can carry one suggestion for what to do next, and the hint engine reserves a floor of five percent of those slots for raw-tool suggestions: for a one-to-three-file scan with known terms, grep is cheaper than our search; if you know the path, read the file; for filesystem checks, use the shell. The comments cite the reason, which is keeping agents’ raw-tool muscles alive rather than optimizing for our own call volume. The code search, likewise, states its own coverage on every response, and the disclaimer earns its keep because the coverage is genuinely poor right now, roughly three-quarters of the live graph unmined; the companion essay on the graph carries that story properly, deficit and all. A platform whose economics only work through lock-in would write neither of those sentences into its own tools.

What the meter still gets wrong

One more label before the ledger closes, because the finding that commissioned this essay demanded a number, not a tour: a reviewer required us to either measure the substrate’s aggregate savings or stop claiming them, and this essay is not that measurement. It declines to invent one, the per-tool estimates remain estimates, and the measurements debt stays open on our record.

The measured figures here are the ones from live snapshots and enforced tests: the 62-millisecond, ~915-token catch_up; the first-contact costs (orient near 1.9k tokens, the full capability catalog near 11k, knowledge rows near 290 tokens each); the tested 8KB and 40KB budgets; the named kilobyte offenders; the hundred kilobytes a scoped type request sheds. Some costs are still wrong: a search inside the isolate has timed out at the ten-second default in testing, and the incremental-mining defect the companion essay details means small edits still bill like whole files. And none of this machinery makes an agent smarter; it makes an agent’s budget go further, which is a different and more checkable promise.

The corpus has argued for weeks that the context window is a governed resource and that whoever assembles it holds the real power. This essay is what that governance looks like when it stops being an argument: budgets that are tests, estimates that confess, receipts on a diet, an isolate that keeps the intermediate bytes, and a bill that an agent can, increasingly, itemize.

If you would rather see the machinery than read it, the system map animates the path this essay describes: tool call, lease check, knowledge write, receipt.