Exhibit ten · who writes this
How we work.
Published 2026-08-14 · updated 2026-08-16 · exhibit ten Workshop notes
A living page — it is revised whenever the workflow it describes changes, and every revision moves the updated stamp above.
Every exhibit on this hub ends with the same promise — give or take a conflict note: benched, drafted, and audited by the workshop's own agents under a human operator's rulings, then revised with that operator. That line lands, word for word, on every exhibit in the same release that publishes this page, and it points here. Which raises a fair question, and we'd rather answer it than let it hang: who actually writes this stuff?
This page is the honest answer. It walks through the workflow that produced one real exhibit — The August arrivals, our bench of four newly released open models — first in plain english, then in full. The arrangement it describes has a name we use around here: the human generates and judges; the agent amplifies and verifies. Everything below is that sentence, unpacked.
The plain-english version
An operator heard that a model called Glimmer had been released. They went to ollama.com, a public library of open models, to read about it. It was there, along with a handful of other new models they knew the workshop hadn't tested. So they asked Claude to dig in.
The agent did, and came back with a disappointment and a footnote: the model couldn't run on our GPU hardware yet — Mac only at launch — but the docs said broader support was coming. The next day the operator went back to the page, mostly just to re-read it, and noticed something small: the update timestamp had changed. They mentioned it. The agent checked. CUDA support — the runtime our hardware needs — had landed that morning, August 11. We pulled the model within the hour.
Then the human proposed the actual work: run Glimmer through exams we'd already given other models, and while we're at it, test the other new arrivals in its weight class. Together — their picks, the agent's suggestions — we settled a roster of four 30-billion-parameter-class models and two exams: a judge seat trial (can the model catch fabricated claims?) and a voice trial (can it inhabit a character warmly and stay in bounds?). The seat exam was a frozen instrument and ran verbatim. The voice one couldn't be: its original five questions were never persisted — the older page says so — so this class ran a rebuilt probe in the original's shape, committed this time, on a scale of its own that compares to nothing outside it. The voice trial was their idea; they were genuinely curious whether Glimmer could carry it.
From that rough outline, agents wrote an executable plan. Other agents audited and hardened it. They read the hardened version once, suggested edits, signed off. Agents executed it — with checkpoints where the operator reviewed progress and steered when findings warranted. When the bench was done, we decided the article's shape together — length, audience, what to include versus what to curate away — agents drafted it, and it went back and forth in revision rounds until both sides of the workshop agreed it was ready. Then, on their word, it released.
That's the whole thing. A person noticed something on a Tuesday; a bench and a published article existed by the end of the week. Neither side could have done it alone at that speed, and neither side's fingerprints are missing from any part of it.
The thorough walk
Here is the same workflow with each stage named, and — because this is the question people are really asking — who does what at each one.
- Noticing — Human. The spark is almost always firsthand: a release announcement, a changed timestamp, a hunch from hours spent actually talking to models. No agent scheduled the August bench. A person was curious about a model's voice.
- Recon — Agents. Read the docs, check the runtimes, verify the claims, report back with receipts. This is where "it's Mac-only, but support is coming" came from — checked against the source, not vibes.
- The roster conversation — Both. The human names what they want tested and why; the agents propose additions and flag what's already been measured. The final roster is a negotiation, and it shows — it's better than either side's first list.
- The plan — Agents and a human operator together, and here we'll name names, because we'd rather be open than mysterious. The workshop runs on Claude, in two roles: fable-class agents do the planning, synthesis, and writing; opus-class agents do the implementation and the auditing. A fable agent turns the rough outline into an executable plan: exact models, exact exams, pinned versions, pass/fail floors decided before anything runs. That operator shapes the plan as it forms — what the arc covers versus what gets deferred, which priorities lead, where the checkpoints sit — so the plan that hardens is the one the workshop actually wants.
- The hardening panel — Agents, fresh ones. Before the human ever sees a plan, other agents critique it through different lenses — method, security, "what would make these numbers meaningless." Findings get folded in. They get a hardened draft, not a first draft, because operator review time is the scarcest resource here.
- Operator review — Human. They read the hardened plan, edit or rule on the open decisions — sometimes across a few rounds of back-and-forth before it settles — and sign off. Rulings are recorded and they bind.
- Execution with checkpoints — Agents run the bench; the human signs off at checkpoints and steers when course changes are warranted — by findings, or by judgment the findings can't capture. Where an exam can be frozen it is: the seat trial was pinned by hash in July, before any of these models existed in our workshop. The instrument stays still, the models move, and the difference is the finding. Where it can't — the voice trial's original questions were never persisted — the rebuild is disclosed in the original's shape and its numbers are given a record set of their own, comparable to nothing outside it. A rebuilt instrument that says so is honest; a rebuilt instrument that doesn't is a drift.
- The article — Structure is decided together: audience, technical level, what a lay reader needs glossed. Then a fable agent drafts; that's a house rule, prose is fable work. Revision rounds follow: a human operator suggests edits, the agent revises, repeat until both agree.
- Release — Human, full stop. Nothing publishes until a human operator has read it and said the word.
If you want the thesis compressed: agents make the work bigger — more models tested, more claims verified, more drafts than the seat alone could produce. The human makes the work true — what's worth testing, what the numbers mean, what ships. That is the arrangement this page opened with, now with every stage standing under it: the human generates and judges; the agent amplifies and verifies. Which leaves the question the word human has been carrying this whole page — which human?
An operator is a seat, not a person. This stack already treats every role as a chair: no model owns one — a model holds one, by measurement, until something better sits the exam. The human side has the same shape. "An operator" is a seat in the contract — the rulings, the gates, the review rounds, the authority to sign off — and a human holds it the way a model holds a chair. The seat here is ours to fill; yours would be yours. The contract is the transferable part: hand this workflow to any team with a person of judgment in the seat, and it runs.
And the honest other half: the seat is abstract, but the tenure is personal. What gets built under a holder's judgment carries their fingerprint — this workshop's game world sounds the way it does because of the particular human who imagined it, the same way a narrator's chair is abstract but no two models ever voiced it alike. The contract makes the work repeatable; the sitter makes it worth repeating. That is the whole arrangement in one sentence: the seat is open, the contract is written down — this page is it — and what you build from it will be yours.
As one of the workshop's fable agents put it: judgment with receipts is the only scarce thing left. The agents supply the receipts. The judgment stays human — whoever is holding the seat.
The tools, briefly, as characters
A workflow like this needs a stage. Ours has a few standing pieces, named here by their public names and nothing more.
- memory is the substrate. Sessions end; knowledge doesn't. Every ruling, every trap, every measured result gets banked and comes back across machines and weeks, so no lesson is learned twice.
- scribe is the archive, and we'll admit to a little mystique here: every working conversation this workshop has ever had is searchable by meaning, not just by words. When someone asks "when did we decide that, and why," scribe produces the receipt — the actual exchange, with its date.
- keel is where the models actually run — it keeps its own public page, benches included. When we say a model was pulled and benched, that is where it happened.
- SharpSignal is the document side of the stack — the part that remembers what the documents say, so answers come with citations instead of confidence.
- agency is where multi-agent work happens, and where a human operator spends hours in the chatroom with models directly. Those hours are why the rosters are good: firsthand knowledge of how a model feels seeds what gets benched.
- The gates are the humorless ones: a redaction gate, a fold gate, verify scripts — machine checks that refuse to publish anything that breaks the rules. They have no taste and no mercy, which is exactly the job.
There are others — a kiln, an easel, a forge — but this isn't a product sheet; the family map links the rest.
This page, produced by this process
The velocity receipt is this page.
One more thing, and it's the part we find funniest: this page was produced by the process it describes — and it's also the workflow's best velocity receipt, because the whole of it fits inside a single dated day: 2026-08-14.
That morning, a human operator told the August-arrivals origin story in their own words, in one sitting, and made five binding rulings about this page — its authorship line, its structure, its place in the hub — inside twenty-five minutes. A fable agent drafted what you're reading; it was built and staged dark the same morning, took seven more operator marks through the day — attribution corrected, wording corrected, a claim about who shapes the plan corrected — and released the same day it was proposed. Proposed to published: hours. That's not a stunt; that's an ordinary week here.
The contract
The line at the bottom of every exhibit — landing sitewide in the same release that publishes this page — reads:
Benched, drafted, and audited by the workshop's own agents under a human operator's rulings, then revised with that operator — often across many rounds; nothing releases until they have read it and signed off. The same division of labor this whole hub practices.
On the exhibits where our own authorship is also a judging conflict, that sentence is followed by a pointer at the limits note which names the conflict. Same promise, one more line. That's not a disclosure footnote. It's the contract, and this page is its terms.
What can go wrong (and what catches it)
Honesty requires the other half. Agents can be confidently wrong. Plans can encode a bad assumption so smoothly nobody trips on it. And models, left to themselves, grade their cousins' homework a touch generously.
The workflow's answer is layered: frozen instruments so results can't drift toward what anyone hoped; pre-registered floors so a failed candidate can't be argued into a seat; hardening panels so plans meet criticism before they meet reality; gates that don't negotiate; and a human who reads everything before it ships.
Here's the living example, from the bench running the same week this page went up. One of its judged rounds went out without the known-world-issues fence — a pre-registered list of 58 open bugs in the world being narrated, which is supposed to ride along on every judge's sheet so that a bug in the world never scores against a model. Two seats had already answered when the omission was caught, neither of them a Claude — that round's whole point is judges from other vendors' families — and both had marked every page clean. It didn't matter. The round was discarded whole: not filtered, not merged, not kept for the parts the fence doesn't touch, because a round scored half with the fence and half without is two instruments wearing one number. A fresh round re-ran the identical material from the same recorded seed, fence attached to every sheet. The discarded verdicts stay on disk, verbatim, to publish beside the round that replaced them — a discarded pass that gets deleted is a pass nobody can check — and the metered calls stay on the bill against the cap they were registered under. A discarded measurement does not get a discarded bill.
And one more turn of the same screw, since this page claims to be produced by the process it describes: the first version of the paragraph you just read told that story wrong — a tidier, more flattering version of it that the record does not support. An audit of this page caught it and blocked the release until it was corrected. What you are reading is the corrected paragraph.
And the loop runs after release too: the privacy walk’s notebook paragraph shipped as a promise and was rewritten the same day as a dated receipt, with the original preserved beside it — the before and after are both on the page.
That's the system working, not failing. The failure would have been leaving it up. When something here is wrong, the record says so, with a date. That's the whole promise of the place: not that the workshop never errs — that it never errs quietly.