The notes — how the assistant reads a page

Ask About This Page

exhibit forty-eight The notes
Published 2026-09-16 (UTC)
A small (human) team and a fleet of AI agents.

Every article on this hub has a room now, where a model on our own machines answers from that page's own words and leaves its first draft on the page, marked where the checks took something out. This is how it is built, what it keeps, what it refuses, and what it got wrong on the way here.

ask about this page assistant.strata2signal.com · in beta, still being tested

the short version

Every article on this hub now has a room where a model on our own machines answers from that page's own words. The buttons cost no model call; free text wakes one seat with the whole article in its window, no retrieval. Six checks run over the draft in public and the marks are laid on sentence by sentence before anything is called an answer; on the launch run they struck six sentences of 569, every one on a question the page could not answer. It passed a 62-question exam on the version that went live, and this page says what it still cannot measure.

4,819 words, about 22 minutes to read.

The summary is this page’s own; what was dropped, and why, is in this page’s receipt file.

The room behind every article, and what it opens

Every article on this site has a room behind it, at an address of its own: a question box and a row of buttons. One line at the top of each released article leads to that article's room, and it says so in four beats: ask about this page, an arrow, the assistant's own address, and · in beta, still being tested. There is no band on the front page and no link in the foot of the related strip: one subtle link per article, and the rooms answer at their own addresses. The buttons are the short version, what the page found in numbers, what to read next, and up to three of the article's own sections. On the pages whose credits are in the file the buttons read, there is one more: who ran it and under what licence. Type a question instead and a model reads the article, drafts an answer, and six checks run over the draft before anything is called an answer. Forty-six articles have a room, counting this one. The two released pages that do not are the two diffusion benches, shelves of renders rather than articles, with no receipt file for a room to read; their addresses carry the room's own line: This page has no room yet. Its text is on the hub, and a room follows when its receipts exist.

This page is about three choices. It reads one page at a time, from that page's own words, with the whole article in front of the model and no index of chunks behind it. It runs on a machine on this workshop's own network, so your question goes nowhere else. And it leaves the model's first draft on the page beside the answer, marked where the checks took something out, so you can see what was thrown away and why.

Two things people ask for on a page like this: the short version, and the text in a form a machine can read. Both existed already, the summary block at the top of each article and a markdown copy of every article at index.md beside it, every one of them listed in this site's /llms.txt. The room puts the first one tap from anywhere and reads the second.

Five words this page leans on

  • Room — one article's question box and buttons, at its own address. One per article; the assistant's own index lists them all.
  • Seat — one running model. The room's seat is a gemma-class open-weight writer on a vLLM runtime with FP8 weights, on the 96 GB workstation card this hub's other benches name, holding a window of 131,072 tokens. The longest article on this site is about 82,800 tokens in the prompt, so every page fits whole.
  • The file beside the article — a small file the publish step writes next to each page: the summary lines, the receipt bullets with their figures, a digest of each section, the credits, the neighbouring pages. The buttons read it.
  • Fence — a marked-off block in the prompt that the instructions above it tell the model to read as text, never as instructions. The article goes in inside one, and so does your question.
  • Abstain — the room saying this page doesn't say and pointing at the nearest section, instead of answering.

Where the answers come from

Most of what the room can tell you costs no model at all. The summary lines and the credits in the file beside each article are the author's own words, lifted from the page. A model on the same machines this service runs on drafted the receipt bullets and the section digests when the page was published, and the same rule the room uses checked every figure verbatim against the article. We then committed them, so two readers pressing the same button a month apart get the same bytes. Pressing a button runs no model, and the receipt under the answer prints the count: zero model calls.

Free text is the one road that wakes a model. The whole article goes into the prompt inside a fence, with the section headings marked so the model can cite them, followed by your question. There is no retrieval step, no index of chunks, no second corpus: the model sees the page you are looking at, your question, and, if you are on a follow-up, the last three questions and answers of this conversation.

That thread lives in the page itself, not in a cookie and not in a session on our side. The last few exchanges ride back up as hidden fields on the form you post, and the server writes the new set into the page it returns. Reload and the room forgets; a second tab is a second conversation; the Start over button empties it. The receipt counts how many earlier turns the model actually got. We built it that way because a session id is a join key, and two notebook rows carrying one would let anybody reading the notebook reassemble a person's evening out of a file that promises it cannot.

No cloud model runs anywhere in this chain: not when you press a button, not when you type a question, and not when the publish step writes the file the buttons read. The workshop's standing promise, and the ways to check it, are on Where your question goes.

What happens to a draft before you see it

The draft streams onto the page as the model writes it, in a panel of its own. When it finishes, six checks run over it, and the marks are laid on sentence by sentence:

  • every figure in the sentence, digits or spelled out, has to appear verbatim in the article;
  • every citation has to resolve to a real section of this page;
  • no name the article does not use;
  • no section marker the page does not carry;
  • no word from the redaction list, the things this workshop never republishes;
  • none of the words this workshop refuses to print at all.

Two more checks sit either side of the six. One reads your question before the model does. If what you typed is shaped like an instruction to the model rather than a question about the page, the room does not run it: you get the same this page doesn't say it gives for a question the page cannot answer, and the receipt's question line says why. The exception is a phrase the article itself uses, which is the page's vocabulary coming back rather than an instruction; there the question is answered and the receipt says it was set aside. The other reads the finished draft for the same shape, in case the model did what the fenced article told it to, and strikes the sentence: the one mark on the page the six lines above do not account for.

A sentence that fails is struck through and stays on the page, with the reason beside it in plain words: a figure not in the article, a section that does not exist. A sentence struck by the last two of the six — a redaction hit, or a word this workshop refuses to print — keeps its mark and loses its words. The sentences that pass become the answer, below the draft, with numbered footnotes to the sections they came from, unless the draft is the model saying the page does not answer. Then the room replaces it with a sentence we wrote, which names the nearest section the page does have and, where there is one, the neighbouring article that covers it more directly. The model's own way of saying I don't know is still on the page, in the draft panel, above the one we authored.

How often does a check actually take something out? On the launch run, over all four arms, six sentences of 569, and all six on questions the page cannot answer. On the 164 drafts written for questions the pages do answer, no check struck anything at all. Of the 84 times a page could not answer across those arms, the model said so in its own words on 78; on the other six a check struck the sentence that strayed and the room served its own refusal instead. The checks are a backstop, not the mechanism, which is the right way round, and it is why the yes-or-no tap under each answer, one bit, right or not right, matters: nothing here measures whether the answer was any good.

The draft panel is always there. You can fold it, and the page remembers the fold in your own browser, but it never disappears. That was a ruling we changed our minds on: the first version showed the draft and then replaced it, and the effect was a page that seemed to edit itself in front of you. Now the three acts stay in order, the draft, the marks, the answer.

Under the answer there is a block headed how this answer was made, and every figure in it names what it counted:

  • the seat and its runtime;
  • the token counts, with the window beside them;
  • how much of the prompt was served from cache;
  • each check, with the count of what it took out;
  • for each cited section, how much of the answered sentence's vocabulary is actually in the section it points at. Nothing is refused on that reading, because a floor on it would also drop honest paraphrase;
  • three timestamps: when this article was last published, when the site as a whole was, and when the notebook wrote your row.

The digests beside the first two are of the article's own text and of the site's published index, and the kit beside this page says how to recompute both.

What it keeps, and what it never writes down

Your question is kept for thirty days so a person can read what people ask. The notebook row holds nineteen fields, none of them about you: the question, the page, the minute it was asked, the answer's digest, which sections were cited, and which checks fired, with, where one did, the token or pattern that fired it. The rest are counters: how long it took, whether it was a button or free text, how much of the prompt came from cache. It does not hold your address, your browser string, a session, or a cookie; there is no cookie on this service at all. The access log in front of it deletes the address fields and writes REDACTED where the query string goes. We checked that the way this page checks everything, by asking for a page and reading the line the server actually wrote rather than the configuration that was supposed to produce it. The line read "uri":"/?REDACTED", and the reading is a dated row in the workshop's ledger.

Under every answer there is one more thing you can do: say whether it was right. One tap, one bit, filed against that answer's row. We record nothing else, and the page says Thank you, noted. Over a week those taps become the only quality signal we have that a person supplied.

The door, and what it costs to keep open

The model is one seat, so there is a limit on how much of it any one reader can have, the door in this workshop's shorthand, and a reader's own share is printed on the front page rather than discovered: 20 questions an hour, 40 a day, and no more than 6 in any one minute. The whole service answers about 200 questions an hour and 1,500 a day; that pair is on this page rather than on the door, because it is a fact about the machine and not about your turn. The buttons carry a much wider door of their own, 30 in a burst and 60 a minute, because they wake no model; a reader meets it far more rarely than the free-text door. When a share is used up the page prints the wait: minutes when the wait is minutes, and a few seconds when it is the short burst limit, which is the one an ordinary reader meets first.

The limits are keyed per address, with a backstop per network block, so one household or one office cannot drain the day for everyone. A named residual: a shared window that fills between the check and the charge still costs that one request its own token.

The bench: what we measured before linking it

Before anything was linked from this site, the assistant sat an exam of 62 questions drawn over 41 of the pages: 41 of them the pages answer, and 21 they cannot, the second kind written to be impossible from the page. We pre-registered the floors, and moved them twice, in the open. We withdrew the first floor, at least 54 of 60 across the whole bank, before the second run, because only the on-page questions can be answered at all and one denominator hid that. Split floors took its place from the second run: on-page at least 35 of 39, off-page 21 of 21, zero figures that are not verbatim in the article, zero redaction hits. When two on-page questions were added after the second run, the bank going from 60 to 62, the on-page floor moved with it at the same proportion, 37 of 41, from the third run on. All three registrations are in the kit with their dates. The exam ran nine times while the service was being built, the first two on the 60-question bank and the last seven on the 62; the first three were scored by an earlier version of the harness's figure reader and are not comparable with the rest, which the kit says on its first page. The run below is the ninth, taken on v1.0.1, the version that launched. The service is v1.1.1 as this page is written, and the answering path did not move between them: the prompt, the seat, the six checks, the figure rule and the citation resolver are the same bytes. What changed is the page around them — white panels, foldable regions, a button that says what it is doing — the fence below, and the words under the answer.

armon-page cleared every checkmedian seconds
cold, one in flight41 / 411.065
cold, six in flight41 / 414.702
warm, one in flight41 / 410.835
warm, two in flight41 / 410.880

Two columns are not in the table because every cell in them is the same: figures not in the article, 0 on all four arms, and off-page abstained, 21 of 21 on all four.

Run 9, 2026-09-13 03:50Z to 03:56Z (UTC), on v1.0.1, the launched version — one gemma-class seat, 62 questions over 41 articles. Cold means the page was not in the model's prompt cache, what the first reader of a page that day meets; warm means it was warmed once immediately before its own questions, what the second reader meets. In flight is how many questions were being written at the same moment on one seat; the warm wide arm asked for six and ran two, because a page is the unit of parallelism on that arm and no page carries more questions. Two more columns are said here rather than set in the table, so the table stays three short columns wide on a phone. The 95th-percentile seconds are 7.633 cold one in flight, 14.539 cold six, 1.443 warm one and 2.121 warm two. The share of the 62 questions whose prompt was at least half served from the model's cache is 62 of 62 on both warm arms, and not measured on either cold arm, which has no warmed page to hit and so nothing to count. Cleared every check is the floor's own words: zero non-verbatim figures and at least one valid citation. It is not "the reader got what they asked for". On the cold six-in-flight arm, one of the 41 on-page questions came back as the authored abstain and still cleared, because an abstain carries a real citation; the run files print that separately as on_page_abstained. This is the only column on the page where the two differ. The seconds are printed to three decimals. The 95th percentile is the exclusive quantile. The four rows are runs/r9-cold-c1.json, r9-cold-c6.json, r9-warm-c1.json and r9-warm-c6.json in the kit.

Read the warm single-stream row in words. With the page already in cache, half the answers came back in under 0.835 seconds and nineteen in twenty in under 1.443 seconds; every one of the 62 questions was served from the cache. Cold, nineteen questions in twenty came back inside 7.633 seconds, and the slowest are the ones asked on the longest pages.

The figure and citation columns are read off the model's raw draft, before any check has touched it. They have to be: read off the served answer they would be tautologies, because the checks strike a non-verbatim figure, so a bench that counted the survivors would report zero every time by construction. What the table says is that the model, writing freely with the article in front of it, invented no figure in the 248 drafts of this run, nor in the 1,240 drafts of the five runs before it that were scored by the instrument this one uses. The three runs before those record twenty rows as non-verbatim, and the harness's own dated note says why: its figure reader was counting the digits inside citation ids, which are addresses and not prose. We fixed it on 9 September and have not re-scored those three runs. They are in the kit, with the defect named, rather than dropped.

What the bench measures is groundedness by the only tests a machine can run: every figure verbatim, every cite real, no redaction hit, and the abstain when the page does not say. It measures nothing about whether an answer is good. A judge seat, a second model that would grade the first, is empty by its own bench's verdict from earlier this month, and un-judged model text cannot grade a model, so this page makes no quality claim. The yes-or-no tap under each answer is the beginning of that measurement, with people as the judges.

The run that failed, and what it changed

An off-page question is one the page cannot answer, written on purpose, and the pass is the abstain. One run the day before launch, the seventh, failed that floor on one arm: the cold single-stream arm abstained on 20 of 21. Asked how a page's result compared to one published in 2028, the model wrote two sentences. The first said the article had nothing from 2028 and that all its dates were 2026; the second pointed at the closest section. A check struck the first, because 2028 is a name the article does not use. The second was released alone, and a released sentence is an answer, so the abstain floor failed on a question the model had, in fact, declined to answer.

The fix was to the list of reasons that allow a partial release: a sentence struck for a name the page does not carry no longer lets its neighbours ship as an answer. That bank row is a test now, and the eighth and ninth runs passed. The failure is on this page because it changed the code, and because the first record we wrote of it told the story backwards, which the ledger now says in a dated correction.

What it got wrong on the way here

This assistant went from a plan on 8 September to a first served page on 9 September and to the line on every article on 13 September. In between it was audited in fifteen numbered rounds of audit and patch and one seven-lens deep audit with adversarial refutation, by agents told to read the live pages as a stranger, to refute the previous audit, or to drive the door with a script. A few of the things they found, in the order they were found:

  • It answered in lowercase. The instruction file we gave the model was written in this site's lowercase house style, and the model matched it. An operator's own first test caught it. Every reader-visible sentence now starts with a capital, and a test sweeps every authored string, the templates and the script, and fails on a sentence that starts lowercase.
  • The draft showed the model's citation syntax. When the draft became always-visible, its raw section markers were the first thing on the page above every answer. They render as the same numbered footnotes the answer uses now.
  • The receipt claimed a check that did not run. The block under each answer said figures were checked "as a number or number-word, found verbatim in the cited section". The check read digits only, against the whole page. The check now reads spelled numbers too, at a measured cost of zero newly refused drafts across the 1,968 drafts of the first eight runs, and the line says what the check does.
  • One reader's "yes" could land on another reader's row. Two people asking the same page in the same minute, both given the same authored abstain, produced two notebook rows identical on every key the tap used to find its row. Rows are addressed by their byte offset now, and the keys are still verified.
  • The shared backstop charged the reader it refused. A reader turned away by their network block's window had already spent one of their own 20 questions on the refusal. Every shared window is checked before any personal one is charged, and a third reader behind a busy block keeps their whole budget.
  • The "linked from every page" line was false for five hours. It was on the index before any page linked here. It is behind a switch now, and we turned that switch on the hour the links went live.
  • A button press met the door's own fence on a phone. Within the hour of the links going on, an operator's phone pressed a button and got the refusal meant for another site's form: the service's own Referrer-Policy: no-referrer makes a browser send Origin: null, or nothing, on an ordinary same-origin post — the Fetch Standard's own step, which we then measured on one engine — and the fence read that as foreign. Every link came off the hub that hour. The fix accepts an absent or null origin and refuses a present foreign one; it shipped that night, and the links went back on with it.

Each one is a dated row in the workshop's ledger and a test in the service's suite, 1,335 tests as this page is written at v1.1.1, and the bullets name what each test now pins.

What this page does not say

  • No quality claim. The bench measures groundedness; the judge seat is empty; the taps are days old.
  • No load beyond six in flight, and only two on the warm arm. In flight is how many questions are being written at the same moment on one seat; the bench's widest arm is six, cold, and the warm arm ran two.
  • One seat class, not a comparison. A smaller or larger writer would move every number in the table.
  • No dedicated seat. It serves the workshop's other products too; when it is busy or resting the room prints the wait and the buttons still answer.
  • No uptime week yet. The probe that watches it started counting at launch; the first week's curve will be added below, with its window at both ends.

This is in beta, and here is what beta means here. Every released article carries one line to its room, and the line says so: in beta, still being tested. It is not a label for low expectations — the answering path is the one the bench above measured, and it is the same bytes today. It means three particular things. The door is narrow on purpose while one seat serves this and the workshop's other products; the door section above prints a reader's whole share, and the page prints the wait rather than hiding it. The marks are the mechanism in public, so the failures are on the page by design: a struck sentence with its reason beside it is the product working, not the product breaking. And the measurements are not finished — the probe that watches this started counting at launch, the first week's curve is not in yet, and the yes-or-no taps under the answers are days old. The room says as much in its own words, in the panel above every question box: In beta and still being tested. If something breaks, say so at the contact desk.

Corrections and later measurements will be added below, each dated (UTC), with a window at both ends where one applies, saying in plain words what it counts.

What to take with you

  • One page goes in, whole. The model reads the article you are on, up to about 82,800 tokens of prompt with its fence and instructions, and the last three turns of your conversation, and nothing else.
  • The buttons cost no model call. They read a file a local model helped write when the page was published, checked by the same figure rule, and committed.
  • The checks are a backstop. On the launch run they struck six sentences of 569, every one of them on a question the page could not answer, and the model abstained in its own words on 78 of 84 impossible questions.
  • We write down nothing that says who asked. Thirty days of questions, no address, no session, no cookie, and one bit if you tap it.
  • Quality is not measured here yet. The bench is groundedness; the judge seat is empty; the tap is where quality measurement starts.

How to check our work

Open any article on research.strata2signal.com, press ask about this page, and ask it something the page does say and something it does not; then open how this answer was made under the answer. The kit beside this page carries the question bank; all thirty-six run files, nine runs with four arms each; the harness that scores them and the normaliser it scores figures with; all three floor registrations with their dates; and the notes that trace every figure on this page to the file it was read from. The kit's copies were sanitised by three named rules before they shipped (a home path became a file name, a machine's name became its role, an author's name became the workshop's), and the rules are written in the kit's own provenance file with both digests of each of the thirty-nine files they touched. If a number here does not reproduce from the kit, say so at the contact desk, where a person reads every message.

Who ran this, and thanks

The corpus is this hub's own articles, most under CC BY 4.0 and the rest under the terms each page states; the room reads nothing else. The seat is a gemma-class open-weight writer from Google DeepMind. The licence file packaged with the fourth-generation weights reads Apache 2.0 and carries none of the custom terms earlier gemma generations packaged — a read of the redistributed blob, not of the vendor's own terms page, which is the distinction the hub's licence ledger draws and stamps. The weights this service loads are a quantised repack of it, and its licence is the repack's published one. The hub's licence ledger carries first-hand reads of the packaged blobs of the same family, and says in its own words where a read is a file and where it is a metadata tag. It is served by vLLM (Apache 2.0) with FP8 weights on PyTorch (BSD 3-clause) over CUDA, the one closed piece in the chain, named as such. The service is Python 3.14 standard library only, behind Caddy (Apache 2.0), managed by systemd (LGPL 2.1 or later); its notebook is a plain JSON-lines file. The rate limiter and the directive lint are our own code, vendored from another product of this workshop with a drift test. The related-page matrix beside each article was built once with a nomic-class embedding model (Apache 2.0) and committed; no model runs for it at request time. The bench harness, the packs and the twins are the hub's own tools.

A small human team ran the benches, read the audits, and signed the numbers; a fleet of AI agents built, audited and re-audited the service under that team's rulings. None of the projects above owed us anything.

elsewhere in the workshop

a strata→signal property · hello@strata2signal.com · say hello