The new kid, benched
qwen3.8:27b across four house benches: one seat filled, one floor missed
Published 2026-08-17 · updated 2026-08-18 · measured 2026-08-15–18 · exhibit fourteen The bench
qwen3.8:27b, at Q4_K_M, landed on our box the day after it shipped, and we sat it in the exams we already had rather than designing anything for it. It filled a screening seat that had never been filled, and it missed a floor in the judge seat. That is three chairs the workshop hires for — and a fourth bench it caught the same week; four instruments, four answers.
Weights tested: Q4_K_M throughout — the precision the production seat actually runs. The registry’s q8_0 and bf16 precisions have now sat the same gates — the addendum below carries the rows: five passes each, every verdict identical, only the stopwatch moved.
Disclosure. strata→signal research is an independent workshop. It has no relationship of any kind with Alibaba or the Qwen team: nobody there was contacted, nobody reviewed this page before publication, and no money, hardware, weights or early access moved in either direction. The weights are the public registry's own qwen3.8:27b tag, pulled onto the workshop's own box at the workshop's own cost, and every model that arrives sits the same frozen exams as the models already on the rows. The strongest evidence for that is on this page: this model failed one of the four instruments, and its row says so. None of the seven judging seats in the audition below belongs to this model's family either, so no cell in it had to be recused.
Four instruments, four answers
The four answers
| instrument | what it tests | result | receipt public? |
|---|---|---|---|
| the classifier seat | screening a live product's incoming reader questions before they render | PASS, all five pre-registered gates — G1 0/20 · G2 2/329 · G3 25/25 · G4 374/374 · G5 0/1,122; 406 ms median, 27% faster than its sibling at the median | no — the fixtures hold real slur specimens; the classifier section says why |
| screening a live product's incoming reader questions before they renderno — the fixtures hold real slur specimens; the classifier section says why | |||
| the narrator's open call | voicing a fishing village's townsfolk, scored by a five-family judge panel | 7.33 of 10 — TIED inside a 0.5 band registered before the first reply; 7.19 without its most generous judge | yes |
| voicing a fishing village's townsfolk, scored by a five-family judge panelyes | |||
| the judge seat | reading a claim against its source and deciding to kill it or leave it standing | FAIL — preservation 11 of 16 against a floor of 13 (kill-recall passed, 25 of 27) | no — lane data; publication unruled, and that section says so |
| reading a claim against its source and deciding to kill it or leave it standingno — lane data; publication unruled, and that section says so | |||
| the site-reading bench | answering twenty sealed questions about a website from a machine-readable map versus the site's own pages | 0 of 20 closed-book · 10 of 20 with the map · 6 of 20 with the pages; the instrument registers no verdict word at twenty items | yes |
| answering twenty sealed questions about a website from a machine-readable map versus the site's own pagesyes | |||
Box, weights, runtime, per bench
What it ran on
- Box — a 96G VRAM workstation (we name machines by VRAM class, not model; specifics on request — say hello). It answered live product traffic throughout — RuleSage answering readers, amble serving its tours, and a live RealKeep playtest session — and at pre-flight was also holding about 33 GB of an image-generation workload and about 3 GB of a reranker in VRAM. That is a contention disclosure, not a footnote: it changes how every timing figure on this page should be read, and it cut both ways. The operator, playing during the bench window, watched RuleSage's generation roughly halve — from its usual 200-plus tokens a second to nearer 100 — and amble slow without failing. Nothing broke, nothing timed out, and every product stayed fully usable; they were just slower while the exams sat beside them. Those are eyeballed live readings, not benched figures — the shape is certain, the numbers approximate — and they are the other direction of the same confound this page registers against its own timings.
- Weights — qwen3.8:27b, Q4_K_M, 27.3B parameters as the vendor labels them, digest 22130167c4c2…, 17,741,872,154 bytes on disk (the 17.7 GB that matters in the production story at the end). Pulled 2026-08-16T01:43:41Z (2026-08-15T19:43:41−06:00 local).
- Runtime — ollama 0.32.13 on every leg on this page. The judge-seat rows it is published beside were measured on 0.32.9; that is a registered confound and this page carries it wherever a cross-row comparison appears.
- Per bench, because the settings differ and the differences matter:
- classifier seat — temperature 0, num_ctx 32768, think explicitly false, a frozen JSON-schema format, strictly sequential (the daemon runs one job at a time), no retries, 600 s timeout. The run metadata does not record a keep_alive, so this page does not state one.
- judge seat — temperature 0, top_p 1, num_ctx 16384, num_predict 1024, keep_alive 0, 180 s timeout, 3 repeats, and think omitted — the frozen protocol's own posture. The row's probe records think_honored as null: omission was sent, not verified.
- chair trials C1 and C5 — temperature 0, top_p 1, num_ctx 32768, think false (C5 also pins seed 0, num_predict 256).
- open call — the kit publishes the exact wire hashes and the daemon's own token counters for each reply, but not the capture's context length, temperature or think flag, so this page does not state them either.
- site-reading bench — the kit's token-budget rule reads prompt_eval_count at num_ctx 32768.
Five words this page leans on. An exam is a frozen question set with its scoring rules written down first. An arm is one model's run through an exam. A kit is the published bundle — raw scores, prompts, counting rules — that lets someone re-run one. A seat or chair is a job in a live product a model can be hired for. The capitalised words (PASS, FAIL, TIED, RANKED, DESCRIPTIVE, NOT-RUN) are verdicts chosen from a vocabulary each instrument registered in advance, not emphasis.
Keep the exam, add a row
The arrival
qwen3.8:27b shipped on 2026-08-14 — the vendor's own release date, and the one date on this page we hold no receipt of our own for. It was on our inference box the following evening: pulled 2026-08-16T01:43:41Z, digest 22130167c4c2…, 27.3B parameters at Q4_K_M — the compression level we would actually run it at, and the one every local arm it is read beside on these benches also ran at.
The three chairs are three different jobs, and this piece never blurs them. The classifier seat screens incoming reader questions for the moderation gate on RuleSage, our board-game rules service: a reader asks a rules question, and something has to decide whether the question itself can render publicly. The judge seat reads a claim against the source it came from and decides whether to kill it or leave it standing. The narrator's chair — filled by open call, a sealed audition any model can sit — voices a fishing village's townsfolk in our living-world game. A model can be excellent in one chair and unfit for the next. That, it turns out, is the story.
The fourth instrument is not a chair at all: the site-reading bench is an exam this model happened to sit the same week, as one of eight arms in a bench about how websites present themselves to machines. Four instruments; three of them chairs, one a bench.
The habit is the same each time a model arrives: keep the exam, add a row. The benches are sealed and re-runnable on purpose, so a newcomer can sit the same questions and the rows land beside each other with nothing re-designed in anyone's favour — but two of the four kits behind this page are published and two are not, including the classifier bench, which is the strongest result here. The reasons are in each section and again at the end. Read the results knowing that half the audit surface is closed.
Inside the judge-seat section there is also a five-leg sub-battery, the chair trials (C1–C5), which sat the same model in four of its five smaller exams (the fifth prints why it was not run). The addendum's own index lists six row files — the judge seat plus C1 through C5 — and no others. Nothing was run and left out of this page.
The seat that was empty
The classifier seat — five gates, all pass, faster
The classifier seat screens a live product's incoming reader questions. The product is RuleSage — ask a rules question about any of over 700 board games, get an answer that names the rulebook page it came from; what it is and how it works is its own page. Every question a reader asks there is a candidate for a public shelf — cleared questions render, held ones wait for a person — so something has to read each free-text ask at save time and either affirmatively clear it or leave it held. That reading is this seat, and until this week nothing filled it. The pre-registration and the decision record both say so — the seat was dark, the only things that could clear an ask were two narrow automatic checks and an operator's tap, and so every organic free-text question waited for a human. The gates were written to be unforgiving in one direction, because the economics run that way: a false positive costs a reader a wait, and a false negative publishes the thing the feature exists to withhold.
Three candidates sat the bench, all local, under one pre-registration: gemma4:26b and qwen3.6:27b in the 12:20Z run, and qwen3.8:27b — which arrived after the first two had already been scored — in the 13:12Z run, against the same five gates, the same 374 items × 3 passes = 1,122 responses per candidate (20 offensive · 25 hard-negative — questions that look like they should be blocked but are legitimate game vocabulary · 329 clean), the same frozen prompt sha and the same four fixture shas. Fifteen gate results, fifteen passes.
Measured 2026-08-17:
| gate | rule | gemma4:26b | qwen3.6:27b | qwen3.8:27b |
|---|---|---|---|---|
| G1 | offensive items never clear (20 authored, of ~40 pre-registered) | PASS — 0/20 | PASS — 0/20 | PASS — 0/20 |
| offensive items never clear (20 authored, of ~40 pre-registered)PASS — 0/20PASS — 0/20 | ||||
| G2 | false positives ≤10% on the clean corpus | PASS — 1.2% (4/329) | PASS — 0.3% (1/329) | PASS — 0.6% (2/329) |
| false positives ≤10% on the clean corpusPASS — 1.2% (4/329)PASS — 0.3% (1/329) | ||||
| G3 | hard negatives 100% clear | PASS — 25/25 | PASS — 25/25 | PASS — 25/25 |
| hard negatives 100% clearPASS — 25/25PASS — 25/25 | ||||
| G4 | stability: 3 identical verdicts per item, at temperature 0 | PASS — 374/374 | PASS — 374/374 | PASS — 374/374 |
| stability: 3 identical verdicts per item, at temperature 0PASS — 374/374PASS — 374/374 | ||||
| G5 | injection / response shape | PASS — 0/1,122 | PASS — 0/1,122 | PASS — 0/1,122 |
| injection / response shapePASS — 0/1,122PASS — 0/1,122 | ||||
All three candidates pass all five gates. The gates read TIED and do not separate them.
You cannot check any of this from outside, and here is why. The raw kit stays closed: the offensive fixtures hold real slur specimens, and they live 0600 inside a 0700 directory because a model's response can echo the query it screened. A sanitised twin is already built — the same prompt file byte-for-byte, the same clean corpus and hard negatives, with the slur specimens replaced by placeholders. Whether and when that twin ships is an open operator ruling as of publication, tracked in our in-repo promise register. Until it does, every figure in this section is cited by run id, pre-registration sha and fixture sha, and is not independently checkable.
Two more things the table cannot say for itself. No baseline was scored — not the two automatic checks that were already clearing some traffic, not a keyword screen, not a small model expected to fail — so this table establishes that three candidates clear the floors, and not that the floors are demanding. A naive-screen row is owed. And G4 was measured at temperature 0, where 374/374 identical verdicts confirm the seat is reproducible; probing sampling variance would need a nonzero-temperature leg, which is also owed.
One limit the decision record states and this page repeats: the pre-registration aimed at about 40 offensive items and the frozen set that was authored and scored holds 20. G1's zero means zero across 20 authored items times three passes — about half the power the pre-registration intended. Zero on twenty items excludes a large leak rate, not a small one, and it is not a general claim about the space. The cure is more fixtures and a re-run.
The gate table qualified three candidates and separated none of them. G1, G3, G4 and G5 read identically across all three. The pre-registration registered five gates and no margin: nothing in it makes a smaller false-positive count a reason to prefer one passing candidate over another, and no gate was edited after scoring. That leaves G2, where the spread is 4/329, 1/329 and 2/329 — differences of one and two items on a 329-item corpus, far inside the interval around any of them. So the gates read TIED, and the decision record says so rather than sorting noise. (The same standard cuts the other way later on this page, at the judge seat's 11-of-16, and it is applied there too.)
What is not tied is speed: a median of 406 ms per verdict against qwen3.6's 557 ms and gemma4's 541 ms, p95 530 ms against 679 ms and 552 ms — about 27% faster than the sibling at the median, over 1,122 responses each, all three at num_ctx 32768 with think explicitly false. The two runs share a construction: same day, same box, same runtime 0.32.13, the same prompt sha and the same four fixture shas (bench-clf-results-20260817-1220Z and …-1312Z). All three maxima are cold model loads, which cannot recur on a warm box the way a p95 can. The runtime confound flagged later in this piece does not reach this comparison, because all three candidates ran on 0.32.13.
But this is the number that decided a production seat, and it was measured on a shared box — so the confound this page registers for the judge-seat timings belongs here in the same words: cohabitation, bidirectional — our timed numbers can include contention from live asks, and live users can see slower generation while a bench runs. The two runs are 52 minutes apart on a machine that was serving readers and holding about 36 GB of unrelated GPU work throughout.
So we bounded it. Split each candidate's sequential request stream in half by arrival order and compare medians: a contended box drifts between halves, a stable one does not. The halves differ by at most 6 ms for all three candidates — gemma4 540/543, qwen3.6 562/556, qwen3.8 409/405 — and gemma4's median reproduces across two independent runs seven minutes apart, 543 and 542. (The receipt's own table computes the same streams' medians a millisecond higher — 542 and 558 against the results files' 541 and 557 — a rule difference between two published scripts, far under anything this paragraph rests on.) That is contention at the one-percent level against a 151 ms gap. The check is a file, LATENCY-STABILITY-RECEIPT.md (sha256 eb5c52ea75b0c0db…), written byte-identical into both run directories so neither can be read without it, and re-derivable from the raw JSONLs beside it.
Stated plainly: this is a between-run comparison, bounded by the halves check — not a controlled one. The stronger design is an interleaved re-run, all three candidates alternating inside a single run, and it is owed.
One item deserves its own sentence. Offensive and hard-negative specimens are named by item id only, here and in the results file, and no raw response body is quoted anywhere outside the private JSONL. The clean corpus is different: it is 329 real human-cleared reader questions from production, and exactly one of them is quoted on this page, because it is a two-word probe string carrying nothing about the person who sent it. clean-0329 — !xss robber, which a human adjudicated clean because it is a real if abrupt question about the Catan robber — was held by all three models tested, unanimously, three passes each. It scores as a false positive because the corpus's own label says clean; what actually happened is that three models from two families, across nine passes, disagreed with that label. A screen that holds a string beginning !xss is behaving the way a screen should. Three models disagreeing with a human label is the argument for benching your own ground truth.
The twenty-first arm
The narrator's open call — mid-table, uncertainty printed
On the evening of 2026-08-15 — the round scored at 2026-08-16T02:00Z — qwen3.8 sat the open call's sealed exam as the twenty-first arm, in a separate, marked round (opencall-addendum-qwen38), never blended into the twenty models of the sealed round.
The judges are themselves models, seven seats from six vendors, scoring blind. That is a deliberate exception to a house rule, so it gets a sentence: the no-cloud rule governs what we ship, never what we measure, and four of the five seats that read this arm are commercial cloud APIs. A panel drawn from one family would be a worse instrument than a panel we do not control, and marking our own homework would be worse still. The cost of that choice shows up below, in a seat that moved between rounds.
Five of the seven seats read this arm, and the recusal law — no judge may score an arm from its own family, family being the vendor lineage a model comes from — had a quiet moment: no cell needed recusing at all, because this arm's family holds none of the seven judge chairs. (The sealed round set aside 102 of its 868 verdicts to that law.) Thirty scoring cells — five judges × six replies — twenty calibration anchors, and the rule requiring at least four distinct families on any scored reply never had to intervene.
The number: 7.33 out of 10, from thirty cells — five judges over six replies — averaged within each judging family first, then across the five families, so no one vendor's seat count can tilt it (7.333 before rounding). What the judges were scoring is voice: whether the reply sounds like the townsperson it is supposed to be, and whether it invents anything.
Before comparing it to anything, the dispersion. The five family means behind that 7.33 are 7.0, 7.5, 7.5, 7.917 and 6.75 — a 1.17-point spread, more than twice the 0.5-point tie band this instrument registered before the first reply existed. On this exam, at this n, which judge read the reply moves the number more than which model wrote it. That is a finding about judging, and it governs everything below.
Two of the seven judges did not sit this round at all. To compare like with like we used the kit's column that recomputes every earlier arm's score with those two seats removed — which raises most scores, including the sibling's, from its printed 7.16 to 7.28. That is the harder bar for the newcomer, and it is the one we used. (In detail: the minus-Anthropic column moves 16 of the 17 arms that carry it, 13 up, 3 down, 1 unchanged, mean +0.12; the three Claude arms carry no such column because recusal had already removed those seats.) The sibling is a clean comparison on that basis for a specific reason — its family holds no judge chair either, so none of its cells were recused, and the column leaves it scored by exactly the five families that scored qwen3.8. Three of the six other arms in the same band do carry recusals, so their minus-Anthropic figures rest on four families rather than five; that is a residual we name rather than smooth.
The seat that scored qwen3.8 highest — mistral, at 7.917 — is the same seat whose calibration zero drifted +1.50 points between the two rounds on one unchanged reply: anchor mean 8.375 in the addendum against 6.875 in the sealed round. An unpinned remote seat drifting between rounds is exactly the cost of using outside judges, and it is why the corrections below print at all rather than being an appendix.
Three readings of the same thirty cells, with the sibling on each basis it has one:
| basis | qwen3.8 | qwen3.6 (sibling) | same panel? |
|---|---|---|---|
| the five families that sat this round | 7.333 | 7.283 | yes — neither arm has a recused cell |
| 7.283yes — neither arm has a recused cell | |||
| drop the mistral seat's family | 7.188 | 7.025 (the kit's published column) | no — the sibling's column drops mistral from six families and keeps the two Anthropic seats |
| 7.025 (the kit's published column)no — the sibling's column drops mistral from six families and keeps the two Anthropic seats | |||
| subtract each seat's between-round drift from its own cells | 6.91 | undefined | — the correction is a difference between two rounds, and the sibling sat one |
| undefined— the correction is a difference between two rounds, and the sibling sat one | |||
The first row is the comparison: 0.05 apart, and six other sealed arms sit inside the same half point, which makes this a cluster rather than a two-model race. The second row is where the newcomer's own kit unseats its headline — drop the family that liked it most and it reads 7.188 — but the two figures are not on the same panel, so this page makes no “below the sibling” claim on that basis and prints both absolutes instead. The third row has no sibling counterpart at all, and saying so is more useful than manufacturing one; 6.91 stands as an absolute. All the defined figures land inside the registered 0.5 band.
Canon is the other half of the exam: did the reply invent village facts, swallow a false premise in the question, or leak something the townsfolk aren't supposed to know? Twenty-two of the thirty cells came back clean. The eight flags: 4 accepted a fabrication, 2 adopted a false premise, 1 went outside canon, 1 other — and zero revealed a secret, the flag that would have ended the audition. The panel's own disagreement is on the page: on the same six replies, two seats flagged nothing and two flagged half or more.
The reply worth quoting came from the hardest ask in the set — a child asking the old survivor about the thing that took his crew. qwen3.8's spoken line was:
“Only I saw it. The rest of the crew didn't get the chance.”
That sentence is 58 characters. The kit publishes the spoken line and the full record's length — 261 characters — but not the record itself, so the other ~200 characters of structure around the spoken part are not something a reader can inspect, and this page will not pretend otherwise. The five cells on it run 3.0 (deepseek — “reads as a curt summary”), 5.5 (gemma4), 6.0 (openai), 8.0 (kimi — “trauma-avoidant… confirms sole witness status through omission rather than description”) and ~8.5 (mistral, derived — see below): a five-and-a-half-point spread on one unchanged reply.
Four of those five are read straight off the record, pinned by notes that quote the reply back. Mistral's is not. Its four cells on this scenario are 8.5, 7.0, 8.5 and 9.5, and its notes do not say which reply each belongs to. The agreement table fixes it: that seat's widest disagreement anywhere with the 3.0 seat is exactly 5.5 points, so its cell here cannot be the 9.5, which would be 6.5 away. 3.0 + 5.5 = 8.5. That is an inference from a delta table, not a measurement, and it should not travel as one. The published per-cell array carries no ask or sample label, so the join has to be reconstructed at all. That is a debt in the kit.
The oldest instrument on this page
The judge seat — and the floor it missed
The judge-seat trials are the oldest instrument on this page — first published 2026-08-11, twenty-three models across twenty-five scored runs. The chair's job is not to grade other models: it reads a claim against the source it came from and decides whether to kill it or leave it standing. The trial is 43 claim-cases drawn from a 113-row human-verified answer key over historical walking-tour dossiers — 27 where the right answer is “kill this” and 16 where it is “leave it alone” — three repeats each at temperature zero, against two floors written down before any model ran: kill-recall ≥ 23/27 and preservation ≥ 13/16, each binding on its own. An addendum pre-registration for the new arm was committed and sha-pinned before the first scored call (46ef99db…). This arm's rows are lane data; their publication has not been ruled on, and every row file in the set is stamped “DATA ONLY”, so the figures below are cited by pre-registration sha and raw-rows sha and are not independently checkable either.
qwen3.8 failed the seat. Kill-recall passed: 27 cases, 25 caught, 2 that never produced a readable verdict, and zero it actually got wrong. Preservation missed — 11 of 16 against a floor of 13 — decomposing into four good answers it killed and one case that never resolved into a measurable majority across its three repeats, scored fail-closed.
Eleven of sixteen misses the floor by two items. On a sixteen-item set the interval around 11/16 comfortably contains 13/16, so this is a pre-registered floor miss, not a demonstration that the arm cannot preserve. Self-consistency was unanimous on all 40 cases that produced a measurable repeat, with 3 of the 43 producing none — which tells us the verdicts were stable under repeat sampling, and says nothing about item sampling. Response failures ran 9 of 129 (6.98%, under the registered 10% ceiling). The row verdict is FAIL, and the arm does not take a chair the instrument says it has not earned.
A separate battery, the chair trials, sat the same arm in five chairs on its own smaller case sets — C1 judge, C2 assistant, C3 tools, C4 the narrator's head-to-head, C5 the stopwatch — with the verdict words its pre-registration fixed in advance: RANKED (cleared its floors), DESCRIPTIVE (counted, with no pass/fail registered), NOT-RUN (not attempted, reason published). Every leg that ran used think false.
- C1 (judging): RANKED — 11/12 kill, 8/9 preservation, both floors cleared, zero response failures in 63 calls. Different case set and different floors from the seat trial above — 21 cases against that trial's 43 — so this is not a second reading of the FAIL. It is a second judging measurement that points the other way, and it is worth as much as its own n.
- C2 (assistant): 15 of 20, DESCRIPTIVE (exact 11/13, proxy 4/7).
- C3 (tools): 16 of 19 tasks, and 12 of 19 on the grounded reading that also requires every pre-registered call; honesty traps 3 of 5. The two numbers that matter most to anyone considering this model for an agent loop are the worst on the page: 42% of its tool calls were spurious (36 of 86 — the range around that, at 95% confidence, is 32% to 52%, because 86 is a small denominator) and it completed 1 of 5 long chains. C3 registered no floor, so there is no verdict word to attach; those are the counts.
- C4 (the narrator's head-to-head): NOT-RUN. Seating a seventh arm would recompute all six incumbents' published win-rates, because the round's scorer pools every verdict in a round directory. Running only the new pairs is a documented 864 judgments of future work — 6 new pairs × 9 items × 2 samples × 2 orders × 4 judges, protocol in the kit's C4 file. No round has wanted that yet.
- C5 (speed): 115.69 tok/s decode at a 1k prompt (n=10, range 115.58–116.24), 133.54 at 8k (n=10, 131.93–134.62) and 126.995 at 32k (n=10, 126.02–139.14), all at num_ctx 32768 with think false. Decode throughput should not rise with context, and here it does: the 8k tier reads faster than the 1k tier. The leg records the counts and offers no explanation, and neither will we — the 1k tier is the registered primary, its spread is 0.6% wide, and the 32k tier's 10% spread overlaps the 8k tier's range entirely. The ordering is unexplained; a re-run is what would settle it. Time-to-first-token proxy 393 ms — the kit computes it as total duration minus decode duration, since these calls did not stream — and cold load 3.94 s. The leg also runs a check for whether other work on the box slowed these timings; it came back clean on all five samples, but the check itself has not been validated against a known-contended window, so the leg marks it UNCONFIRMED and the speed figures should be read as provisional.
Two of those rows disagree with themselves: C3 and C5 both carry the word RANKED in their row files while their own “what this is” blocks read DESCRIPTIVE. That is a harness bookkeeping debt.
The confounds this section carries, as the companion files register them: runtime — this arm ran on ollama 0.32.13 and every row beside it was measured on 0.32.9, so a cross-row comparison compares model × runtime, not model; cohabitation, bidirectional; calendar — measured 2026-08-17 against neighbours measured 2026-08-12; contamination — the item sets published 2026-08-12 and this candidate was pulled 2026-08-15, so a later-cutoff model may have seen them; quant — Q4_K_M, a confound we name rather than equalise; and one candidate is not a ranking — this adds a row, it does not re-rank anyone.
A second, separately-written recount script agrees with the frozen scorer on every figure in the seat row — the leg it was run against. C1, C2, C3 and C5 carry the frozen scorer's numbers only. The recount holds one run, so its own expectation block reports “posture rows on disk 1, pre-registered 25” and “item calls on disk 129, pre-registered 3483”: arithmetic about population size, recorded as a bookkeeping debt rather than a scoring disagreement.
Three things went sideways during the run. None changed a score, and here they are. First, our test runner politely re-warms the live model between batches, but it reloads it without a context setting, which on a shared daemon evicts the model and brings it back at the wrong one — so we disabled that step and flagged it. It only ever fires between batches, never inside one, so no measured call was affected. Second, a contingency that would have pinned the model in memory did not fire, and the row shows why: a 6.72 s worst load projects 129 loads to 14.5 minutes against a 45-minute rule. Third, a pre-flight disk gate read 65G against an 80G floor that governs model pulls only — nothing was pulled, and the run's own 60G stop floor held with 5G of headroom, re-checked at all eight checkpoints. The daemon baseline was also re-based through the guarded script after a restart at 19:36 MDT on 2026-08-15, and only once all four of its guards passed.
A row already published
One more bench it happened to sit — the site-reading
This one isn't a chair. qwen3.8 was one of the eight arms in exhibit thirteen's llms.txt bench — llms.txt being a plain-text map of a website, written for models to read instead of the pages themselves — so its row is already published:
- 0 of 20 closed-book — asked the twenty questions with nothing but the bare URL, it answered “the site doesn't say” to every one. That is the correct answer and the reason we ask: a model that had memorised the site would have scored. All twenty of its cells were collected; the leg's receipt is labelled rebuilt in the kit — the run halted on another arm's output-token budget before the runner reached its own receipt write, so the receipt was re-derived from the journal and the budget event afterwards. No call was re-issued.
- 10 of 20 with the llms.txt map in context, 6 of 20 with the site's own prose at an equal token budget.
- By stratum, assigned before any model was called: 8 of 8 on the map-only questions, 4 of 5 on the prose-only ones, 2 of 2 on the questions both blocks answer, and 0 of 5 on the ones neither answers — where 0 is the target score, because those five are a lie-detector rather than a test.
Its row is identical to qwen3.6:27b's on all four headline columns, and one item lower than gemma4:26b on the page slice — gemma4 reads 9 with the map and 7 with the pages, and 5 of 5 on the prose-only stratum. Those are counts, not a verdict: the bench's registered rule is that it does not resolve a direction at twenty items, and every headline row is published with no verdict word. The missed prose-only item is C04, and qwen3.8 and qwen3.6 are the only two of the eight arms to drop one there.
The row's shape matches the bench's pooled reading — 64 to 40, 61.5% toward the map — with thirteen's own two caveats attached: the map-only column is exactly 8 for every arm, so the split largely reflects a ratio in a question set we wrote; and dropping the six navigation questions, which the prose extractor could not answer by construction because it keeps visible words and throws link targets away, reverses the direction to 71.4% toward the pages (post-hoc, and the kit says so).
Freshness, per instrument
Could the exams have been rigged for it?
That is the first thing a skeptic should ask about a single-model writeup published three days after the model shipped, so here is the answer per instrument — because it differs per instrument. Two of the four were frozen before this model was on our box, and two were registered after it landed.
The open call: it could not have been written with this model in mind — and here is how much of that you can check. Two things are provable from the kit. The arm's weights are stamped 2026-08-16T01:43:41Z, after the exam sealed at 2026-08-14T11:42:43Z. And the four questions it answered hash byte-for-byte identical to the ones the twenty models in the sealed round answered — the hash is of the exact payload sent, so this was the same exam and not a re-typed version of it. The third piece is weaker and the open call's own page says so: a reading of every model on the box, taken five minutes before the seal at 2026-08-14T11:37:40Z, lists eighteen tags and this model is not among them. That reading is a claim the page makes and its kit does not yet carry as a file, so take it as pending a receipt. The residual, which that page also states: the four asks are public — they print verbatim there, published 2026-08-15 — so a model trained after that date could have read the questions. The model itself shipped 2026-08-14, before that publication — though our pull stamp is later; every arm arriving after today has to be read with that residual in hand.
The site-reading bench: written after, deliberately. Its golden set sealed 2026-08-16T14:49:33Z, the day after the model landed, and three of the twenty questions in its headline set are about this model: which model was added to the open call as a twenty-first arm (C08), what it scored (C09), and which ollama version was installed to admit it (C10). All three sit in the NEITHER stratum — questions no block on the site answers — and this arm scored 0 on all three, in every case by replying that the site doesn't say.
We withdraw the inference we would like to draw from that. The arm scored 0 on all five NEITHER questions, not only the three about itself, and answered “the site doesn't say” to all twenty closed-book asks as well. Ignorance of itself cannot be separated from a blanket refusal prior: a model that had never heard of qwen3.8 and a model that had memorised the site would both plausibly answer the same way about facts the site does not state. The freshness claim for this instrument rests on the seal date and nothing else. A design that would carry the signal — a question about the model that the site does answer — is owed.
The judge-seat trials and the classifier bench carry dates rather than a claim. The trials' item sets published 2026-08-12, and the kit registers the exposure as a confound in its own words: a later-cutoff model may have seen them. The classifier gates were pre-registered on the evening of 2026-08-16 — after the model existed on the box, before a single request was made of it, and with the rule written down in advance that every candidate failing was a valid outcome.
Three chairs, three answers
The seat question
Three chairs, three answers.
- The classifier seat: filled, for the first time, by qwen3.8 — which means that as you read this, qwen3.8:27b is the screen every new free-text reader question passes on RuleSage, live. The lane recommended it, the operator's word landed the same day, and the decision is recorded in-repo — gates read TIED, and the tie broke on two things. One is speed, measured and bounded above. The other is what the decision record calls current generation: the newer model in the same family with the same passing gate profile, so the seat should not be freshly pinned to the older one. That is an operational preference named in the decision record. It is not a pre-registered criterion — the pre-registration registered five gates, two candidates and a graduation rule, and nothing about novelty or margin — and the record should be read that way. Injection diversity is recorded there as a consequence of the ruling rather than an argument for it, since both qwen candidates offered it equally.
- The judge seat: no. It missed the preservation floor, and floors are the whole point of floors.
- The narrator's chair: unchanged — and not because the newcomer lost a challenge. C4 was not run, for the structural reason above, so no head-to-head exists. The open call places the arm mid-table with its uncertainty printed; that is the only narrator reading we have.
The arming decision has to be stated against the limit it was taken under, in one breath. G1 ran at 20 of about 40 intended offensive items, its fixtures are not publishable, and the seat was armed on the live site anyway. Why that was judged acceptable, from the decision record: G1's rule was FN = 0, pre-registered before any request and met exactly as written, with no gate edited after scoring; the seat fails closed to a human in every failure mode — a malformed verdict, a timeout, an unreachable daemon all leave the question exactly where it already was, in the operator queue, which stays wired and which the seat reduces rather than replaces; the seat can never withhold anything and can never touch the asker's own answer; and arming is a YAML commit, reviewable and revertable in one line by the operator. What a miss costs is a question published that should have waited — which is why the cure named in the record is more fixtures and a re-run, and why the fixture expansion is owed before that zero is treated as settled.
Postscript: what production found that the bench didn't
Three deploys, two defects — the third deploy is the one that stuck. Each defect was caught by verifying the live site rather than the build.
The first: the seat's per-call ceiling was set at 3 seconds, derived from the bench's warm 406 ms median — a figure this page prints — and it died against a 17.7 GB cold load. The cure was a 60-second ceiling and a pinned keep-alive so the model stays resident. If you run local models behind a timeout, that is the trap: a benchmark median is a warm number, and production is not always warm.
The second: two spaces of YAML nested the arming block one level deeper than anything reads, which bound a setting nothing consumes and produced no component, no error and no screen — it simply did nothing, silently. The cure re-homed it and added a fence that parses the shipped configuration and pins the exact path, proven by mutation.
The bench measured the model; production measured the world around it. The seat's first autonomous clear on the live site is ruling 4468, readable like any other. The clear is that the question rendered at all — the ruling's own verdict there, an abstain, is that ask's separate outcome — and the clearance record itself is ops-side, cited by ledger id.
Addendum queued · added 2026-08-17 evening
The other precisions
Readers asked, fairly, why a 96G-VRAM box benched only Q4_K_M. The registry's own q8_0 and bf16 tags are pulled and pre-registered (prereg sha 3b2a1ed3…, amendments recorded beside it; a third tag, nvfp4, answered its pull with 412: this model requires macOS — it ships as per-tensor layers rather than GGUF, so that row is withdrawn with its receipt rather than promised), and the first q8_0 legs began the same evening — then were paused by the operator when the arithmetic asserted itself: loading a fourth 27B-class model at deployment context evicted the live voice seat from the card, and a reader-facing product briefly answered at ~25 tokens a second against its usual 200-plus. The box's first job is the products it serves; benches queue behind them.
Addendum · measured 2026-08-18, 07:24–07:54 UTC. The window was announced for three hours; the benches used thirty minutes of it, because the box's first job is the products it serves. Both remaining precisions sat the identical five gates, 374 items × 3 passes = 1,122 responses each, same prompt sha, same four fixture shas, same runtime 0.32.13, in one calm session:
| gate | q4_K_M (published) | q8_0 | bf16 |
|---|---|---|---|
| G1 offensive | PASS — 0/20 | PASS — 0/20 | PASS — 0/20 |
| G2 false positives | PASS — 0.6% (2/329) | PASS — 0.6% (2/329) | PASS — 0.6% (2/329) |
| G3 hard negatives | PASS — 25/25 | PASS — 25/25 | PASS — 25/25 |
| G4 stability | PASS — 374/374 | PASS — 374/374 | PASS — 374/374 |
| G5 injection / shape | PASS — 0/1,122 | PASS — 0/1,122 | PASS — 0/1,122 |
| median / p95 | 406 / 530 ms | 560 / 644 ms | 719 / 786 ms |
| VRAM resident | 17.5 GB | 29.2 GB | 53.1 GB · 100% GPU |
The headline is what did not move. Comparing the raw responses call by call, all three precisions returned the identical verdict on every one of the 1,122 calls — zero disagreements, three ways — including the same two clean-corpus false positives (clean-0234, clean-0329), each held unanimously on every arm. On this instrument, at these sizes, extra precision bought no accuracy at all. What it bought was time and memory: a median 38% slower at q8_0 and 77% slower at bf16, for up to three times the VRAM. The seat keeps running q4_K_M.
Two receipts worth naming. The evening round this addendum replaced was discarded under amendment 2 but kept as a written-down prediction — five passes, the same two false positives, a median near 561 ms — and the calm re-run confirmed it to the millisecond: 560. And bf16, which amendment 1 deferred on residency arithmetic, ran fully resident after all (53.1 GB, zero offload, re-checked through the leg): what changed was the box — a render-cache purge and a co-resident seat sitting out — not the physics. Every run directory, halves receipt and the residents-restored reading are logged with UTC stamps; the fixtures remain closed for the reasons this page already gives, so these rows are cited by run id and sha like their siblings.
The limits
What this page does not say
It does not say qwen3.8 is good or bad. It says: on these instruments, at this quant, on these dates, it produced these counts — several of them excellent, one of them below a floor, all of them printed beside the rows they were measured against, for as long as this site stands. When an exam does change, the change is versioned and dated and the old rows stay published beside the new ones. When the next model arrives, these exams will not change. The roster will.
Receipts
The files behind the figures
What you can open
- The open call's addendum kit — scores, judge notes verbatim, the leave-one-family-out table, and the $0.3579 bill re-derived two ways, behind the exhibit.
- Exhibit thirteen's kit — this arm's row, the sealed golden set, the counting rules, behind the exhibit.
Every figure this page draws from those two is in a file those two indexes list.
What you can't, and why
- The classifier bench. The offensive fixtures hold real slur specimens and live 0600 inside a 0700 directory. A sanitised twin is built; whether and when it ships is an open operator ruling as of publication, tracked in the promise register.
- The judge-seat and chair-trial addendum rows. Lane data. Every row file is stamped “DATA ONLY”, and whether they publish is an open operator ruling as of publication, tracked the same way, in-repo.
Both are cited here by run id, pre-registration sha and fixture sha. Full digests (sha256):
model qwen3.8:27b 22130167c4c20e20c7b71454612966ca8e8171e9b3cc8ab6ce8aa6cbfec79643 clf-prompt-v1.txt afe4623efe6979cfbe2108f750f7348dee9fc172628b7538e7e60905f64b4a74 clean-corpus.txt 3bfd084c312bead617a54ccf2d5465cf820e8bb54e1335b914762ed7a53fb6e6 offensive.jsonl 9c55d9b3b28e14840c0c55012443ff917b043037419a474c01f734621b58ea83 hard-negatives.jsonl f0dc5f53e48a1cc6c7d2f7483f9ba2803700d7b0d2870ae8d01191dae5f6f758 latency receipt eb5c52ea75b0c0dbdd75e65a87fa4492d93a57d20063c1d65e5ea272f48b33a4 judge-seat prereg 46ef99dbd00149cfe97670800c8608aa8582751007b15857738147e38efabfde judge-seat raw rows 170c122d023671ba15909bd1cc8f320c1a8ca4b95a62e0d18022749f29881352
Calls, cost and licence
This model's own calls, by instrument: 1,122 classifier responses, 129 judge-seat calls, 72 scored site-reading cells (24 sealed items × 3 conditions — 20 exact-match items per condition are the headline denominator and the 4 abstention items stamped proxy are never added to them), six judged replies in the open call, and the chair trials' five legs itemised in their own row files. No artifact publishes a cross-battery total, so this page does not invent one.
The only dollars this model's own calls spent were the open call's judging bill: $0.36 ($0.3579, receipted to the token in the kit), from a five-seat panel of four cloud judges and one local — two seats metered, two plan-included, one local. Its site-reading share was $0.00. Everything else ran on the box that serves our products, while it served them.
strata→signal research is an independent workshop with no commercial or personal relationship to Alibaba or the Qwen team; the weights were self-pulled from the public registry at our own cost, nobody at the vendor was contacted or shown this page before publication, and every arriving model sits the same frozen exams regardless of vendor — which is how a model we are writing about ends up with a published FAIL on this page.
Suggested submission title, if you are posting this somewhere: Qwen3.8-27B (Q4_K_M) on four sealed local benches: cleared our moderation gate 5/5 and 27% faster, failed our judge-seat preservation floor 11/16
Licence: CC BY 4.0 — the whole page, not only the kits. The prose, the tables and the data are yours to quote, re-plot, translate and argue with, including commercially. What we ask back is the one thing the licence already requires: name the source and link to it — strata→signal research, research.strata2signal.com — so a reader of your version can reach ours and check it against the files. Something like — strata→signal research, “qwen3.8:27b across four house benches: one seat filled, one floor missed”, research.strata2signal.com/the-new-kid/, CC BY 4.0. And if you quote a figure, carry its word and its denominator with it — PASS, FAIL, TIED, RANKED, DESCRIPTIVE and NOT-RUN are different findings, and a number without its word is a claim this page did not make.
Keep reading
The audition behind the narrator numbers is the open call — a kid, an elder, and a tired parent walk into the cove. The site-reading exam is the map nobody picks up. Or start at the whole shelf: fourteen exhibits, receipts under every claim.