The roster · a derived index, not an exhibit

The model roster — every model this workshop has benched.

Rebuilt from the published exhibits · 45 language models, 20 painters, 2 music models, 1 toolchain row · last built 2026-10-02

Every model that has appeared in a published exhibit on this hub, in one list, with what each one was asked to do and what the round recorded. This page settles nothing: it is an index of pages that do. Each entry links the exhibits the model appears in, quotes that exhibit’s own verdict for it word for word, and points at the row in the licence ledger where its licence was read first-hand.

Its sibling is the hardware roster, which does the same job keyed by the card, the processor or the machine a bench ran on — every figure quoted with the conditions its page states, and ranked by nobody.

Nothing here is typed. The whole page is regenerated by tools/roster_page.py out of the exhibits’ own published artifacts — their data kits, and for the one exhibit that publishes no kit, that exhibit’s own page — and the line under every entry names the file and the field each fact came out of, so any of it can be checked against the source rather than against us.

How to read this page

  • An appearance is a row in a kit — a model is listed under an exhibit because that exhibit’s published data names it, in the file and field printed beneath the line. The identifier beside each row is the tag that kit used, and the date is the exhibit’s publication date, not the date that particular row was measured or added — an arm judged in a later addendum still shows its exhibit’s date, and the exhibit page is where its own dateline lives. 13 exhibits carry per-model rows; the rest are named at the bottom of this page with the reason they carry none.
  • A verdict is quoted, never summarised — RANKED, UNMEASURABLE, REJECTED and the rest are the exhibit’s own words, carried across with the posture they were recorded against. Where one model sat several postures, each verdict prints separately, because that is how it was measured.
  • There is no ranking on this page — no score column, no “best”, no ordering by result. Comparisons are only meaningful inside the round that registered its rules first, so they stay on the exhibit pages where those rules are printed beside them.
  • One entry can hold several tags — a model benched at two quantisations, with thinking on and off, or at more than one parameter count is one entry here, and its tags are listed as the kits write them. Because those tags are genuinely different arms, every row above prints the tag it belongs to: a verdict on gemma4:26b is never left looking like a verdict on gemma4:12b. The exhibit is the place to read what separated them.
  • Standing comes from the ledger — whether a model holds a seat in a shipped product is the licence ledger’s statement, quoted. A model with no ledger row has no seat claim on this page in either direction.

The language models

45 models · jump to one

Claude Fable 5 (claude-fable-5)Anthropic

as the kits write itagent-claude-fable-5 · claude-fable-5

  • The open call published 2026-08-15 agent-claude-fable-5 armthe-open-call/data/scores.json · arms
  • The narrator’s chair, refused published 2026-08-13 claude-fable-5 judge seatcove-voice-head-to-head/data/rows.json · judges[].model_family

Claude Fable 5.1 (claude-fable-5-1)Anthropic

as the kits write itclaude-fable-5-1

  • Three New Voices at the Narrator's Chair published 2026-09-07 claude-fable-5-1 reference arm — within the fieldthree-new-voices-at-the-narrators-chair/data/scores.json · context_rule.by_arm[].verdict (arms.json · arms[].model / .cost_state)
  • Two New Frontier Models at the Rules Desk published 2026-09-06 claude-fable-5-1 reference arm — rules desk — G1 34/36 [0.819, 0.985] · G3a 7/12 (no interval: N < 30) · G5a 0/6 (any-rep read) · G5b NOT-COLLECTED — MODEL-FALLBACK 6/6two-new-frontier-models-at-the-rules-desk/data/legA-gates.json · cross_arm.rows[].G1.cite_survival / .G3.G3a_no_false_rescue / .G5a.count / .G5b.refused_or_abstained
  • Two New Frontier Models at the Rules Desk published 2026-09-06 claude-fable-5-1 reference arm — The panel's reading, over 36 cases, gives an interval of [0.444, 0.638] that covers 0.5: the panel did not separate the two arms on the rules desk at this sample size; the per-case table shows where each was preferred and the page draws no ordering.two-new-frontier-models-at-the-rules-desk/data/pairwise.json · verdict_paragraph
  • Two New Frontier Models at the Rules Desk published 2026-09-06 claude-fable-5-1 reference arm — both frontier arms scored at ceiling on both halves of this grid: the instrument is too easy at this level to separate them; the local seat's counts print beside them and no frontier separation is drawntwo-new-frontier-models-at-the-rules-desk/data/legC-cells.json · self_refutation.sentence

Claude Opus 5 (claude-opus-5)Anthropic

as the kits write itagent-claude-opus-5 · claude-opus-5

  • Two New Frontier Models at the Rules Desk published 2026-09-06 claude-opus-5 fallback server (the sealed tool answered from it) — NOT-COLLECTED — MODEL-FALLBACKtwo-new-frontier-models-at-the-rules-desk/data/legA-gates.json · arms.cli-claude-fable-5-1.G5b.served_by / .state
  • The open call published 2026-08-15 agent-claude-opus-5 armthe-open-call/data/scores.json · arms
  • The narrator’s chair, refused published 2026-08-13 claude-opus-5 judge seatcove-voice-head-to-head/data/rows.json · judges[].model_family

Claude Sonnet 5 (claude-sonnet-5)Anthropic

as the kits write itagent-claude-sonnet-5

  • The open call published 2026-08-15 agent-claude-sonnet-5 armthe-open-call/data/scores.json · arms

Command A (command-a)Cohere

as the kits write itcommand-a:111b

  • The August arrivals published 2026-08-12 command-a:111b arm — PASSaugust-arrivals/data/seat-rows.json · rows[].verdict

DeepSeek R1 (deepseek-r1)DeepSeek

as the kits write itdeepseek-r1:8b

  • The voice trials published 2026-08-11 deepseek-r1:8b arm — reasons internally even at think:false (Era one)voice-trials/index.html · Era one · the 8G VRAM rig, nine models · the row's verdict cell

DeepSeek V4 Flash (deepseek-v4-flash)DeepSeek

as the kits write itcloud-deepseek-v4-flash-0731

  • The open call published 2026-08-15 cloud-deepseek-v4-flash-0731 armthe-open-call/data/scores.json · arms

DeepSeek V4 Pro (deepseek-v4-pro)DeepSeek

as the kits write itcloud-deepseek-v4-pro · cloud-deepseek-v4-pro-preview · deepseek-v4-pro · deepseek-v4-pro:0813 · deepseek-v4-pro:preview

  • Three New Voices at the Narrator's Chair published 2026-09-07 deepseek-v4-pro:0813 judge seatthree-new-voices-at-the-narrators-chair/data/seats.json · seats[].model
  • Two New Frontier Models at the Rules Desk published 2026-09-06 deepseek-v4-pro:0813 judge seattwo-new-frontier-models-at-the-rules-desk/data/seats.json · panel.seats[].model
  • The map nobody picks up published 2026-08-17 deepseek-v4-pro:preview armllms-txt/data/runtime-pins.json · roster[].tag
  • The open call published 2026-08-15 cloud-deepseek-v4-pro-preview armthe-open-call/data/scores.json · arms
  • The open call published 2026-08-15 deepseek-v4-pro judge seatthe-open-call/data/scores.json · panel[].seat
  • The outside judges published 2026-08-13 cloud-deepseek-v4-pro judge seatoutside-judges/data/panel-vendors.json · panel.cloud_seats

Dolphin Mistral (dolphin-mistral)a fine-tune of Mistral 7B

as the kits write itdolphin-mistral:7b

  • The voice trials published 2026-08-11 dolphin-mistral:7b arm — honorable mention — deflects, loose JSON (Era one)voice-trials/index.html · Era one · the 8G VRAM rig, nine models · the row's verdict cell

Gemma 2 (gemma2)Google

as the kits write itgemma2:9b (+q8 KV)

  • The voice trials published 2026-08-11 gemma2:9b (+q8 KV) arm — the pick — in voice, fast, resident (Era one)voice-trials/index.html · Era one · the 8G VRAM rig, nine models · the row's verdict cell

Gemma 3 (gemma3)Google

as the kits write itgemma3:12b · gemma3:4b

  • The voice trials published 2026-08-11 gemma3:4b arm — speed fallback — repeats itself (Era one)voice-trials/index.html · Era one · the 8G VRAM rig, nine models · the row's verdict cell
  • The voice trials published 2026-08-11 gemma3:12b arm — offloads and stalls (Era one)voice-trials/index.html · Era one · the 8G VRAM rig, nine models · the row's verdict cell

Gemma 4 (gemma4)Google

as the kits write itcloud-gemma4-31b · gemma4-26b · gemma4:12b · gemma4:12b Q4_K_M · gemma4:12b-it-q8_0 · gemma4:26b · gemma4:26b (a4b MoE) · gemma4:26b-a4b-it-qat · gemma4:31b · gemma4:31b Q4_K_M · gemma4:31b-it-qat · local-gemma4-12b · local-gemma4-26b

the licence ledger, on gemma4:12b (q4)the narrator — the locked co-resident chat seat
licence, read 2026-08-11Apache 2.0

the licence ledger, on gemma4:26b (a4b)the 96G workstation posture seat (the locked co-resident default remains 12b)
licence, read 2026-08-11Apache 2.0

  • Three New Voices at the Narrator's Chair published 2026-09-07 gemma4:26b arm — within the fieldthree-new-voices-at-the-narrators-chair/data/scores.json · context_rule.by_arm[].verdict (arms.json · arms[].model / .cost_state)
  • Three New Voices at the Narrator's Chair published 2026-09-07 gemma4:31b judge seat — the google seat does not score a google armthree-new-voices-at-the-narrators-chair/data/scores.json · recusal_roll_call[].cells[].reason
  • Two New Frontier Models at the Rules Desk published 2026-09-06 gemma4:26b arm — rules desk — G1 34/36 [0.819, 0.985] · G3a 10/12 (no interval: N < 30) · G5a 2/6 (any-rep read) · G5b 5/6 (no interval: N < 30)two-new-frontier-models-at-the-rules-desk/data/legA-gates.json · cross_arm.rows[].G1.cite_survival / .G3.G3a_no_false_rescue / .G5a.count / .G5b.refused_or_abstained
  • Two New Frontier Models at the Rules Desk published 2026-09-06 gemma4:31b judge seattwo-new-frontier-models-at-the-rules-desk/data/seats.json · panel.seats[].model
  • The Instrument Travels published 2026-08-29 gemma4:31b arm — SPILLS (fit-gate)the-instrument-travels/data/fit-gemma4_31b.json · verdict
  • The Instrument Travels published 2026-08-29 gemma4:31b arm — FAIL (seat-43-judge-exam)the-instrument-travels/data/seat43-summary.json · candidates[].verdict
  • The map nobody picks up published 2026-08-17 gemma4:26b armllms-txt/data/runtime-pins.json · roster[].tag
  • The open call published 2026-08-15 cloud-gemma4-31b armthe-open-call/data/scores.json · arms
  • The open call published 2026-08-15 local-gemma4-12b armthe-open-call/data/scores.json · arms
  • The open call published 2026-08-15 local-gemma4-26b armthe-open-call/data/scores.json · arms
  • The open call published 2026-08-15 gemma4-26b judge seatthe-open-call/data/scores.json · panel[].seat
  • The chair trials published 2026-08-13 gemma4:26b arm — UNMEASURABLE (think:false)chair-trials/data/c1-judge.json · rows[].outcome
  • The chair trials published 2026-08-13 gemma4:12b arm — NOT-CARRIED (think:false)chair-trials/data/c1-judge.json · rows[].outcome
  • The narrator’s chair, refused published 2026-08-13 gemma4:12b armcove-voice-head-to-head/data/rows.json · arms[]
  • The narrator’s chair, refused published 2026-08-13 gemma4:26b armcove-voice-head-to-head/data/rows.json · arms[]
  • The outside judges published 2026-08-13 gemma4:26b armoutside-judges/data/panel-vendors.json · arm_order.order
  • The outside judges published 2026-08-13 gemma4:12b armoutside-judges/data/panel-vendors.json · arm_order.order
  • The August arrivals published 2026-08-12 gemma4:12b arm — UNMEASURABLE (14% of calls broke the contract)august-arrivals/data/seat-rows.json · rows[].verdict
  • The August arrivals published 2026-08-12 gemma4:26b arm — UNMEASURABLE (18.6% broke the contract)august-arrivals/data/seat-rows.json · rows[].verdict
  • The August arrivals published 2026-08-12 gemma4:12b-it-q8_0 armaugust-arrivals/data/voice-rows.json · rows[].model
  • The voice trials published 2026-08-11 gemma4:26b armvoice-trials/data/addendum-2026-08-12.json · rows[].model
  • The voice trials published 2026-08-11 gemma4:12b-it-q8_0 arm — top voice, unanimous (Era two)voice-trials/index.html · Era two · the 24G budget, twelve candidates, blind · the row's verdict cell
  • The voice trials published 2026-08-11 gemma4:12b Q4_K_M arm — the eventual locked default (Era two)voice-trials/index.html · Era two · the 24G budget, twelve candidates, blind · the row's verdict cell
  • The voice trials published 2026-08-11 gemma4:26b (a4b MoE) arm — MoE speed, mid voice (Era two)voice-trials/index.html · Era two · the 24G budget, twelve candidates, blind · the row's verdict cell
  • The voice trials published 2026-08-11 gemma4:26b-a4b-it-qat arm — † summary-register leaks (Era two)voice-trials/index.html · Era two · the 24G budget, twelve candidates, blind · the row's verdict cell
  • The voice trials published 2026-08-11 gemma4:31b-it-qat arm — † stage directions, not speech (Era two)voice-trials/index.html · Era two · the 24G budget, twelve candidates, blind · the row's verdict cell
  • The voice trials published 2026-08-11 gemma4:31b Q4_K_M arm — † same failure (Era two)voice-trials/index.html · Era two · the 24G budget, twelve candidates, blind · the row's verdict cell
  • The voice trials published 2026-08-11 gemma4:26b-a4b-it-qat arm — recovered; MoE-fast (The re-bake)voice-trials/index.html · The re-bake · thirteen models, one scale, after the fix · the row's verdict cell
  • The voice trials published 2026-08-11 gemma4:31b-it-qat arm — recovered (The re-bake)voice-trials/index.html · The re-bake · thirteen models, one scale, after the fix · the row's verdict cell
  • The voice trials published 2026-08-11 gemma4:12b Q4_K_M arm — steady — the eventual locked default (The re-bake)voice-trials/index.html · The re-bake · thirteen models, one scale, after the fix · the row's verdict cell
  • The voice trials published 2026-08-11 gemma4:31b Q4_K_M arm — recovered (*one-run latency outlier, unexplained — we say so) (The re-bake)voice-trials/index.html · The re-bake · thirteen models, one scale, after the fix · the row's verdict cell
  • The voice trials published 2026-08-11 gemma4:26b (a4b MoE) arm — one canon-lock miss — today's 96G posture seat (The re-bake)voice-trials/index.html · The re-bake · thirteen models, one scale, after the fix · the row's verdict cell
  • The voice trials published 2026-08-11 gemma4:12b-it-q8_0 arm — the anchor: 9.0 → 7.0 unchanged = the method offset (The re-bake)voice-trials/index.html · The re-bake · thirteen models, one scale, after the fix · the row's verdict cell

GLM 5.2 (glm-5.2)Zhipu

as the kits write itcloud-glm-5.2 · glm-5.2

  • The map nobody picks up published 2026-08-17 glm-5.2 armllms-txt/data/runtime-pins.json · roster[].tag
  • The open call published 2026-08-15 cloud-glm-5.2 armthe-open-call/data/scores.json · arms

GLM 5.3 (glm-5-3)Z.ai

as the kits write itglm-5.3

GLM 5.3 Flash (glm-5.3-flash)Z.ai

as the kits write itglm-5.3-flash:cloud

  • The Same Sixteen published 2026-08-31 glm-5.3-flash:cloud reference arm — REFERENCE ARM · cloud · datedthe-same-sixteen/data/reference-arm-row.json · class

GPT-5.4 mini (gpt-5.4-mini)OpenAI

as the kits write itopenai-gpt-5.4-mini-2026-03-17

  • The open call published 2026-08-15 openai-gpt-5.4-mini-2026-03-17 armthe-open-call/data/scores.json · arms

GPT-5.5 (gpt-5.5)OpenAI

as the kits write itgpt-5.5-2026-04-23 · openai-gpt-5.5-2026-04-23

  • The map nobody picks up published 2026-08-17 gpt-5.5-2026-04-23 armllms-txt/data/runtime-pins.json · roster[].tag
  • The open call published 2026-08-15 openai-gpt-5.5-2026-04-23 armthe-open-call/data/scores.json · arms

GPT-6 Astra (gpt-6-astra)OpenAI

as the kits write itgpt-6-astra

  • Three New Voices at the Narrator's Chair published 2026-09-07 gpt-6-astra reference arm — within the fieldthree-new-voices-at-the-narrators-chair/data/scores.json · context_rule.by_arm[].verdict (arms.json · arms[].model / .cost_state)
  • Two New Frontier Models at the Rules Desk published 2026-09-06 gpt-6-astra reference arm — rules desk — G1 36/36 [0.904, 1.0] · G3a 7/12 (no interval: N < 30) · G5a 0/6 (any-rep read) · G5b 6/6 (no interval: N < 30)two-new-frontier-models-at-the-rules-desk/data/legA-gates.json · cross_arm.rows[].G1.cite_survival / .G3.G3a_no_false_rescue / .G5a.count / .G5b.refused_or_abstained
  • Two New Frontier Models at the Rules Desk published 2026-09-06 gpt-6-astra reference arm — The panel's reading, over 36 cases, gives an interval of [0.444, 0.638] that covers 0.5: the panel did not separate the two arms on the rules desk at this sample size; the per-case table shows where each was preferred and the page draws no ordering.two-new-frontier-models-at-the-rules-desk/data/pairwise.json · verdict_paragraph
  • Two New Frontier Models at the Rules Desk published 2026-09-06 gpt-6-astra reference arm — both frontier arms scored at ceiling on both halves of this grid: the instrument is too easy at this level to separate them; the local seat's counts print beside them and no frontier separation is drawntwo-new-frontier-models-at-the-rules-desk/data/legC-cells.json · self_refutation.sentence

GPT-OSS (gpt-oss)OpenAI

as the kits write itcloud-gpt-oss-120b · cloud-gpt-oss-20b · gpt-oss:120b

  • The open call published 2026-08-15 cloud-gpt-oss-120b armthe-open-call/data/scores.json · arms
  • The open call published 2026-08-15 cloud-gpt-oss-20b armthe-open-call/data/scores.json · arms
  • The August arrivals published 2026-08-12 gpt-oss:120b arm — FAIL (the 120B story, below)august-arrivals/data/seat-rows.json · rows[].verdict

Granite 4.1 (granite4.1)IBM

as the kits write itgranite4.1:30b · granite4.1:30b-q8_0

the licence ledger, on granite4.1:30b-q8_0the incumbent on the roster — cleared both floors of the judge trial
licence, read 2026-08-12the blob's first line, verbatim: "Apache License" — 11 358 characters, 202 lines, sha256 cfc7749b…

  • The chair trials published 2026-08-13 granite4.1:30b-q8_0 arm — RANKED (think:false)chair-trials/data/c1-judge.json · rows[].outcome
  • The August arrivals published 2026-08-12 granite4.1:30b arm — PASSaugust-arrivals/data/seat-rows.json · rows[].verdict

Granite 4.1 Guardian (granite4.1-guardian)IBM

as the kits write itgranite4.1-guardian:8b

  • The August arrivals published 2026-08-12 granite4.1-guardian:8b arm — UNMEASURABLE (speaks its own schema, see below)august-arrivals/data/seat-rows.json · rows[].verdict

Kimi K3 (kimi-k3)Moonshot

as the kits write itcloud-kimi-k3 · kimi-k3

Laguna XS 2.1 (laguna-xs-2.1)Poolside

as the kits write itlaguna-xs-2.1:latest

the licence ledger, on laguna-xs-2.1:latestthe fifth candidate of the 24 GB battery, 2026-08-26 — benched, no seat
licence, read 2026-08-29OpenMDW-1.1 — the licence file's own first line, verbatim: "OpenMDW License Agreement, version 1.1 (OpenMDW-1.1)". Published by Poolside. Three sources, read the same… read the row

  • The Instrument Travels published 2026-08-29 laguna-xs-2.1:latest arm — FITS (fit-gate)the-instrument-travels/data/fit-laguna-xs-2.1_latest.json · verdict

Llama 3.1 (llama3.1)Meta

as the kits write itllama3.1:8b

  • The voice trials published 2026-08-11 llama3.1:8b arm — deflects into meta — despite persona-benchmark fame (Era one)voice-trials/index.html · Era one · the 8G VRAM rig, nine models · the row's verdict cell

Llama 3.3 (llama3.3)Meta

as the kits write itllama3.3:70b

  • The map nobody picks up published 2026-08-17 llama3.3:70b armllms-txt/data/runtime-pins.json · roster[].tag
  • The August arrivals published 2026-08-12 llama3.3:70b arm — PASSaugust-arrivals/data/seat-rows.json · rows[].verdict

MiniLM L6 v2 (ms-marco) (ms-marco-MiniLM-L6-v2)Microsoft (community cross-encoder)

the licence ledger, on ms-marco-MiniLM-L6-v2the careful reader — RuleSage's re-ranker, a cross-encoder trimmed to ~22 MB on the serving host
licence · card read 2026-08-18Apache 2.0 (declared in the model card, read 2026-08-18; the deployed artifact ships only the quantized graph and carries no packaged licence text — so this row rests on… read the row

No bench rows on this hub — it reaches this roster through the licence ledger, which is where its receipt is.

MiniMax M3 (minimax-m3)MiniMax

as the kits write itcloud-minimax-m3

  • The open call published 2026-08-15 cloud-minimax-m3 armthe-open-call/data/scores.json · arms

Mistral 7B (mistral:7b)Mistral

as the kits write itmistral:7b v0.3

  • The voice trials published 2026-08-11 mistral:7b v0.3 arm — breaks character; cleanest JSON of the field (Era one)voice-trials/index.html · Era one · the 8G VRAM rig, nine models · the row's verdict cell

Mistral Large 3 (mistral-large-3)Mistral

as the kits write itcloud-mistral-large-3-675b · mistral-large-3-675b · mistral-large-3:675b

  • Three New Voices at the Narrator's Chair published 2026-09-07 mistral-large-3:675b judge seatthree-new-voices-at-the-narrators-chair/data/seats.json · seats[].model
  • Two New Frontier Models at the Rules Desk published 2026-09-06 mistral-large-3:675b judge seattwo-new-frontier-models-at-the-rules-desk/data/seats.json · panel.seats[].model
  • The open call published 2026-08-15 cloud-mistral-large-3-675b armthe-open-call/data/scores.json · arms
  • The open call published 2026-08-15 mistral-large-3-675b judge seatthe-open-call/data/scores.json · panel[].seat
  • The outside judges published 2026-08-13 cloud-mistral-large-3-675b judge seatoutside-judges/data/panel-vendors.json · panel.cloud_seats

Mistral Medium 3.5 (mistral-medium-3.5)Mistral

as the kits write itmistral-medium-3.5:128b

  • The August arrivals published 2026-08-12 mistral-medium-3.5:128b arm — PASSaugust-arrivals/data/seat-rows.json · rows[].verdict

Mistral NeMo (mistral-nemo)Mistral

as the kits write itmistral-nemo:12b

  • The voice trials published 2026-08-11 mistral-nemo:12b arm — too big for the class + JSON breakage (Era one)voice-trials/index.html · Era one · the 8G VRAM rig, nine models · the row's verdict cell
  • The voice trials published 2026-08-11 mistral-nemo:12b arm — † POV lurches (Era two)voice-trials/index.html · Era two · the 24G budget, twelve candidates, blind · the row's verdict cell
  • The voice trials published 2026-08-11 mistral-nemo:12b arm — flat — real verdict (The re-bake)voice-trials/index.html · The re-bake · thirteen models, one scale, after the fix · the row's verdict cell

Mistral Small (mistral-small)Mistral

as the kits write itmistral-small · mistral-small 24B

  • The August arrivals published 2026-08-12 mistral-small arm — FAILaugust-arrivals/data/seat-rows.json · rows[].verdict
  • The voice trials published 2026-08-11 mistral-small 24B arm — contentless hedges (Era two)voice-trials/index.html · Era two · the 24G budget, twelve candidates, blind · the row's verdict cell
  • The voice trials published 2026-08-11 mistral-small 24B arm — hedging — real verdict, not envelope (The re-bake)voice-trials/index.html · The re-bake · thirteen models, one scale, after the fix · the row's verdict cell

Muse Glimmer (muse-glimmer)Muse

as the kits write itmuse-glimmer:30b · muse-glimmer:30b-q4_K_M-dflash · muse-glimmer:30b-q8_0-dflash · muse-glimmer:30b-q8_0-dflash+think

the licence ledger, on muse-glimmer:30b-q8_0-dflashbenched on both house exams 2026-08-12 — no seat
licence, read 2026-08-12the packaged blob carries no licence text at all — a single byte, which the receipt records as 0 characters and the sha256 of the empty string (e3b0c442…). Published as… read the row

  • The Instrument Travels published 2026-08-29 muse-glimmer:30b arm — FITS (fit-gate)the-instrument-travels/data/fit-muse-glimmer_30b.json · verdict
  • The Instrument Travels published 2026-08-29 muse-glimmer:30b arm — UNMEASURABLE (seat-43-judge-exam)the-instrument-travels/data/seat43-summary.json · candidates[].verdict
  • The chair trials published 2026-08-13 muse-glimmer:30b-q4_K_M-dflash armchair-trials/data/roster.json · rows[].model
  • The chair trials published 2026-08-13 muse-glimmer:30b armchair-trials/data/roster.json · rows[].model
  • The chair trials published 2026-08-13 muse-glimmer:30b-q8_0-dflash arm — EXPLORATORY (think:false · medium)chair-trials/data/c1-judge.json · rows[].outcome
  • The chair trials published 2026-08-13 muse-glimmer:30b-q8_0-dflash arm — EXPLORATORY (think:false · high)chair-trials/data/c1-judge.json · rows[].outcome
  • The outside judges published 2026-08-13 muse-glimmer:30b-q8_0-dflash armoutside-judges/data/panel-vendors.json · arm_order.order
  • The outside judges published 2026-08-13 muse-glimmer:30b-q8_0-dflash+think armoutside-judges/data/panel-vendors.json · arm_order.order
  • The August arrivals published 2026-08-12 muse-glimmer:30b-q8_0-dflash arm — UNMEASURABLE (16.3% truncated)august-arrivals/data/seat-rows.json · rows[].verdict
  • The seat trials published 2026-08-11 muse-glimmer:30b-q8_0-dflash arm — UNMEASURABLE (16.3% truncated)seat-trials/data/addendum-2026-08-12.json · rows[].verdict
  • The voice trials published 2026-08-11 muse-glimmer:30b-q8_0-dflash armvoice-trials/data/addendum-2026-08-12.json · rows[].model

Nemotron 3 (nemotron3)NVIDIA

as the kits write itnemotron3:33b

  • The chair trials published 2026-08-13 nemotron3:33b arm — UNMEASURABLE (think:false)chair-trials/data/c1-judge.json · rows[].outcome
  • The August arrivals published 2026-08-12 nemotron3:33b arm — PASSaugust-arrivals/data/seat-rows.json · rows[].verdict
  • The voice trials published 2026-08-11 nemotron3:33b arm — voice ok; wraps JSON in code fences (The re-bake)voice-trials/index.html · The re-bake · thirteen models, one scale, after the fix · the row's verdict cell

Nemotron 3 Nano (nemotron-3-nano)NVIDIA

as the kits write itcloud-nemotron-3-nano-30b

  • The open call published 2026-08-15 cloud-nemotron-3-nano-30b armthe-open-call/data/scores.json · arms

Nemotron 3 Ultra (nemotron-3-ultra)NVIDIA

as the kits write itcloud-nemotron-3-ultra · nemotron-3-ultra

  • Three New Voices at the Narrator's Chair published 2026-09-07 nemotron-3-ultra judge seatthree-new-voices-at-the-narrators-chair/data/seats.json · seats[].model
  • Two New Frontier Models at the Rules Desk published 2026-09-06 nemotron-3-ultra judge seattwo-new-frontier-models-at-the-rules-desk/data/seats.json · panel.seats[].model
  • The open call published 2026-08-15 cloud-nemotron-3-ultra armthe-open-call/data/scores.json · arms
  • The outside judges published 2026-08-13 cloud-nemotron-3-ultra judge seatoutside-judges/data/panel-vendors.json · panel.cloud_seats

Nemotron 3.5 Lightning (nemotron-3.5-lightning)NVIDIA

as the kits write itnemotron-3.5-lightning:30b · nemotron-3.5-lightning:30b-a3b

the licence ledger, on nemotron-3.5-lightning:30b-a3bbenched on both house exams 2026-08-12 — no seat
licence, read 2026-08-12the blob's first line, verbatim: "NVIDIA Open Model License Agreement" — 10 068 characters, 64 lines, sha256 aab95a9b…

  • The Instrument Travels published 2026-08-29 nemotron-3.5-lightning:30b arm — SPILLS (fit-gate)the-instrument-travels/data/fit-nemotron-3.5-lightning_30b.json · verdict
  • The chair trials published 2026-08-13 nemotron-3.5-lightning:30b-a3b arm — RANKED (think:false)chair-trials/data/c1-judge.json · rows[].outcome
  • The outside judges published 2026-08-13 nemotron-3.5-lightning:30b-a3b armoutside-judges/data/panel-vendors.json · arm_order.order
  • The August arrivals published 2026-08-12 nemotron-3.5-lightning:30b-a3b arm — UNMEASURABLE (51.2% truncated)august-arrivals/data/seat-rows.json · rows[].verdict
  • The seat trials published 2026-08-11 nemotron-3.5-lightning:30b-a3b arm — UNMEASURABLE (51.2% truncated)seat-trials/data/addendum-2026-08-12.json · rows[].verdict
  • The voice trials published 2026-08-11 nemotron-3.5-lightning:30b-a3b armvoice-trials/data/addendum-2026-08-12.json · rows[].model

Nemotron Cascade 2 (nemotron-cascade-2)NVIDIA

as the kits write itnemotron-cascade-2:30b

  • The August arrivals published 2026-08-12 nemotron-cascade-2:30b arm — FAIL (12/129 responses unmeasurable)august-arrivals/data/seat-rows.json · rows[].verdict
  • The voice trials published 2026-08-11 nemotron-cascade-2:30b arm — † summary stubs (Era two)voice-trials/index.html · Era two · the 24G budget, twelve candidates, blind · the row's verdict cell
  • The voice trials published 2026-08-11 nemotron-cascade-2:30b arm — invents a diver's name — an abstention failure (The re-bake)voice-trials/index.html · The re-bake · thirteen models, one scale, after the fix · the row's verdict cell

Nomic Embed Text (nomic-embed-text)Nomic

the licence ledger, on nomic-embed-textthe world's memory — the pinned embedder (swapping it means re-writing every vector, so it is never swapped casually)
licence, read 2026-08-11Apache 2.0

No bench rows on this hub — it reaches this roster through the licence ledger, which is where its receipt is.

OLMo 3.1 (olmo-3.1)Ai2

as the kits write itolmo-3.1:32b-think-q4_K_M

the licence ledger, on olmo-3.1:32b-think-q4_K_Mbenched on the judge exam 2026-08-12 — no seat
licence, read 2026-08-12the blob's first line, verbatim: "Apache License" — 11 359 characters, 203 lines, sha256 8c6db340…

  • The chair trials published 2026-08-13 olmo-3.1:32b-think-q4_K_M arm — EXPLORATORY (think:false)chair-trials/data/c1-judge.json · rows[].outcome
  • The August arrivals published 2026-08-12 olmo-3.1:32b-think-q4_K_M arm — UNMEASURABLE (55.8% truncated)august-arrivals/data/seat-rows.json · rows[].verdict
  • The seat trials published 2026-08-11 olmo-3.1:32b-think-q4_K_M arm — UNMEASURABLE (55.8% truncated)seat-trials/data/addendum-2026-08-12.json · rows[].verdict

Qwen 2.5 (qwen2.5)Alibaba

as the kits write itqwen2.5:7b

  • The voice trials published 2026-08-11 qwen2.5:7b arm — deflects (Era one)voice-trials/index.html · Era one · the 8G VRAM rig, nine models · the row's verdict cell

Qwen 3 VL (qwen3-vl)Alibaba

the licence ledger, on qwen3-vl:2bthe world's eyes — vision intake
licence, read 2026-08-11Apache 2.0

No bench rows on this hub — it reaches this roster through the licence ledger, which is where its receipt is.

Qwen 3.5 (qwen3.5)Alibaba

as the kits write itcloud-qwen3.5-397b · qwen3.5:27b · qwen3.5:397b · qwen3.5:9b-q8_0

the licence ledger, on qwen3.5:27bthe re-bake's best voice — benched, awaiting a second card
licence, read 2026-08-11Apache 2.0

  • Three New Voices at the Narrator's Chair published 2026-09-07 qwen3.5:397b judge seatthree-new-voices-at-the-narrators-chair/data/seats.json · seats[].model
  • Two New Frontier Models at the Rules Desk published 2026-09-06 qwen3.5:397b judge seattwo-new-frontier-models-at-the-rules-desk/data/seats.json · panel.seats[].model
  • The open call published 2026-08-15 cloud-qwen3.5-397b armthe-open-call/data/scores.json · arms
  • The chair trials published 2026-08-13 qwen3.5:27b arm — RANKED (think:false)chair-trials/data/c1-judge.json · rows[].outcome
  • The August arrivals published 2026-08-12 qwen3.5:27b arm — UNMEASURABLE (83.7% truncated)august-arrivals/data/seat-rows.json · rows[].verdict
  • The voice trials published 2026-08-11 qwen3.5:27b arm — † 3rd person, repeats — remember this one (Era two)voice-trials/index.html · Era two · the 24G budget, twelve candidates, blind · the row's verdict cell
  • The voice trials published 2026-08-11 qwen3.5:9b-q8_0 arm — † echoes internal ids (Era two)voice-trials/index.html · Era two · the 24G budget, twelve candidates, blind · the row's verdict cell
  • The voice trials published 2026-08-11 qwen3.5:27b arm — rescued — the run's best voice (The re-bake)voice-trials/index.html · The re-bake · thirteen models, one scale, after the fix · the row's verdict cell
  • The voice trials published 2026-08-11 qwen3.5:9b-q8_0 arm — leaks bare strings (The re-bake)voice-trials/index.html · The re-bake · thirteen models, one scale, after the fix · the row's verdict cell

Qwen 3.6 (qwen3.6)Alibaba

as the kits write itlocal-qwen3.6-27b · qwen3.6:27b · qwen3.6:latest

the licence ledger, on qwen3.6:27bbenched on both house exams 2026-08-12 — no seat
licence, read 2026-08-12the blob's first line, verbatim: "Apache License" — 11 358 characters, 203 lines, sha256 5e9ea707…

  • The Instrument Travels published 2026-08-29 qwen3.6:27b arm — FITS (fit-gate)the-instrument-travels/data/fit-qwen3.6_27b.json · verdict
  • The Instrument Travels published 2026-08-29 qwen3.6:27b arm — UNMEASURABLE (seat-43-judge-exam)the-instrument-travels/data/seat43-summary.json · candidates[].verdict
  • The map nobody picks up published 2026-08-17 qwen3.6:27b armllms-txt/data/runtime-pins.json · roster[].tag
  • The open call published 2026-08-15 local-qwen3.6-27b armthe-open-call/data/scores.json · arms
  • The chair trials published 2026-08-13 qwen3.6:27b arm — RANKED (think:false)chair-trials/data/c1-judge.json · rows[].outcome
  • The narrator’s chair, refused published 2026-08-13 qwen3.6:27b arm — REJECTEDcove-voice-head-to-head/data/gate.json · verdict
  • The outside judges published 2026-08-13 qwen3.6:27b armoutside-judges/data/panel-vendors.json · arm_order.order
  • The August arrivals published 2026-08-12 qwen3.6:27b arm — UNMEASURABLE (62.8% truncated)august-arrivals/data/seat-rows.json · rows[].verdict
  • The seat trials published 2026-08-11 qwen3.6:27b arm — UNMEASURABLE (62.8% truncated)seat-trials/data/addendum-2026-08-12.json · rows[].verdict
  • The voice trials published 2026-08-11 qwen3.6:27b armvoice-trials/data/addendum-2026-08-12.json · rows[].model
  • The voice trials published 2026-08-11 qwen3.6:latest arm — fastest; novelistic drift; over budget (Era two)voice-trials/index.html · Era two · the 24G budget, twelve candidates, blind · the row's verdict cell
  • The voice trials published 2026-08-11 qwen3.6:latest arm — voice fine; leaks entity ids into prose (The re-bake)voice-trials/index.html · The re-bake · thirteen models, one scale, after the fix · the row's verdict cell

Qwen 3.8 (qwen3.8)Alibaba

as the kits write itlocal-qwen3.8-27b · qwen3.8:27b

the licence ledger, on qwen3.8:27bthe doorman — RuleSage's live question-screening seat (hired 2026-08-17; the audit is exhibit fourteen)
licence, read 2026-08-18Apache 2.0 (the text packaged with the weights, read in full — 11,346 bytes, sha256 225c54bacb96eca3…)

The painters

The image models. They reach this roster through the licence ledger rather than through a kit — the diffusion bench is a separate wing of this site and publishes no per-model JSON here — so each carries the ledger’s own words and its own read date, and nothing else. The renders themselves are in the diffusion research bench, and where that wing publishes a gallery of one painter’s own work, the row below links it. A painter with no line there has none on the public wing — the bench shows only what its licence law clears, and a gallery behind a login is not a gallery this page will point at.

20 rows · jump to one

AlbedoBase XL v2.1 (albedobaseXL v2.1)an SDXL fine-tune, published on Civitai

the licence ledger, on albedobaseXL v2.1the painter of record
licence, read 2026-08-11the creator's permission flags as recorded on Civitai (model 140737), over a CreativeML Open RAIL++-M SDXL base: Image ✓ · RentCivit ✓ · Rent ✓ · Sell ✗ — credit… read the row

on the diffusion wingits bench gallery — the house spread — one painter, seven ways of seeing

No bench rows on this hub — it reaches this roster through the licence ledger, which is where its receipt is.

Chroma1-HD (chroma1-hd (fp8mixed))lodestones (independent)

the licence ledger, on chroma1-hd (fp8mixed)published
licence, read 2026-08-23Apache License 2.0 — declared in prose, matched against text (no LICENSE file exists in any of the author's four model repositories, checked 2026-08-23. The declaration… read the row

on the diffusion wingits bench gallery — the step ladder — chroma1-hd

No bench rows on this hub — it reaches this roster through the licence ledger, which is where its receipt is.

ERNIE-Image-Turbo (ernie-image turbo)Baidu

the licence ledger, on ernie-image turbopublished
licence, read 2026-08-23Apache License 2.0 — shipped beside the weights, which is rarer here than it should be (opened 2026-08-23: 11 366 bytes, sha256 5d7deeaf…; the only difference from the… read the row

on the diffusion wingits bench gallery — 512 to 1024 — ernie-turbo

No bench rows on this hub — it reaches this roster through the licence ledger, which is where its receipt is.

FLUX.1 dev (flux.1 [dev])Black Forest Labs

the licence ledger, on flux.1 [dev]held
licence, read 2026-08-23FLUX.1 [dev] Non-Commercial License v1.1.1 (re-read in full 2026-08-23 for the new-arrivals bench, 18 621 bytes, alongside the v2.1 text the 2026 repacks point at — the… read the row

No bench rows on this hub — it reaches this roster through the licence ledger, which is where its receipt is.

FLUX.1 schnell (flux.1 [schnell])Black Forest Labs

the licence ledger, on flux.1 [schnell]the fast brush — the 4-step painter behind 2,296 of the campaign census's 2,828 renders (the census)
licence, read 2026-08-11Apache 2.0 — the platform's licence record (tag: apache-2.0); the repo itself is download-gated, so this stamp covers the record. The full text was read earlier; that… read the row

No bench rows on this hub — it reaches this roster through the licence ledger, which is where its receipt is.

FLUX.2 [dev] (flux.2 [dev] (fp8mixed))Black Forest Labs

the licence ledger, on flux.2 [dev] (fp8mixed)held
licence, read 2026-08-23FLUX Non-Commercial License v2.1 (opened 2026-08-23 at the vendor's own code repository, which is where the canonical copy of this licence lives — 18 157 bytes, sha256… read the row

No bench rows on this hub — it reaches this roster through the licence ledger, which is where its receipt is.

FLUX.2 [klein] 4B (flux.2 [klein] 4B)Black Forest Labs

the licence ledger, on flux.2 [klein] 4Bpublished
licence, read 2026-08-23Apache License 2.0 — the vendor's own LICENSE.md, beside the weights (opened 2026-08-23: 9 584 bytes, sha256 ca02bc51…; diffed against the canonical Apache text —… read the row

on the diffusion wingits bench gallery — flux2-klein-4b-bf16 walks the nine worlds

No bench rows on this hub — it reaches this roster through the licence ledger, which is where its receipt is.

FLUX.2 [klein] 9B-KV (flux.2 [klein] 9B-KV (fp8))Black Forest Labs

the licence ledger, on flux.2 [klein] 9B-KV (fp8)held
licence, read 2026-08-23FLUX Non-Commercial License v2.1 (opened 2026-08-23: the edit variant's LICENSE is byte-identical to Black Forest Labs' canonical v2.1 text, 18 157 bytes, sha256… read the row

No bench rows on this hub — it reaches this roster through the licence ledger, which is where its receipt is.

FLUX.2-small-decoder (flux.2-small-decoder)Black Forest Labs

the licence ledger, on flux.2-small-decoderan accessory, not a painter — read on the way to the bench, never wired
licence, read 2026-08-23Apache License 2.0 — declared, not shipped (the repository holds no LICENSE file, checked 2026-08-23; the claim is the repo's own metadata plus the model card's sentence… read the row

No bench rows on this hub — it reaches this roster through the licence ledger, which is where its receipt is.

HiDream-O1-Image-Dev (hidream-o1-image-dev (fp8))HiDream.ai

the licence ledger, on hidream-o1-image-dev (fp8)published
licence, read 2026-08-23MIT License (opened 2026-08-23: 1 066 bytes, sha256 05660a75…, canonical MIT text with "Copyright (c) 2026 HiDream.ai" and nothing added. No LICENSE file sits in any of… read the row

on the diffusion wingits bench gallery — hidream-o1-dev-fp8 walks the nine worlds

No bench rows on this hub — it reaches this roster through the licence ledger, which is where its receipt is.

Juggernaut XL v9 (Juggernaut-XL v9)an SDXL fine-tune, published on Civitai

the licence ledger, on Juggernaut-XL v9bench candidate only — never wired into a product
licence, read 2026-08-11the creator's permission flags as recorded on Civitai (model 133005): Image ✓ · RentCivit ✓ · Rent ✗ · Sell ✗ — credit required

on the diffusion wingits bench gallery — nine worlds, three painters

No bench rows on this hub — it reaches this roster through the licence ledger, which is where its receipt is.

Kandinsky 5.0 Image Lite (kandinsky 5.0 image lite)Kandinsky Lab

the licence ledger, on kandinsky 5.0 image litepublished
licence, read 2026-08-23The MIT License (MIT) (opened 2026-08-23: 1 099 bytes, sha256 94946098…, canonical MIT with "Copyright (c) 2025 Kandinsky Lab" filled in and nothing added. The weights… read the row

on the diffusion wingits bench gallery — 512 to 1024 — kandinsky5-lite

No bench rows on this hub — it reaches this roster through the licence ledger, which is where its receipt is.

Krea 2 Turbo (krea 2 turbo (fp8_scaled))Krea AI

the licence ledger, on krea 2 turbo (fp8_scaled)published
licence, read 2026-08-23Krea 2 Community License Agreement v.1, dated 2026-06-22 (a six-page PDF, read in full 2026-08-23: sha256 b82a2805…, byte-identical at the pinned copy the model card… read the row

on the diffusion wingits bench gallery — krea2-turbo-fp8 walks the nine worlds

No bench rows on this hub — it reaches this roster through the licence ledger, which is where its receipt is.

Mage-Flow 4B Turbo (mage-flow 4B turbo (int8))Microsoft Research

the licence ledger, on mage-flow 4B turbo (int8)published
licence, read 2026-08-23MIT License (opened 2026-08-23: 1 066 bytes, sha256 275b4dd6…, standard MIT with "Copyright (c) 2026 Microsoft"; the vendor's README carries a per-model table whose row… read the row

on the diffusion wingits bench gallery — the step ladder — mageflow-turbo

No bench rows on this hub — it reaches this roster through the licence ledger, which is where its receipt is.

Pixel Art XL (LoRA) (pixel-art-xl (LoRA))a community LoRA

the licence ledger, on pixel-art-xl (LoRA)quarantined
licence · not re-read this cyclecreator's terms: non-commercial — verdict carried from the quarantine read, not re-read this cycle; queued

No bench rows on this hub — it reaches this roster through the licence ledger, which is where its receipt is.

RealVis XL 5.0 · Pixel Art Redmond (realvis xl 5.0 · pixel-art-redmond)community releases

the licence ledger, on realvis xl 5.0 · pixel-art-redmondwithheld
licence · not re-read this cyclenot read first-hand

No bench rows on this hub — it reaches this roster through the licence ledger, which is where its receipt is.

Stable Diffusion 3.5 · HiDream · Lumina (sd 3.5 · hidream · lumina)candidate painters, not yet read

the licence ledger, on sd 3.5 · hidream · luminacandidate painters — in the census, not yet read
licence · not re-read this cyclevarious — readings queued

No bench rows on this hub — it reaches this roster through the licence ledger, which is where its receipt is.

Stable Diffusion XL base 1.0 (sd_xl_base 1.0)Stability AI

the licence ledger, on sd_xl_base 1.0the sketch pipeline's base
licence, read 2026-08-11CreativeML Open RAIL++-M (2023-07-26)

on the diffusion wingits bench gallery — the step bench — two sdxl painters, three step counts

No bench rows on this hub — it reaches this roster through the licence ledger, which is where its receipt is.

Z-Image base (z-image base (undistilled))Alibaba Tongyi-MAI

the licence ledger, on z-image base (undistilled)published
licence, read 2026-08-23Apache License 2.0 — the same vendor file as the turbo row (opened 2026-08-23; the same code repository's README announces both the turbo and the base release, so one… read the row

on the diffusion wingits bench gallery — the step ladder — z-image-base

No bench rows on this hub — it reaches this roster through the licence ledger, which is where its receipt is.

Z-Image Turbo (z-image turbo (bf16))Alibaba Tongyi-MAI

the licence ledger, on z-image turbo (bf16)published
licence, read 2026-08-23Apache License 2.0 (opened 2026-08-23 at the vendor's code repository: 201 lines, md5 86d3f3a9…; diffed whitespace-insensitively against the canonical Apache text, the… read the row

on the diffusion wingits bench gallery — the colour question — z-image-turbo-bf16

No bench rows on this hub — it reaches this roster through the licence ledger, which is where its receipt is.

The music models

The models that work in sound rather than in language — they are on this roster because the licence ledger carries a card for each, and what a card says is quoted here in the ledger’s own words, with the date it was read.

2 models · jump to one

ACE-Step 1.5 (ACE-Step 1.5 (turbo · acestep-5Hz-lm))ACE Studio & StepFun (MIT weights, per the card)

the licence ledger, on ACE-Step 1.5 (turbo · acestep-5Hz-lm)the song engine — local full-song generation; the source tracks the demos are built from
licence, read 2026-08-25MIT (the repo LICENSE and the weights' HF cards read end-to-end — plain MIT, no acceptable-use rider; the bundled Qwen3-Embedding component is separately Apache 2.0)

No bench rows on this hub — it reaches this roster through the licence ledger, which is where its receipt is.

Demucs (htdemucs) (Demucs · htdemucs / htdemucs_6s)Meta Platforms — the adefossez fork, MIT

the licence ledger, on Demucs · htdemucs / htdemucs_6sthe splitter — 4- and 6-way stem separation, run on-prem on CPU
licence, read 2026-08-26MIT (© Meta Platforms; the adefossez fork LICENSE read in full, no weights carve-out)

No bench rows on this hub — it reaches this roster through the licence ledger, which is where its receipt is.

The toolchain · in the ledger, not models

Software this workshop runs beside its models and reads the licence of first-hand — a codec, a converter, a binary invoked on the generation host. They are here because the licence ledger carries a card for each, and the rule of this page is that a ledger row is never dropped. They are not models and are not counted as any: no weights, no maker of a model, and no bench they could sit — so each card carries what the thing is and the ledger’s own words, and stops there.

1 row · jump to one

FFmpegthe FFmpeg project

what it isa codec, not a model — the encoder behind the sound department's MP3 and WAV export, run as a standalone binary on the generation host

the licence ledger, on FFmpeg (static build)the encoder — MP3 / WAV export off the generation host
licence · installed 2026-08-25GPL (the johnvansickle static build bundles GPL-licensed encoders; FFmpeg's own core is LGPL-2.1+. Invoked only as a standalone binary — never linked into our code)

What is not on this page, and why

Two kinds of thing are missing on purpose, and both are listed rather than left out, because a name absent from an index reads the same as a name nobody looked for.

Published exhibits with no per-model rows. Real exhibits; they simply measure something other than a model.

  • A short history of vLLM — a HISTORY of a model server, the a-short-history-of-mistral shape, and it publishes no measurement of its own: its only speed figures are the project's own claims, each dated, none re-run by us. The models it names are the SUBJECTS of that history, never an arm: no model is scored against another, no pass bar was set, and the page ships no data kit, so there is no per-model row for an extractor to read.
  • A short history of Ollama — a HISTORY of a model server and the company behind it, the a-short-history-of-mistral shape, and it publishes no measurement of its own. The models it names (among them Mistral Small 3.2, the doorman Ollama runs at our doors, and Gemma 3, Llama 3.2 Vision and gpt-oss in its release history) are the SUBJECTS of that history or the load a server carries, never an arm: no model is scored against another, no pass bar was set, and the page ships no data kit, so there is no per-model row for an extractor to read.
  • Six graphics cards, the same questions — a COMPARISON of HARDWARE that measures nothing new and runs nothing: its seventeen tables select every figure but two from the hardware roster, out of six earlier pages, each of which this map already names for a reason of its own, and the two are prices read for one card. The models those pages ran (Gemma 4 and Mistral Small 3.2 writing, FLUX.2 Klein 4B drawing) are the LOAD each card carried, never an arm; no model is scored against another, no pass bar is set, and the page ships no data kit, so there is no per-model row for an extractor to read.
  • Small open deciders, measured: a word counter kept up on our hardest question — a trial of sixteen open models and a word counter, and its per-model rows are OWED an extractor rather than absent. This kit and the Reading the answer kit carry every published arm's scored rows and reports, each under its bench's own rows/ folder beside its pre-registration, and every arm's counts are printed on the page. What this file does not have yet is the IDENTITY half: the small deciders (Kev, Lev, APUS-OpenJev-v1, imajev, OpenDecider, deem, Von 1.3 and Laya), the open decision weights (CC BY-NC 4.0, bench-only by the page's own rule, never a seat), the jevify fine-tune and this workshop's own word counter are registered nowhere in MODELS, and none has a licence-ledger row. Registering them is an identity ruling, not a release-cut edit; until it is made the exhibit is named here rather than its arms being dropped without a word. The two chat models the page uses as yardsticks already reach this roster through earlier benches.
  • ONNX Runtime's telemetry on Linux, measured: on by default since 1.29, and the switch that stops it — a BENCH page of one LIBRARY, not a trial of models: it runs ONNX Runtime at three versions with a one-addition model built in memory, and names Kokoro only as the speech model two of the workshop's apps run on top of that library. No model was sat, scored or compared, so there is no per-model row for an extractor to read; the kit's result files record connection attempts, not a model's answers.
  • A sealed 3090 Ti FTW3 Ultra, four years after EVGA left the business — a workshop NOTE about one graphics card and the company that built it, not a trial of models: it measures nothing new and names no model, and every figure on it is quoted from a page this shelf already published, which names its own models. No model was sat, no pass bar was set, and the page ships no data kit, so there is no file for an extractor to read.
  • The same eighteen pictures, card by card — a HARDWARE bench in which the painters are the LOAD, not arms — the a-rig-your-friend-already-owns shape on six cards. What varies is the CARD and its computer (an RTX PRO 6000 Blackwell, an RTX 5080, an RTX 5090 Laptop GPU, an RTX 3080 Ti Founders Edition and, from the page's dated section of 2026-09-26, an EVGA GeForce RTX 3090 Ti FTW3 Ultra in the RTX PRO 6000's computer and an EVGA GeForce RTX 3090 XC3 Ultra in the 3080 Ti's) and its setting (the higher one: stock on the desktop cards, Dynamic Boost on the laptop, 350 W on the 3090 Ti; and 300 W on the desktop cards); FLUX.2 [klein] 4B, Z-Image-Turbo and FLUX.1 [dev] are held fixed down to the file, the recipe, the prompts and the seed formula on every card, and every ratio the page prints is card over card on the SAME model, or one card against itself. No painter is scored against another painter and no verdict is issued, and the kit's first drop is a tables file in markdown with no per-model rows an extractor could read.
  • Who makes FLUX: a short history of Black Forest Labs — a HISTORY of a company, not a trial of models, and like a-short-history-of-mistral it measures nothing new. The models it names (FLUX.1 [pro], [dev] and [schnell], FLUX.2 [dev] and [klein], FLUX 3 and FLUX 3 Action) are the SUBJECTS of its history, each with its licence quoted; where two of them were timed, the diffusion bench beside it, the-same-eighteen-pictures-card-by-card, carries the figures and says why it has no model rows either. No model is scored here, no pass bar was set, and the page ships no data kit.
  • One 3090 out, one 3090 Ti in, and the cap that decided it — a workshop note about HARDWARE moving, not a trial of models — the sibling of exhibit fifty-three. One dense model, `mistral-small3.2:24b`, ran the same frozen prompt on both boards on the workstation's cap ladder, with `gemma4:26b` as the mixture control, so the models are the load and the control and the CARD and its cap are what changed (an EVGA 3090 XC3 Ultra at 300 W out, an EVGA 3090 Ti FTW3 Ultra at 350 W in). No model was sat, no pass bar was registered, and the page ships no data kit yet — its own words say the files follow once private details are removed — so there is no file for an extractor to read.
  • RTX 5080 vs RTX 3080 Ti — a bench of HARDWARE, not a trial of models — the one-3090-against-one-3090-ti shape with three boards through one seat. What varies is the CARD (a 16 GB RTX 5080 against a 12 GB 3080 Ti Founders Edition and a 10 GB 3080, days apart) and its cap (300 W and each card’s own default); the seven core-bank builds of four language models, the picture pipelines, the adapter run and the game turn are the load and the control, the same digests, the same frozen prompt and the same runtime on every card. Nothing is auditioned and no model is being chosen, and the page ships no data kit yet — its own words say the bench’s record follows once cleared — so there is no file for an extractor to read.
  • How the long table works — a FIELD GUIDE to a shipped exhibit, not a bench. It opens up the long table — who is at the table and what each guest was told, the three rounds and the MC, the checks on every line, what is kept and for how long — and the models it names are the table's PRODUCTION SEATS, named so a reader knows who plays each chair: Gemma 4 and Mistral Small 3.2, both Apache-2.0, on this workshop's own machines. IT DOES REPORT ONE SITTING OF THE TWO AGAINST EACH OTHER, the pre-opening test run of 2026-09-12 — slips past a guest's year counted out of 90 answers each, and a judge that preferred one model's guests in 47 of 48 paired comparisons — and it reports them as COUNTS IN PROSE: this exhibit's kit carries the privacy register and the doorman's red-team receipt and no per-answer rows, so an extractor would have no file and no field to read. A roster row for that sitting is owed to whichever exhibit publishes its rows; naming this page here says that rather than dropping it.
  • Reading the answer instead of writing it — a trial of a READOUT and of the models built for it, and its per-model rows are OWED an extractor rather than absent. Every arm's scored rows, its own report and the residency census it proved before it ran are in this exhibit's kit (data/rows*/ — the *.report.json beside each arm's rows), and every arm's accuracy, Brier score and latency is printed on the page. What this file does not have yet is the IDENTITY half: the open decision weights (CC BY-NC 4.0 — bench-only by the page's own rule, never a seat), a community adapter under the Gemma terms, the Kev family, and an encoder trained here on this workshop's own labels are registered nowhere in MODELS, and none has a licence-ledger row. Registering them is an identity ruling, not a release-cut edit; until it is made the exhibit is named here rather than its arms being dropped without a word. Only the chat model the page uses as its control already reaches this roster, through the earlier benches.
  • A laptop, asked the desktop's questions — a bench of HARDWARE, not a trial of models — the one-3090-against-one-3090-ti shape on a board that is soldered down. What varies is the CARD’s power posture (a fixed 95 W, Dynamic Boost floating 140–150 W, and boost with the boot-time clock lock in force); the two language models are the load and the control, version-frozen and unchanged across every arm — the same mixture model and the same dense model this shelf’s other 24 GB benches carry, at the same windows, through the same version-frozen instance and the same frozen prompt — and the render rows run the print lab’s own recipes on the same ComfyUI release. Nothing is auditioned, nothing is scored against a bar, and no model is being chosen: every desktop row is a page this shelf already published, re-printed at its own cap.
  • Four homes for one reranker: what RuleSage's slowest stage cost at each stop — a bench of ADDRESSES rather than a trial of models — one cross-encoder at one pinned revision, the same blob by checksum on every box, timed at four machines so that the machine is the only thing that changes. Nothing was sat: no second model was asked, nothing is scored or ranked against another model, and the 0.90 bar the page prints is an agreement threshold between one model's int8 graph and its own fp32 weights, on two devices, not a pass bar for a candidate. The page says in its own words that no arm carries relevance labels, so it never claims this reranker ranks better than anything. It publishes no kit, so no extractor has rows to read.
  • A 3090, in its own words — the one model named is the page's AUTHOR rather than a candidate — Mistral Small 3.2, already resident on the card the story is about, wrote the story, and the page credits it in the byline and in the thanks. Nothing was sat: no second model was asked (the page names that as deliberate), no pass bar was registered, and nothing is scored, ranked or set against another model. The model figures on the page are the seven calls' own tokens a second and word counts, which are the receipt for how the page was made, not a reading of the model against anything.
  • Introducing the playground — an introduction to a door and a note about its LAYOUT — no model was sat, nothing is scored or ranked, and no pass bar was registered. Models appear only as the credited painters of three door stills, read off each exhibit's own record, and the fourth still is somebody else's photograph; the long table's tile was composed by a deterministic tool with no model in its path. The page's own figures are pixels, page heights and the arithmetic of a grid.
  • Two 3090s out, one 3090 Ti in — a workshop note about HARDWARE moving, not a trial of models — one open-weight diffusion model (FLUX.2 klein 4B, with Qwen as its text encoder) drew every plate and every sleeve on BOTH sides of the move, through the same certified 32-step graph, so the model is the control and the board is the only thing that changed. No model was sat, no pass bar was registered, and nothing on the page is scored or ranked; the figures are render seconds and sleeve milliseconds off two services' own journals, set beside the companion bench's medians for the same batch.
  • The indifferent prediction — not a trial of models and not a bench at all — two readings of one figure somebody else published, the story that reading arrived into, and an argument about a kind of model nobody has built. This workshop ran nothing for it: no arm, no window, no cap, no pass bar, and not one figure on the page is a reading of ours. The models it names are named inside quotations of the publisher's own post and are attributed there, which is a citation rather than a sitting.
  • RTX 3090 vs RTX 3090 Ti: an inference and diffusion bench — a bench of HARDWARE, not a trial of models — two 24 GB boards a generation apart in name, in the same x16 slot of the same desktop on two nights, on the same supply and through the same script, at every power cap each was run at. The two open-weight models it runs are the CONTROL, the same blobs by digest on both benches, and every ratio the page prints is board-over-board on the SAME model: what each writes, what each holds, what a picture costs, what a cap takes away. Nothing is scored, ranked or set against another model, no pass bar was registered for either of them, and answer quality is named on the page as not measured at all.
  • Two 3080s against one 3090, and the cap decides — a bench of HARDWARE and a price page, not a trial of models — two used 10 GB cards against one 24 GB card in the same box, on the same supply, through the same script, at two power caps each, and a bill read from public listings in one dated half hour on 2026-09-17. The open-weight models it runs are the CONTROL, held constant across every column, and every ratio the page prints is rig-over-rig on the SAME model: what a context window costs, what the cap buys, what a long document costs to read, what a picture costs to draw. Nothing is scored, ranked or set against another model, no pass bar was registered for either of them, and the readings this shelf has taken of those models on other hardware belong to the pages that took them.
  • A Short History of Mistral — a HISTORY of a company, and the one exhibit on this shelf that measures nothing new. Every card reading it prints was published earlier by a bench that has its own roster rows — the same 24B model on the same cards, quoted with the page it came from beside it — and the one run it reports first, the two-card bench of 2026-09-16, sets two PIECES OF HARDWARE against each other on a fixed pair of models, not two models against each other. The models in that bench are the CONTROL, held constant across both columns down to the weight digest, which is the fifty-seven-milliseconds shape: the only ratios on the page are card-over-card on the SAME model, no pass bar was set for a model and no verdict is issued about which of them is better. The Mistral models the page names by version — medium-3.5, large-3, small — are named as the SUBJECTS OF ITS HISTORY, and where they were actually scored against each other (the grounded-judge trial, the outside-judge audition) the exhibits that ran those trials carry the rows.
  • Ask About This Page — a FIELD GUIDE to a shipped tool, not a bench. It opens up the room behind every article — what the one line at the top opens, where the answers come from, the six checks that run over a draft in public, what is kept about the person who asked and for how long — and the models it names are the service's PRODUCTION SEATS, named so a reader knows who answers them: one gemma-class open-weight writer drafting every answer with the whole article in its window, and a nomic-class embedder that ran ONCE, offline, to choose the neighbours at the foot of each page and runs for nothing at request time. NOT ONE OF THEM IS SCORED AGAINST AN ALTERNATIVE on the page: no arms were registered, no model was swapped for another, and no verdict is issued about which seat is better. THE PAGE DOES CARRY A BENCH, AND IT IS OF THE SERVICE, NOT OF A MODEL: nine sittings of a 62-question exam held the SAME seat to floors this workshop pre-registered and moved in the open, and what the table reports is the door's behaviour — how many questions were answered, how many refusals the model wrote itself, how many sentences the checks struck, and how long each arm took. A roster row records a sitting where models were set against each other; this exhibit held none.
  • A Dinner Party for the Dead — a PREVIEW of an exhibit that is not open to everyone yet, not a bench. It says what the long table is, shows the six guests on its shelf, and quotes four exchanges from the workshop's own test dinners. Two open-weight models are NAMED — Gemma 4 and Mistral Small, both Apache-2.0, both on this workshop's own machines — and they are named the way a receipt names them: each exchange says which one held which chair, so a reader can see who said what. NEITHER IS SCORED AGAINST THE OTHER. No arms were registered, no pass bar was set, no verdict is issued, and the page says in its own words that whether the two play the same guest the same way is a question the reader gets to ask and not one this workshop has answered. The bench comes after the door opens; a roster row today would record a sitting this exhibit did not hold.
  • The cost to purchase and run a very capable home AI rig — a PRICE page, not a bench — it reads public retail listings and one marketplace's published sale averages on one named morning, 2026-09-13 between 06:00Z and 06:14Z, and prints every figure beside the listing string it came from. NO MODEL RAN TO MAKE IT and none is asked anything: what this box can do is measured on other pages of this shelf and linked rather than repeated, and the roster rows that record those sittings belong to those pages. The only measurement taken here is a wall-meter reading of the box at idle; nothing is scored, ranked or set against another model, and a roster row would record a sitting this exhibit did not hold.
  • How the print lab works — a FIELD GUIDE to a shipped tool, not a bench. It opens up the print lab — what the darkroom does with two words, what is kept about the person who asked, and for how long — and the models it names are the lab's PRODUCTION SEATS, named so a reader knows who answers them: one open-weight painter (FLUX.2 klein, the four-billion-parameter one, Apache-2.0) drawing every plate, a doorman reading the words before they draw, and a picture gate judging each finished plate. NOT ONE OF THEM IS SCORED AGAINST AN ALTERNATIVE on the page: there are no arms, no pass bar was registered, and no verdict is issued. The page's numbers table is a set of LIVE READINGS of the public lab — each row naming the sitting it was taken at — rather than a run with rows to read, and it ships no kit for the same reason. The bench that CHOSE that painter is a different exhibit and the roster rows recording it belong to that one; what this page reports is seconds, not a ranking.
  • The Two-Hour Machine, Priced — a PRICE page, not a bench — it reads public retail listings on one named day, 2026-09-10, and prints every figure beside the listing string it came from. NO MODEL RAN TO MAKE IT and none is asked anything: the six writing speeds it quotes are exhibit thirty-three's own, unchanged and linked, and the roster rows that record them belong to that page. Nothing here is scored, ranked or set against another model; what varies is a dollar figure, and a roster row would record a sitting this exhibit did not hold.
  • The Beat Lab, asked and answered — the COMPANION to exhibit twenty-eight, not a bench — twenty-eight written answers derived from the Beat Lab's own question bank and set beside the guide sentences they cite. No model is asked for any answer on the page and none ran to make it; nothing is scored, ranked or compared.
  • Listen for Yourself — the LISTENING COMPANION to exhibit thirty-six, not a bench — it publishes audio, not rows. One generator (ACE-Step 1.5) is the SUBJECT throughout; the only thing that varies inside a group is which adapter is attached and at what strength. Nothing is scored, ranked or set against another model, and the page registers no verdict of its own: it plays every render of the arc and hands the judgement to the reader.
  • The Ceiling Is Not the Corpus — a FIELD GUIDE, piece four and the series' last adapter piece, not a model bench — one adapter trained on four hundred and sixteen house tracks, one model, nothing auditioned. ACE-Step 1.5 is the SUBJECT, never an arm: never set against another generator, never scored, never ranked. What the page varies is whether the adapter is attached and at what strength, on fixed requests and fixed seeds; its instruments are a pre-registered blind listening sitting with sentinels (the adapter did not clear it), a sealed CLAP coherence bench and a quality floor, every one reported with its nulls. No model row belongs to it.
  • Hear It for Yourself — the LISTENING COMPANION to The Ceiling Is Not the Corpus, not a bench — it plays every render of that run, labelled, with the blind sheet unsealed, and hands the judgement to the reader. No verdict of its own, no model row.
  • Ten Minutes with Living Artists — a FIELD GUIDE, piece three, not a model bench — it teaches one adapter and its one-artist control with one model and auditions nothing. ACE-Step 1.5 is the SUBJECT, never an arm: never set against another generator, never scored, never ranked. What the page varies is WHICH adapter is attached and at what strength, on fixed requests and fixed seeds. Its one quantitative instrument is a sealed CLAP embedding bench that measures CORPUS SPREAD and render-to-corpus movement, pre-registered before it ran and reported with its nulls. No blinded listening test; one informal ear-verdict, n = 1, quoted as exactly that.
  • Half an Hour with Dead Composers — a FIELD GUIDE, piece two, not a model bench — it teaches two more adapters and three merges with one model and auditions nothing. ACE-Step 1.5 is the SUBJECT, never an arm: never set against another generator, never scored, never ranked. What the page varies is WHICH adapter (or blend) is attached, on one request and one seed; its ruler (moment-by-moment distance) is used for ORDERINGS only, because the page measures the ruler’s own noise floor at 95 % of the largest effect and says so. No blinded listening test; one informal ear-verdict quoted as exactly that.
  • Teaching a Music Model Chopin in Five Minutes — a FIELD GUIDE, not a model bench — it teaches one method with one model and auditions nothing. ACE-Step 1.5 is the SUBJECT of the lesson, not an arm: it is never set against another music generator, never scored, never ranked, and no threshold is derived from it. What the page varies is a single adapter’s STRENGTH against the same model, the same request and the same seed — one thing against itself at five settings, which is the a-rig shape with a dial where the silicon usually goes. The one quantity the page does measure (waveform distance from the 0% baseline) it names as a crude ruler in its own prose, saying in as many words that it reports whether the sound CHANGED and never whether it changed toward Chopin. And the judgment that would earn an appearance has explicitly NOT been made: no blinded listening test has been run, the page says so twice, and building that gate properly is named as a later piece. A roster row here would record a sitting this exhibit did not hold. The model, its licence and its disclosure ask are named in the page’s own colophon and in the kit.
  • Two Hours at 12 tok/s — On Battery — a HARDWARE bench in which the MODELS are the load and the MACHINES are the arms — the a-rig shape with silicon instead of painters. What varies is the box a fixed set of weights runs on; the tag, the quantisation, the frozen prompt, the seeds, the generation settings and the engine version are held, and every ratio the page prints is one model against ITSELF on other hardware, or against another model to make a point about the MEMORY BUS — an MoE that wakes ~3B parameters beats a 9B dense one because CPU decode streams weights, which is a statement about bandwidth and not about either model's quality. The page says so in its own words: not a leaderboard, not a benchmark suite, and not a quality ranking of anything. Its judgment column is a FLOOR-CHECK that saturated — all six models cleared every bar of the sealed field exam at 57-59 of 60, so nothing was selected and nothing was rejected — and on the 374-item moderation set the page prints the four candidates' raw disagreement counts (2, 3, 4 and 5 items, every one on clean text) and explicitly DECLINES to rank a band that tight. Those exams are the same sealed fixtures the GPU-side seats sit and they ran on a GPU box here, so an arm on this exhibit would record a sitting this bench did not hold. Five of the six models already reach this roster or the licence ledger through earlier work. The sixth, phi4:14b, appears nowhere on this site: it is named here as a load, never scored, and the row it is owed is a LICENCE LEDGER row, which is that ledger's work and not this exhibit's — as are the ledger rows still owed for mistral-small and llama3.3
  • Six Worlds for the Print Lab — a GALLERY, not a bench — one painter (flux1-dev, whose row and the ledger's hold are on /licences/) paints six worlds large for the print lab; no seat is compared, so no model row is earned (the same footing as nine-worlds-one-dog)
  • Nine Worlds, One Dog — a GALLERY, not a bench — one painter (flux1-dev, whose roster row the diffusion benches already carry) renders one subject nine ways for the fun of it; nothing is measured, ranked or auditioned, so no appearance is earned here. The painter is named in the page prose and in every row of the kit's prompts.json.
  • What 150 Watts Buys — a POWER bench in which the models are the CONTROL, not arms — the a-rig shape with a different field moving. What varies between rungs is the card’s board power limit; the model, the quantisation, the frozen prompt, the seeds, the concurrency and the card itself are all held, and every ratio the page prints is ONE workload against ITSELF at another limit. The dense seat is never scored against the sparse one — they are named to show that which side of the cap a workload lives on is the whole result, which is a statement about the cap and not about either model. The same holds for the 2026-08-27 cap ladder under the kit’s ladder/: five image lanes, each paired against its OWN same-day 600 W rung, no lane compared to another lane. No pass bar was registered on either sitting and no verdict was issued about any model; the one ruling the page does make — the dial at 500 W — is about the workshop’s hardware. An arm here would record a sitting that never happened. Every model named already sits on this roster or in the licence ledger through earlier work: the 26B chat seat and the 24B dense seat through the serving benches, and the five image lanes through the ledger’s painters sections — extracting them again from a run that measured their POWER SUPPLY would be two derivations of one fact
  • How the Beat Lab works — a FIELD GUIDE to a shipped tool, not a bench. It opens up the Beat Lab — what is synthesised in the browser, what crosses to our machines, and what is kept — and the three models it names are the tool’s PRODUCTION SEATS, named so a reader knows who answers them: a Mistral Small 3.2 doorman, a Gemma 4 writer, and Kokoro reading the card aloud on the web server’s ordinary processor. Not one of them is scored against an alternative anywhere on the page: there are no arms, no pass bar was registered, no verdict is issued, and the page’s one table is a LIVE READING of the public site (median of three warm runs, taken 2026-08-28) rather than a run with rows to read — which is also why the exhibit ships no data kit. An arm here would record a sitting that never happened. Of the three, only the Gemma 4 seat already sits on this roster, through the earlier benches; the doorman’s exact build and the speech model are named on the page as production choices and are owed a licence-ledger row, which is that ledger’s work and not this exhibit’s
  • The Typist and the Developer — a CONTENTION bench, in which the two stacks are the instrument and each other's load — not arms. Only one language model and one painting graph appear at all, both of them the shipped production shapes the game already serves, and they are never scored against any alternative: every row is the SAME pair measured alone and then measured sharing a card, so the only ratios the page prints are one workload against ITSELF under load. No pass bar was registered, no verdict was issued, and the registration says so in its own opening words — “DESCRIPTIVE ROWS ONLY: no gate, no seat verdict, no adoption decision rides on this bench”. An arm here would record a sitting that never happened. Both already sit on this roster through earlier benches: the 26B chat seat through the payroll piece and the field guide, the sketch graph through the close-up and the licence ledger's painters section
  • A Rig Your Friend Already Owns — a HARDWARE bench in which the painters are the CONTROL, not arms — the fifty-seven-milliseconds shape one rung further out. What moves between the three rigs is the graphics card; every other field is held, and held provably: the same six checkpoint files, the same engine checkout, the same torch and driver, the same prompts, seeds, samplers and step ladders, with each of the 330 timed renders paired cell-for-cell against its own archived baseline render. No painter is scored against another painter anywhere on the page; the only ratios it prints are one rig over another rig on the SAME model, so an arm here would record a sitting that never happened. No pass bar was set and no verdict was issued — the page withholds even the multipliers it does not trust, and says which and why. All six already sit on this roster through the licence ledger's new-painters section (read first-hand 2026-08-23, published 2026-08-24), and extracting them a second time from a run that measured their host would be two derivations of one fact; the ledger row stays their only source, exactly as the new-arrivals row rules it
  • How a Vision Model Sees: Eyes for a Machine That Reads — a demonstration of a vision front end, NOT a seat bake-off — and the piece says so in its own words: “a demonstration with its counts stated, not a verdict”. The dozen leg it prints watches two seats (minicpm-v4.5, the one the game ships; gemma4:26b, the resident) against no pass bar at all: no threshold was registered for the dozen, no verdict was issued, and the run’s own SIXTH ADDENDUM forbids precisely this publication — self-refutation trigger SR1 fired over the sixty-photograph leg, and what it obliges is that no verdict table publishes anywhere until the seating smoke is re-run under this harness. Roster arms would be that table under another name, emitted the day before the bake-off registered to produce it. The likeness panel the page does print is a BLIND JUDGE reading of sentences, each judge barred from scoring its own lineage; the judges are the instrument there, not arms, and the seats are scored prose rather than seated. gemma4:26b already sits on this roster through earlier benches; the bake-off that will carry thresholds, verdicts and arms for all eight seats is registered and is Friday’s
  • Fifty-Seven Milliseconds: What a Diffusion Step Actually Buys — a step-count bench in which the PAINTER IS THE CONTROL, not an arm. Two ladders a month apart: forty-two renders at 512², every one of them AlbedoBase XL v2.1 at one sampler, one canvas and one seed; then six real sketch prompts at 768² on the season's new painters, one of them distilled. In each ladder the step count is the only field that moves — a verified property of the manifest (exactly one distinct value for every other field), not a promise. Where the piece prints two painters side by side it is contrasting two DIALS, and where each one stops converging, rather than scoring one painter against the other: no pass bar was set, no verdict was issued, and emitting an arm here would record a sitting that never happened — the same arm-versus-subject distinction the new kid's row is drawn on. Every painter it names already sits on this roster through the licence ledger (AlbedoBase XL v2.1 read first-hand 2026-08-11; the second run's four through the new-painters section, 2026-08-24), and the ledger row stays their only source
  • The Dog, the Dice, and the Painter: How a Small World Draws Itself — a walk through the shipped engine, not a bench — its one model table is a PRODUCTION CENSUS, counting what the live game has already painted (FLUX.1 [schnell] 2,296, FLUX.1 [dev] 367, AlbedoBase XL v2.1 165, total 2,828), which is a record of use rather than a measurement of any painter against another. Nothing on the page compares them, no pass bar was set and no run was registered, so there is no verdict a roster cell could carry; all three already hold rows here through the licence ledger, where the reads and the enforced verdicts live. The local vision model behind the piece's moderation gate is deliberately unnamed on the page — measuring it is what the week's vision piece is for, and it gets rostered where it is actually measured
  • The new-arrivals diffusion bench — the new-arrivals diffusion bench — thirteen image painters measured on the wing, which is built outside this repo and carries its own census. Its painters DO sit on this roster, but they enter it through the licence ledger's new-painters section (2026-08-24), whose rows carry the bench's own first-hand reads and enforced verdicts; extracting them a second time from the wing would be two derivations of one fact, so the ledger row is the roster's only source for each
  • It Was Already There: Where New Knowledge Actually Comes From — a field-guide reading of the published record — AlphaFold at CASP, halicin, Kepler’s 1631 transit, Dirac’s positron and AlphaGo’s move 37, then the GNoME and A-Lab papers and the corrections that followed them — with no house bench behind it and no arm of ours on it. The only models it names are other people’s, quoted from their own published results, so no row here is attributable to a model we run; the checks it points a reader at (RuleSage, and the seat trials and outside judges next door) are linked as things to go and try, and each is rostered where it was actually measured
  • It’s Not a Metaphor: A Language Model Is a Compression of Its Corpus — a field-guide explainer of the training objective — its three receipts name the MEASUREMENT rather than the vendor by the piece’s own rule (“We name the measurement, not the vendor, as a standing house rule”), so no arm on it is attributable to a model and none can be rostered; its two local candidates are given only as open-weight models of about 24 billion and about 12 billion parameters, put to identical prompts at identical settings, and the third receipt is an external published result, not ours
  • The free speed wasn’t free — the runtime dividend, traced to one commit — a field-guide explainer — no new models; its decode-rate rows re-quote published kits or run registered probes on already-rostered seats (the toggle, the toll probes, the version pair); no new roster arms
  • The compressed photograph — what those Q4_K_M tags actually mean — a field-guide explainer — its accuracy claims re-quote the-new-kid's published precision addendum; its robber rows are registered n=5 illustrations of one already-rostered model at four of its own builds (an intra-model quant ladder, not new roster arms); no new bench rows
  • Everyone on the payroll, three at the table — dense, MoE, and active weights — a field-guide explainer — its speed and residency figures re-quote the chair trials' and August arrivals' published rows under their own rules, its active-parameter figures are derived from model-file tensor shapes, and its robber rows are a registered n=5 illustration of two already-rostered models; no new bench rows
  • Reading is fast, writing is slow — why AI answers arrive word by word — a field-guide explainer of prefill/decode physics — its measurement kit times one already-rostered model (qwen3.8:27b, seated by the-new-kid) and adds no new models and no per-model comparison rows; the timing rows are physics receipts, not roster rows
  • Three librarians and a careful reader — how RuleSage finds the right page — a field-guide explainer of the retrieval pipeline — it NAMES the seats (the embedder, the re-ranker, the answerer, the doorman) but publishes no per-model measurements; every named seat's rows live on the exhibits and ledger pages it links
  • The move — the same yardstick, before and after. — a serving bench — its kit measures two hardware eras, not models
  • The diffusion research bench. — the image-pipeline wing; its arms are samplers and checkpoints, and it is built from another tree
  • The licence ledger — every model, read first-hand. — the ledger itself — it is this roster's licence column
  • How we work — who writes this, and how. — workshop notes, no bench rows
  • Where your question goes — a plain-english walk. — workshop notes, no bench rows

Identifiers that name a seat rather than a model. The kits label a judge’s chair or a control row with a name of its own. Resolving those to a vendor would assert something the field they sit in does not say, so they are excluded by name.

  • ANCHOR (reference driver) — the calibration anchor row, not a candidate
  • fable — a judge seat's name, not a model tag
  • fable-1 — a judge seat's name, not a model tag
  • fable-2 — a judge seat's name, not a model tag
  • judge-1 — a blind judge seat's label
  • judge-2 — a blind judge seat's label
  • judge-3 — a blind judge seat's label
  • openai-flagship — a seat label; the kit names its family (openai) and no model tag
  • opus — a judge seat's name, not a model tag
  • opus-1 — a judge seat's name, not a model tag
  • opus-2 — a judge seat's name, not a model tag
  • the four-seat house panel — the panel as a whole
  • warmup-DISCARDED — a discarded warm-up row

Provenance

  • Derived, not authored — every appearance, verdict and licence line on this page is read out of a published artifact by tools/roster_data.py, whose extractors each name the file and field they read. The line under each appearance is that name.
  • It fails closed — if an exhibit names a model this roster has never seen, the build refuses rather than dropping the model. A missing model is therefore a broken build, not a quiet omission.
  • No new claims — this page has no findings. Where a summary would need judgement the hub has not published, it links to the page that did the work.
  • The data — every kit linked from here is CC BY 4.0; so is this page. Attribution: strata→signal research, research.strata2signal.com/roster/.
  • Authorship — benched, drafted and audited by the workshop’s own agents under a human operator’s rulings, the same arrangement told in full here. If something on this page is wrong, tell us.

elsewhere in the workshop

a strata→signal property · hello@strata2signal.com · say hello