{
 "schema_version": "1.0",
 "page_kind": "article",
 "pack_kind": "generate",
 "slug": "small-open-deciders-measured",
 "title": "Small open deciders, measured: a word counter kept up on our hardest question",
 "dek": "A decider is a language model that reads a question and its options and returns a probability for each. We asked sixteen open models five questions from our own products. On the hardest, the top five small deciders labelled Apache-2.0, given no examples, could not be told apart from a word counter trained on our own lines. One of them, APUS-OpenJev-v1 9B, passed a pre-registered screening test that the model then guarding our exhibit missed by three lines, letting through only kinds of line its rules did not name.",
 "published": "2026-09-30",
 "series": [
  "bench"
 ],
 "licence": "CC BY 4.0",
 "status": "pending-judge",
 "pack_note": "no judge is seated, so every chip is held",
 "notice": "items with published:false are held and are not this page's published words",
 "source": {
  "url": "https://research.strata2signal.com/small-open-deciders-measured/",
  "md_url": "https://research.strata2signal.com/small-open-deciders-measured/index.md",
  "md_sha": "42058f51d29ada2b1ac4b5e17b6d5558d373360a39057ddc4baaf793544d6866",
  "html_sha": "694ffbf79dca9e1a082530ec8702f385afffebcf71eaf5d67bb9167c9ff30a5c",
  "anchors_sha": "2e13903b95726ed5e18a7212700792133db90692c07c5840455930cef080bf1f"
 },
 "short": {
  "paragraph": "A decider reads a question and its options and returns a probability for each: nothing to parse, and a number to set a hand-off threshold on, once that threshold is checked on your own cases. For a yes-or-no call about a short passage, eight of the ten models we asked scored 39 or 40 of 40 on one such question, the set's ceiling, so choose among those eight by licence and speed. For questions about long pages, Lev-4B took all 53 we sent and got 52 right. For screening what visitors type, APUS-OpenJev-v1 9B alone of the Apache-2.0 models we asked passed our pre-registered test, refusing 32 of 36 lines written to be refused and 0 of 12 harmless ones. For sorting short texts by how they are written, train a bag of words first: ours got 68 of 108, and the top five small deciders labelled Apache-2.0, at 66 to 72, could not be told apart from it. OpenJev 27B scored highest, 101, and its 4-bit build sits whole on one used RTX 3090, but its weights are licensed for non-commercial use, and commercial use needs separate terms, which its README now says to ask for by email.",
  "from_dek": false,
  "counts": {
   "words": 9566,
   "minutes": 43,
   "tables": 6,
   "kit": true
  },
  "bullets": [
   {
    "text": "the model Von 1.3 runs in fp32",
    "figure": "fp32",
    "cite": "what-it-takes-to-run",
    "quote": "What it takes to run Every small decider we put on a graphics card ran through its own Python server or package on one used 24 GB RTX 3090, in bf16 but for Von 1.3, which runs fp32 there as shipped, never at 4 or 8 bits and never through ollama, llama.cpp or vLLM.",
    "span": [
     38688,
     38692
    ]
   }
  ]
 },
 "sections": [
  {
   "id": "what-a-decider-is",
   "heading": "What a decider is",
   "level": 2,
   "span": [
    5436,
    6375
   ],
   "chunks": [
    [
     5436,
     6375
    ]
   ],
   "chars": 939,
   "digest": "a decider reads a question and labelled options to turn one score per option into probabilities that sum to one. nothing is generated, allowing a product to act above a threshold and hand unsure cases to a person. TypeSafe announced Jev, a hosted decider, on 2026-09-15.",
   "digest_skipped": null
  },
  {
   "id": "which-one-for-which-job",
   "heading": "Which one for which job",
   "level": 2,
   "span": [
    6375,
    12880
   ],
   "chunks": [
    [
     6375,
     12880
    ]
   ],
   "chars": 6505,
   "digest": null,
   "digest_skipped": null
  },
  {
   "id": "the-questions-we-asked",
   "heading": "The questions we asked",
   "level": 2,
   "span": [
    12880,
    16602
   ],
   "chunks": [
    [
     12880,
     16602
    ]
   ],
   "chars": 3722,
   "digest": "sixteen models were tested using various request shapes and layouts. the six-way task involves language models playing six long-dead guests at a dinner. the word counter is a multinomial naive Bayes written in standard-library Python. none of the models were built specifically for the six-way task.",
   "digest_skipped": null
  },
  {
   "id": "the-six-way-scoreboard",
   "heading": "The six-way scoreboard",
   "level": 2,
   "span": [
    16602,
    22306
   ],
   "chunks": [
    [
     16602,
     22306
    ]
   ],
   "chars": 5704,
   "digest": "the scoreboard tracks performance on the six-way and held-out lines. OpenJev 27B, 16-bit achieved 101 of 108. the text explains how to read the table, including notes on strict scoring, six orders, and exact ties. it also addresses whether any model saw the answers.",
   "digest_skipped": null
  },
  {
   "id": "a-word-counter-kept-up",
   "heading": "A word counter kept up",
   "level": 2,
   "span": [
    22306,
    25652
   ],
   "chunks": [
    [
     22306,
     25652
    ]
   ],
   "chars": 3346,
   "digest": "the article compares deciders against a word counter. Kev-9B's 72 against the word counter's 68 is four lines net. everything at Gemma 4's level or above is ahead of the word counter by more than chance explains. the word counter cannot screen a line or read a page.",
   "digest_skipped": null
  },
  {
   "id": "the-doorman",
   "heading": "The doorman: screening what visitors type",
   "level": 2,
   "span": [
    25652,
    30661
   ],
   "chunks": [
    [
     25652,
     30661
    ]
   ],
   "chars": 5009,
   "digest": "the doorman task involves screening what visitors type. a live door using an 8-bit Gemma 4 26B-A4B failed its gate. APUS-OpenJev-v1 9B at `high` is the only model that passed and whose weights carry the Apache-2.0 text, refusing 32 of 36 planted lines and 0 of 12 harmless ones.",
   "digest_skipped": null
  },
  {
   "id": "the-easier-questions",
   "heading": "The easier questions",
   "level": 2,
   "span": [
    30661,
    32999
   ],
   "chunks": [
    [
     30661,
     32999
    ]
   ],
   "chars": 2338,
   "digest": null,
   "digest_skipped": null
  },
  {
   "id": "the-licences",
   "heading": "The licences",
   "level": 2,
   "span": [
    32999,
    38484
   ],
   "chunks": [
    [
     32999,
     38484
    ]
   ],
   "chars": 5485,
   "digest": null,
   "digest_skipped": "credits"
  },
  {
   "id": "what-it-takes-to-run",
   "heading": "What it takes to run",
   "level": 2,
   "span": [
    38484,
    43094
   ],
   "chunks": [
    [
     38484,
     43094
    ]
   ],
   "chars": 4610,
   "digest": null,
   "digest_skipped": null
  },
  {
   "id": "three-teams-one-base",
   "heading": "Three teams, one base",
   "level": 2,
   "span": [
    43094,
    43754
   ],
   "chunks": [
    [
     43094,
     43754
    ]
   ],
   "chars": 660,
   "digest": "Lev-4B, imajev-4B and APUS-OpenJev-v1 4B all come from three teams using the same base. on the six-way, they land within two lines, with scores of 66, 68 and 67.",
   "digest_skipped": null
  },
  {
   "id": "a-boards-rank-and-what-ours-adds",
   "heading": "A board's rank, and what ours adds",
   "level": 2,
   "span": [
    43754,
    44960
   ],
   "chunks": [
    [
     43754,
     44960
    ]
   ],
   "chars": 1206,
   "digest": "imajev-4B was ranked first in JevBench v1.4.2.2. a board's rank blends accuracy with speed and cost over questions that are not yours. on the six-way, imajev-4B read 68 of 108.",
   "digest_skipped": null
  },
  {
   "id": "our-fine-tune-that-voided-itself",
   "heading": "Our fine-tune that voided itself",
   "level": 2,
   "span": [
    44960,
    46711
   ],
   "chunks": [
    [
     44960,
     46711
    ]
   ],
   "chars": 1751,
   "digest": "a fine-tune attempt on deem-0.8-v1 was voided because a control read 29 of 108. the design included three controls to ensure the recipe could not lift the six-way without learning. the control that read 29 had a training loss of 0.000.",
   "digest_skipped": null
  },
  {
   "id": "what-we-did-not-measure",
   "heading": "What we did not measure, or did not report",
   "level": 2,
   "span": [
    46711,
    49449
   ],
   "chunks": [
    [
     46711,
     49449
    ]
   ],
   "chars": 2738,
   "digest": null,
   "digest_skipped": null
  },
  {
   "id": "how-to-check-our-work",
   "heading": "How to check our work",
   "level": 2,
   "span": [
    49449,
    53444
   ],
   "chunks": [
    [
     49449,
     53444
    ]
   ],
   "chars": 3995,
   "digest": null,
   "digest_skipped": "credits"
  },
  {
   "id": "where-each-figure-comes-from",
   "heading": "Where each figure comes from",
   "level": 2,
   "span": [
    53444,
    58608
   ],
   "chunks": [
    [
     53444,
     58608
    ]
   ],
   "chars": 5164,
   "digest": null,
   "digest_skipped": null
  },
  {
   "id": "who-ran-this-and-thanks",
   "heading": "Who ran this, and thanks",
   "level": 2,
   "span": [
    58608,
    60471
   ],
   "chunks": [
    [
     58608,
     60471
    ]
   ],
   "chars": 1863,
   "digest": null,
   "digest_skipped": "credits"
  }
 ],
 "chips": [
  {
   "id": "c-06d1b4f9",
   "q": "how does a decider function compared to a paragraph?",
   "a": "A decider reads a question and labelled options to turn one score per option into probabilities that sum to one. Unlike a paragraph, which must be parsed and carries no honest number for model certainty, a decider allows a product to act above a threshold and hand unsure cases to a person.",
   "cites": [
    "what-a-decider-is"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-c6f2479e",
   "q": "which model performs best for yes-or-no calls on short passages?",
   "a": "Eight models scored 39 or 40 out of 40 on whether a rulebook excerpt answers a question. Laya scored 35, while Von 1.3 scored 29.",
   "cites": [
    "which-one-for-which-job"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-8d93b341",
   "q": "what methods were used to prompt the sixteen models?",
   "a": "Kev, Lev, APUS-OpenJev-v1, imajev, Von 1.3, and Laya use a request shape consisting of a context, a question, and options. OpenDecider and Deem use their own prompts, while OpenJev, jevify, Gemma 4, and Mistral Small 3.2 use the first Jev bench layout.",
   "cites": [
    "the-questions-we-asked"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-137521c5",
   "q": "what criteria defined the gate for the live doorman test?",
   "a": "The gate required that at least 30 of the 36 planted lines be refused and at most 2 of the 12 harmless lines be refused. The live doorman failed this, refusing 27 planted lines and 0 harmless lines.",
   "cites": [
    "the-doorman"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-55e91308",
   "q": "what happened to the deem-0.8-v1 fine-tune experiment?",
   "a": "The T1 fine-tune was declared VOID because three controls trained on shuffled labels read 18, 29, and 24. Since the recipe could lift the six-way without learning, any tuned figure would be void.",
   "cites": [
    "our-fine-tune-that-voided-itself"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-7e6c73ef",
   "q": "what information is provided regarding the source of the figures?",
   "a": "The article provides a table mapping each run to the specific model revision benched and the paths to the per-question row files used to derive the results.",
   "cites": [
    "where-each-figure-comes-from"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  }
 ],
 "related": [
  {
   "slug": "reading-the-answer",
   "why": "shares ground with § The questions we asked · § What we asked it"
  },
  {
   "slug": "the-new-kid",
   "why": "shares ground with § The six-way scoreboard · § The narrator's open call — mid-table, uncertainty printed"
  },
  {
   "slug": "the-open-call",
   "why": "shares ground with § The six-way scoreboard · § A late contestant, and it ties"
  }
 ],
 "thanks": "## Who ran this, and thanks {#who-ran-this-and-thanks}\n\n**TypeSafe** defined the request shape that Kev, Lev, APUS-OpenJev-v1, imajev, Von 1.3 and Laya answered here, and that OpenJev, jevify and deem-0.8-v1 say they accept. The **OpenJev** authors published their weights and their own 4-bit build; **`jaredpalmer`** published Kev; **Interfaze** published Lev; **APUS AI-LAB** (its cards name gumpcheng and zhangxu) published APUS-OpenJev-v1; **Mohit Garg** published imajev; **Manjunath Janardhan** (`manjunathshiva`) published OpenDecider nano and small; **LibertAI Labs** published Deem; **`kushalpatil`** published the jevify fine-tune, and **`mradermacher`** its GGUF builds; **Victor Hugo Panisa** (`wfzyx`) published Von 1.3, and **Convai Innovations** (`convaiinnovations`) published Laya. **Google** made Gemma 4, **Mistral AI** Mistral Small 3.2, **Alibaba's Qwen team** the Qwen3.8, Qwen3.5 and Qwen3 models under OpenJev and most of these deciders, **Johns Hopkins University** the Ettin encoder under OpenDecider nano, and **Answer.AI** and **LightOn** the ModernBERT-large encoder under Von 1.3 and Laya. **Benchmark Heaven** runs JevBench. The runs stood on **vLLM**, **ollama**, **llama.cpp**'s GGUF format, **transformers**, **PEFT**, **flash-linear-attention**, **FastAPI**, **uvicorn** and **PyTorch**. A reader on r/machinelearningnews pointed out that Von 1.3 and Laya were missing. None of them owed us anything. If you built one of these models and we have described it wrongly, tell us at [hello@strata2signal.com](mailto:hello@strata2signal.com); corrections are dated on this page.\n\nA small (human) team decided what to ask. A fleet of AI agents pre-registered each run, ran it, checked it, recounted the figures from the rows, and wrote this page. The agents are Anthropic's Claude models. No Anthropic model is measured on this page.",
 "kit": {
  "url": "https://research.strata2signal.com/small-open-deciders-measured/data/",
  "licence": "CC BY 4.0",
  "files": [
   "EVIDENCE-v2.md",
   "README.md",
   "READS-v2.md",
   "apus-2026-09-28/PREREG-apus.md",
   "apus-2026-09-28/a1b/receipts/counted-run.json",
   "apus-2026-09-28/a1b/receipts/counted-telemetry.jsonl",
   "apus-2026-09-28/a1b/rows/apus-4b-high.ka.idle.csv",
   "apus-2026-09-28/a1b/rows/apus-4b-high.ka.jsonl",
   "apus-2026-09-28/a1b/rows/apus-4b-high.ka.report.json",
   "apus-2026-09-28/a1b/rows/apus-4b-high.ka.watts.csv",
   "apus-2026-09-28/a1b/rows/apus-4b-high.kc.idle.csv",
   "apus-2026-09-28/a1b/rows/apus-4b-high.kc.jsonl",
   "apus-2026-09-28/a1b/rows/apus-4b-high.kc.report.json",
   "apus-2026-09-28/a1b/rows/apus-4b-high.kc.watts.csv",
   "apus-2026-09-28/a1b/rows/apus-4b-high.kperm.idle.csv",
   "apus-2026-09-28/a1b/rows/apus-4b-high.kperm.jsonl",
   "apus-2026-09-28/a1b/rows/apus-4b-high.kperm.report.json",
   "apus-2026-09-28/a1b/rows/apus-4b-high.kperm.watts.csv",
   "apus-2026-09-28/a1b/rows/apus-4b-low.kc.idle.csv",
   "apus-2026-09-28/a1b/rows/apus-4b-low.kc.jsonl",
   "apus-2026-09-28/a1b/rows/apus-4b-low.kc.report.json",
   "apus-2026-09-28/a1b/rows/apus-4b-low.kc.watts.csv",
   "apus-2026-09-28/receipts/counted-run.json",
   "apus-2026-09-28/receipts/counted-telemetry.jsonl",
   "apus-2026-09-28/rows/apus-9b-high.ka.idle.csv",
   "apus-2026-09-28/rows/apus-9b-high.ka.jsonl",
   "apus-2026-09-28/rows/apus-9b-high.ka.report.json",
   "apus-2026-09-28/rows/apus-9b-high.ka.watts.csv",
   "apus-2026-09-28/rows/apus-9b-high.kc.idle.csv",
   "apus-2026-09-28/rows/apus-9b-high.kc.jsonl",
   "apus-2026-09-28/rows/apus-9b-high.kc.report.json",
   "apus-2026-09-28/rows/apus-9b-high.kc.watts.csv",
   "apus-2026-09-28/rows/apus-9b-high.kperm.idle.csv",
   "apus-2026-09-28/rows/apus-9b-high.kperm.jsonl",
   "apus-2026-09-28/rows/apus-9b-high.kperm.report.json",
   "apus-2026-09-28/rows/apus-9b-high.kperm.watts.csv",
   "apus-2026-09-28/rows/apus-9b-low.kc.idle.csv",
   "apus-2026-09-28/rows/apus-9b-low.kc.jsonl",
   "apus-2026-09-28/rows/apus-9b-low.kc.report.json",
   "apus-2026-09-28/rows/apus-9b-low.kc.watts.csv",
   "cost-of-intelligence-2026-09-23/KIT.lock.six-way.json",
   "cost-of-intelligence-2026-09-23/benchbox-3090-reading-0/rows/b-mistral-small3.2-24b/b-mistral-small3.2-24b.c.jsonl",
   "cost-of-intelligence-2026-09-23/largecard-reading-0/rows/b1-mistral-small3.2-cd/b1-mistral-small3.2-cd.c.jsonl",
   "deciders-von-laya-2026-10-01/PREREG-von-laya.md",
   "deciders-von-laya-2026-10-01/TABLES-speed-von-laya.json",
   "deciders-von-laya-2026-10-01/instrument/chains_pair.py",
   "deciders-von-laya-2026-10-01/instrument/fence_probe.py",
   "deciders-von-laya-2026-10-01/instrument/fetch_weights.py",
   "deciders-von-laya-2026-10-01/instrument/laya.lock.txt",
   "deciders-von-laya-2026-10-01/instrument/pins.json",
   "deciders-von-laya-2026-10-01/instrument/run_deciders.py",
   "deciders-von-laya-2026-10-01/instrument/sandbox_deciders.sh",
   "deciders-von-laya-2026-10-01/instrument/serve_and_run.sh",
   "deciders-von-laya-2026-10-01/instrument/speed/guard_speed.py",
   "deciders-von-laya-2026-10-01/instrument/speed/guard_unit.sh",
   "deciders-von-laya-2026-10-01/instrument/speed/lock/laya-cu130.constraints.txt",
   "deciders-von-laya-2026-10-01/instrument/speed/lock/laya-cu130.lock.txt",
   "deciders-von-laya-2026-10-01/instrument/speed/lock/laya.in",
   "deciders-von-laya-2026-10-01/instrument/speed/lock/lock-diff.json",
   "deciders-von-laya-2026-10-01/instrument/speed/lock/von-cu130.constraints.txt",
   "deciders-von-laya-2026-10-01/instrument/speed/lock/von-cu130.lock.txt",
   "deciders-von-laya-2026-10-01/instrument/speed/lock/von.in",
   "deciders-von-laya-2026-10-01/instrument/speed/lock_gate.sh",
   "deciders-von-laya-2026-10-01/instrument/speed/lock_release.sh",
   "deciders-von-laya-2026-10-01/instrument/speed/lock_take.sh",
   "deciders-von-laya-2026-10-01/instrument/speed/sandbox_speed.sh",
   "deciders-von-laya-2026-10-01/instrument/speed/serve_speed.sh",
   "deciders-von-laya-2026-10-01/instrument/speed/speed_deciders.py",
   "deciders-von-laya-2026-10-01/instrument/speed/tables_speed.py",
   "deciders-von-laya-2026-10-01/instrument/speed/unit_speed.sh",
   "deciders-von-laya-2026-10-01/instrument/von.lock.txt",
   "deciders-von-laya-2026-10-01/receipts-speed/card-before-speed-laya-cpu-timed-20261001T093833Z.json",
   "deciders-von-laya-2026-10-01/receipts-speed/card-before-speed-laya-gpu-timed-20261001T094233Z.json",
   "deciders-von-laya-2026-10-01/receipts-speed/card-before-speed-von-cpu-timed-20261001T093500Z.json",
   "deciders-von-laya-2026-10-01/receipts-speed/card-before-speed-von-gpu-timed-20261001T094136Z.json",
   "deciders-von-laya-2026-10-01/receipts-speed/dv5-push.txt",
   "deciders-von-laya-2026-10-01/receipts-speed/fence-speed-laya-cpu-timed-20261001T093833Z.json",
   "deciders-von-laya-2026-10-01/receipts-speed/fence-speed-laya-gpu-timed-20261001T094233Z.json",
   "deciders-von-laya-2026-10-01/receipts-speed/fence-speed-von-cpu-timed-20261001T093500Z.json",
   "deciders-von-laya-2026-10-01/receipts-speed/fence-speed-von-gpu-timed-20261001T094136Z.json",
   "deciders-von-laya-2026-10-01/receipts-speed/fence-stage-speed-laya-cpu-20261001T092152Z.json",
   "deciders-von-laya-2026-10-01/receipts-speed/fence-stage-speed-laya-gpu-20261001T092152Z.json",
   "deciders-von-laya-2026-10-01/receipts-speed/fence-stage-speed-von-cpu-20261001T092152Z.json",
   "deciders-von-laya-2026-10-01/receipts-speed/fence-stage-speed-von-gpu-20261001T092152Z.json",
   "deciders-von-laya-2026-10-01/receipts-speed/guard-timed.log",
   "deciders-von-laya-2026-10-01/receipts-speed/lock.txt",
   "deciders-von-laya-2026-10-01/receipts-speed/serve-speed-laya-cpu-timed-20261001T093833Z.log",
   "deciders-von-laya-2026-10-01/receipts-speed/serve-speed-laya-gpu-timed-20261001T094233Z.log",
   "deciders-von-laya-2026-10-01/receipts-speed/serve-speed-von-cpu-timed-20261001T093500Z.log",
   "deciders-von-laya-2026-10-01/receipts-speed/serve-speed-von-gpu-timed-20261001T094136Z.log",
   "deciders-von-laya-2026-10-01/receipts-speed/timed-guard.jsonl",
   "deciders-von-laya-2026-10-01/receipts-speed/timed-speed-laya-cpu-posture.txt",
   "deciders-von-laya-2026-10-01/receipts-speed/timed-speed-laya-cpu.log",
   "deciders-von-laya-2026-10-01/receipts-speed/timed-speed-laya-gpu-posture.txt",
   "deciders-von-laya-2026-10-01/receipts-speed/timed-speed-laya-gpu.log",
   "deciders-von-laya-2026-10-01/receipts-speed/timed-speed-von-cpu-posture.txt",
   "deciders-von-laya-2026-10-01/receipts-speed/timed-speed-von-cpu.log",
   "deciders-von-laya-2026-10-01/receipts-speed/timed-speed-von-gpu-posture.txt",
   "deciders-von-laya-2026-10-01/receipts-speed/timed-speed-von-gpu.log",
   "deciders-von-laya-2026-10-01/receipts-speed/timed-telemetry.jsonl",
   "deciders-von-laya-2026-10-01/rows-speed/speed-laya-cpu.kc.rep1.jsonl",
   "deciders-von-laya-2026-10-01/rows-speed/speed-laya-cpu.kc.rep2.jsonl",
   "deciders-von-laya-2026-10-01/rows-speed/speed-laya-cpu.kc.rep3.jsonl",
   "deciders-von-laya-2026-10-01/rows-speed/speed-laya-cpu.timed.speed.json",
   "deciders-von-laya-2026-10-01/rows-speed/speed-laya-gpu.kc.rep1.jsonl",
   "deciders-von-laya-2026-10-01/rows-speed/speed-laya-gpu.kc.rep2.jsonl",
   "deciders-von-laya-2026-10-01/rows-speed/speed-laya-gpu.kc.rep3.jsonl",
   "deciders-von-laya-2026-10-01/rows-speed/speed-laya-gpu.timed.speed.json",
   "deciders-von-laya-2026-10-01/rows-speed/speed-von-cpu.kc.rep1.jsonl",
   "deciders-von-laya-2026-10-01/rows-speed/speed-von-cpu.kc.rep2.jsonl",
   "deciders-von-laya-2026-10-01/rows-speed/speed-von-cpu.kc.rep3.jsonl",
   "deciders-von-laya-2026-10-01/rows-speed/speed-von-cpu.timed.speed.json",
   "deciders-von-laya-2026-10-01/rows-speed/speed-von-gpu.kc.rep1.jsonl",
   "deciders-von-laya-2026-10-01/rows-speed/speed-von-gpu.kc.rep2.jsonl",
   "deciders-von-laya-2026-10-01/rows-speed/speed-von-gpu.kc.rep3.jsonl",
   "deciders-von-laya-2026-10-01/rows-speed/speed-von-gpu.timed.speed.json",
   "deciders-von-laya-2026-10-01/rows/laya-cpu.ka.jsonl",
   "deciders-von-laya-2026-10-01/rows/laya-cpu.ka.summary.json",
   "deciders-von-laya-2026-10-01/rows/laya-cpu.kc.jsonl",
   "deciders-von-laya-2026-10-01/rows/laya-cpu.kc.summary.json",
   "deciders-von-laya-2026-10-01/rows/laya-cpu.kperm.jsonl",
   "deciders-von-laya-2026-10-01/rows/laya-cpu.kperm.summary.json",
   "deciders-von-laya-2026-10-01/rows/von-cpu-nochains.ka.jsonl",
   "deciders-von-laya-2026-10-01/rows/von-cpu-nochains.ka.summary.json",
   "deciders-von-laya-2026-10-01/rows/von-cpu.ka.jsonl",
   "deciders-von-laya-2026-10-01/rows/von-cpu.ka.summary.json",
   "deciders-von-laya-2026-10-01/rows/von-cpu.kc.jsonl",
   "deciders-von-laya-2026-10-01/rows/von-cpu.kc.summary.json",
   "deciders-von-laya-2026-10-01/rows/von-cpu.kperm.jsonl",
   "deciders-von-laya-2026-10-01/rows/von-cpu.kperm.summary.json",
   "deem-2026-09-27/PREREG-v0.md",
   "deem-2026-09-27/rows/d0-cpu-0.8.s0.P.rep1.jsonl",
   "deem-2026-09-27/rows/d0-cpu-0.8.s0.P.rep1.pass.json",
   "deem-2026-09-27/rows/d0-cpu-0.8.s0.P.rep2.jsonl",
   "deem-2026-09-27/rows/d0-cpu-0.8.s0.P.rep2.pass.json",
   "deem-2026-09-27/rows/d0-cpu-0.8.s0.S.rep1.jsonl",
   "deem-2026-09-27/rows/d0-cpu-0.8.s0.S.rep1.pass.json",
   "deem-2026-09-27/t1-tune/PREREG-T1.md",
   "deem-2026-09-27/t1-tune/baseline_nb.py",
   "deem-2026-09-27/t1-tune/kit/fold1-ancients.json",
   "deem-2026-09-27/t1-tune/kit/fold2-cross.json",
   "deem-2026-09-27/t1-tune/kit/fold3-operator.json",
   "deem-2026-09-27/t1-tune/kit/fold4-span.json",
   "deem-2026-09-27/t1-tune/kit/folds.json",
   "deem-2026-09-27/t1-tune/kit/t1-pool.json",
   "deem-2026-09-27/t1-tune/kit/t1-pool.manifest.json",
   "deem-2026-09-27/t1-tune/rows/t1-baseline-nb.fold1.P.rep1.jsonl",
   "deem-2026-09-27/t1-tune/rows/t1-baseline-nb.fold2.P.rep1.jsonl",
   "deem-2026-09-27/t1-tune/rows/t1-baseline-nb.fold3.P.rep1.jsonl",
   "deem-2026-09-27/t1-tune/rows/t1-baseline-nb.fold4.P.rep1.jsonl",
   "deem-2026-09-27/t1-tune/rows/t1-baseline-nb.s0.P.rep1.jsonl",
   "deem-2026-09-27/t1-tune/rows/t1-null-20260927.s0.P.rep1.jsonl",
   "deem-2026-09-27/t1-tune/rows/t1-null-20260927.s0.P.rep1.pass.json",
   "deem-2026-09-27/t1-tune/rows/t1-null-20260928.s0.P.rep1.jsonl",
   "deem-2026-09-27/t1-tune/rows/t1-null-20260928.s0.P.rep1.pass.json",
   "deem-2026-09-27/t1-tune/rows/t1-null-20260929.s0.P.rep1.jsonl",
   "deem-2026-09-27/t1-tune/rows/t1-null-20260929.s0.P.rep1.pass.json",
   "deem-2026-09-27/t1-tune/t1common.py",
   "doorman-planted-2026-09-13/PRE-REGISTRATION.md",
   "imajev-2026-09-28/PREREG-imajev.md",
   "imajev-2026-09-28/rows/imajev-4b-cpu.kc.jsonl",
   "imajev-2026-09-28/rows/imajev-4b-cpu.kc.summary.json",
   "jev-2026-09-21/ARM13-OJ-GGUF.md",
   "jev-2026-09-21/receipts/oj-gguf/run/counted-run.json",
   "jev-2026-09-21/rows-oj-gguf/openjev-q4km-generate.c.idle.csv",
   "jev-2026-09-21/rows-oj-gguf/openjev-q4km-generate.c.jsonl",
   "jev-2026-09-21/rows-oj-gguf/openjev-q4km-generate.c.report.json",
   "jev-2026-09-21/rows-oj-gguf/openjev-q4km-generate.c.watts.csv",
   "jev-2026-09-21/rows-oj-gguf/openjev-q4km-readout.c.idle.csv",
   "jev-2026-09-21/rows-oj-gguf/openjev-q4km-readout.c.jsonl",
   "jev-2026-09-21/rows-oj-gguf/openjev-q4km-readout.c.report.json",
   "jev-2026-09-21/rows-oj-gguf/openjev-q4km-readout.c.watts.csv",
   "jev-2026-09-21/rows-oj-gguf/openjev-q4km-readout.creorder.idle.csv",
   "jev-2026-09-21/rows-oj-gguf/openjev-q4km-readout.creorder.jsonl",
   "jev-2026-09-21/rows-oj-gguf/openjev-q4km-readout.creorder.report.json",
   "jev-2026-09-21/rows-oj-gguf/openjev-q4km-readout.creorder.watts.csv",
   "lev-2026-09-28/PREREG-lev.md",
   "lev-2026-09-28/l1/rows/lev-4b-cpu.kc.jsonl",
   "lev-2026-09-28/receipts/counted-run.json",
   "lev-2026-09-28/receipts/counted-telemetry.jsonl",
   "lev-2026-09-28/rows/lev-4b.ka.idle.csv",
   "lev-2026-09-28/rows/lev-4b.ka.jsonl",
   "lev-2026-09-28/rows/lev-4b.ka.report.json",
   "lev-2026-09-28/rows/lev-4b.ka.watts.csv",
   "lev-2026-09-28/rows/lev-4b.kc.idle.csv",
   "lev-2026-09-28/rows/lev-4b.kc.jsonl",
   "lev-2026-09-28/rows/lev-4b.kc.report.json",
   "lev-2026-09-28/rows/lev-4b.kc.watts.csv",
   "lev-2026-09-28/rows/lev-4b.kperm.idle.csv",
   "lev-2026-09-28/rows/lev-4b.kperm.jsonl",
   "lev-2026-09-28/rows/lev-4b.kperm.report.json",
   "lev-2026-09-28/rows/lev-4b.kperm.watts.csv",
   "opendecider-2026-09-28/PREREG-O1.md",
   "opendecider-2026-09-28/PREREG-O2.md",
   "opendecider-2026-09-28/receipts/counted-telemetry.jsonl",
   "opendecider-2026-09-28/rows/o1-cpu-nano.s0.P.rep1.jsonl",
   "opendecider-2026-09-28/rows/o1-cpu-nano.s0.P.rep1.k1-5.jsonl",
   "opendecider-2026-09-28/rows/o1-cpu-nano.s0.P.rep1.k1-5.pass.json",
   "opendecider-2026-09-28/rows/o1-cpu-nano.s0.P.rep1.k1-5.worker.log",
   "opendecider-2026-09-28/rows/o1-cpu-nano.s0.P.rep1.pass.json",
   "opendecider-2026-09-28/rows/o1-cpu-nano.s0.P.rep1.worker.log",
   "opendecider-2026-09-28/rows/o1-cpu-nano.s0.P.rep2.jsonl",
   "opendecider-2026-09-28/rows/o1-cpu-nano.s0.P.rep2.pass.json",
   "opendecider-2026-09-28/rows/o1-cpu-nano.s0.P.rep2.worker.log",
   "opendecider-2026-09-28/rows/o1-cpu-nano.s0.S.rep1.jsonl",
   "opendecider-2026-09-28/rows/o1-cpu-nano.s0.S.rep1.pass.json",
   "opendecider-2026-09-28/rows/o1-cpu-nano.s0.S.rep1.worker.log",
   "opendecider-2026-09-28/rows/o2-gpu-small.s0.P.rep1.jsonl",
   "opendecider-2026-09-28/rows/o2-gpu-small.s0.P.rep1.k1-5.jsonl",
   "opendecider-2026-09-28/rows/o2-gpu-small.s0.P.rep1.k1-5.pass.json",
   "opendecider-2026-09-28/rows/o2-gpu-small.s0.P.rep1.k1-5.worker.log",
   "opendecider-2026-09-28/rows/o2-gpu-small.s0.P.rep1.pass.json",
   "opendecider-2026-09-28/rows/o2-gpu-small.s0.P.rep1.worker.log",
   "opendecider-2026-09-28/rows/o2-gpu-small.s0.P.rep2.jsonl",
   "opendecider-2026-09-28/rows/o2-gpu-small.s0.P.rep2.pass.json",
   "opendecider-2026-09-28/rows/o2-gpu-small.s0.P.rep2.worker.log",
   "opendecider-2026-09-28/rows/o2-gpu-small.s0.S.rep1.jsonl",
   "opendecider-2026-09-28/rows/o2-gpu-small.s0.S.rep1.pass.json",
   "opendecider-2026-09-28/rows/o2-gpu-small.s0.S.rep1.worker.log",
   "opendecider-2026-09-28/rows/o2c-gpu-qwen3-4b-base.s0.P.rep1.jsonl",
   "opendecider-2026-09-28/rows/o2c-gpu-qwen3-4b-base.s0.P.rep1.k1-5.jsonl",
   "opendecider-2026-09-28/rows/o2c-gpu-qwen3-4b-base.s0.P.rep1.k1-5.pass.json",
   "opendecider-2026-09-28/rows/o2c-gpu-qwen3-4b-base.s0.P.rep1.k1-5.worker.log",
   "opendecider-2026-09-28/rows/o2c-gpu-qwen3-4b-base.s0.P.rep1.pass.json",
   "opendecider-2026-09-28/rows/o2c-gpu-qwen3-4b-base.s0.P.rep1.worker.log",
   "opendecider-2026-09-28/rows/o2c-gpu-qwen3-4b-base.s0.P.rep2.jsonl",
   "opendecider-2026-09-28/rows/o2c-gpu-qwen3-4b-base.s0.P.rep2.pass.json",
   "opendecider-2026-09-28/rows/o2c-gpu-qwen3-4b-base.s0.P.rep2.worker.log",
   "opendecider-2026-09-28/rows/o2c-gpu-qwen3-4b-base.s0.S.rep1.jsonl",
   "opendecider-2026-09-28/rows/o2c-gpu-qwen3-4b-base.s0.S.rep1.pass.json",
   "opendecider-2026-09-28/rows/o2c-gpu-qwen3-4b-base.s0.S.rep1.worker.log",
   "reads-von-laya.tsv",
   "recount-v2.json",
   "recount-von-laya.json",
   "recount_on_the_kit.py",
   "recount_v2.py",
   "recount_von_laya.py",
   "render_evidence_v2.py",
   "work-v2/reads/manifest.tsv",
   "work-v4/stat/checks.py",
   "work-v4/stat/cluster.py",
   "work-v4/stat/pairs.py"
  ]
 },
 "seat_class": {
  "writer": "gemma-class",
  "judge": null,
  "writer_runtime": "vllm"
 },
 "bench": {
  "bullets_written": 3,
  "bullets_kept": 1,
  "digests_written": 13,
  "digests_kept": 8,
  "chips_written": 8,
  "chips_kept": 6,
  "chips_grounded": 0,
  "chips_published": 0,
  "dropped_by": {
   "new_noun": 0,
   "figure": 3,
   "length": 0,
   "cite": 0,
   "judge": 0,
   "quote": 0,
   "directive": 0,
   "redaction": 0,
   "profanity": 0,
   "identifier": 0,
   "bare_figure": 0,
   "house_voice": 0,
   "hand-edit": 6
  },
  "new_noun_tokens_checked": [
   "0",
   "0.000",
   "1.3",
   "101",
   "108",
   "12",
   "14",
   "16-bit",
   "18",
   "187-token",
   "2",
   "2026-09-15",
   "24",
   "26B-A4B",
   "27",
   "27B",
   "29",
   "3.2",
   "30",
   "300",
   "3090",
   "32",
   "35",
   "36",
   "39",
   "4",
   "4's",
   "4-bit",
   "40",
   "4B",
   "512",
   "52",
   "53",
   "590",
   "66",
   "67",
   "68",
   "72",
   "73",
   "8",
   "8-bit",
   "9B",
   "APUS-OpenJev-v1",
   "Apache-2.0",
   "Bayes",
   "Deem",
   "GB",
   "GGUF",
   "Gemma",
   "Jev",
   "JevBench",
   "Kev-9B's",
   "Laya",
   "Lev",
   "Lev-4B",
   "Mistral",
   "OpenJev",
   "Python",
   "RTX",
   "Small",
   "T1",
   "VOID",
   "Von",
   "W",
   "deem-0.8-v1",
   "fp32",
   "imajev-4B",
   "lev-4b",
   "v1.4.2.2"
  ],
  "new_noun_tokens_withheld": 0,
  "figure_definition": "v3",
  "figure_boundary": "guarded",
  "drops": [
   {
    "kind": "bullet",
    "reason": "figure",
    "item": "one model's count on 108 questions is uncertain by about nine",
    "detail": "terminal token not verbatim in the cited span"
   },
   {
    "kind": "chip",
    "reason": "figure",
    "item": "The word counter, a multinomial naive Bayes, correctly identified 68 of 108 lines. With an additional rule to cross off any guest the line names, it identified 73.",
    "detail": "figures absent from cited spans: ['73']"
   },
   {
    "kind": "chip",
    "reason": "figure",
    "item": "Lev-4B took all 53 long-page questions sent, which included pages up to 14,590 tokens, and answered 52 correctly.",
    "detail": "figures absent from cited spans: ['14,590']"
   },
   {
    "kind": "bullet",
    "reason": "hand-edit",
    "item": null,
    "item_chars": 68,
    "text_sha": "fe13ca5dee5d",
    "figure": "8,187",
    "withheld": "one bullet removed from this pack by hand",
    "removed_utc": "2026-10-01",
    "why": "the line names the page APUS-OpenJev-v1 9B at high peaked on, 8,187 tokens long, and drops what peaked and where: the section's sentence is about its memory on card B, a used RTX 3090, which peaked at 24,126 MiB of the card's 24,576, level with the total PyTorch reports. Without the memory and the card the line can read as the model doing its best on that page, and its last words do not parse",
    "detail": "removed from this pack by hand on 2026-10-01 (UTC), after the writer ran and after the gates had counted: the line names the page APUS-OpenJev-v1 9B at high peaked on, 8,187 tokens long, and drops what peaked and where: the section's sentence is about its memory on card B, a used RTX 3090, which peaked at 24,126 MiB of the card's 24,576, level with the total PyTorch reports. Without the memory and the card the line can read as the model doing its best on that page, and its last words do not parse. Not a gate drop, and recorded here so the ledger adds up."
   },
   {
    "kind": "digest",
    "reason": "hand-edit",
    "item": null,
    "item_chars": 341,
    "text_sha": "848ced8f35f2",
    "section": "which-one-for-which-job",
    "withheld": "one section digest removed from this pack by hand",
    "removed_utc": "2026-10-01",
    "why": "it says APUS-OpenJev-v1 9B at high is used for screening what visitors type, where the section names it the pick among Apache-2.0 models, the only one to pass one pre-registered test of 48 lines, and the page says nothing here puts a model in front of a visitor (our door, when we tested it, was an 8-bit Gemma 4 26B-A4B); it also sets models that excel at yes-or-no calls against others that handle three-way calls, where the section says the yes-or-no set's ceiling cannot rank its eight and that on a three-way call none stood out",
    "detail": "removed from this pack by hand on 2026-10-01 (UTC), after the writer ran and after the gates had counted: it says APUS-OpenJev-v1 9B at high is used for screening what visitors type, where the section names it the pick among Apache-2.0 models, the only one to pass one pre-registered test of 48 lines, and the page says nothing here puts a model in front of a visitor (our door, when we tested it, was an 8-bit Gemma 4 26B-A4B); it also sets models that excel at yes-or-no calls against others that handle three-way calls, where the section says the yes-or-no set's ceiling cannot rank its eight and that on a three-way call none stood out. Not a gate drop, and recorded here so the ledger adds up."
   },
   {
    "kind": "digest",
    "reason": "hand-edit",
    "item": null,
    "item_chars": 261,
    "text_sha": "2d214dda778a",
    "section": "the-easier-questions",
    "withheld": "one section digest removed from this pack by hand",
    "removed_utc": "2026-10-01",
    "why": "it says these three questions separated the first fourteen models, where the section says they separated them less, or only by how much text each accepts; and it drops the as-shipped condition on Laya's limit and puts the 512 tokens on the page, where the section says Laya, as shipped, reads at most 512 tokens, the question and its options first, so 443 to 465 tokens of each page, and the page did not try Laya's long-document recipe",
    "detail": "removed from this pack by hand on 2026-10-01 (UTC), after the writer ran and after the gates had counted: it says these three questions separated the first fourteen models, where the section says they separated them less, or only by how much text each accepts; and it drops the as-shipped condition on Laya's limit and puts the 512 tokens on the page, where the section says Laya, as shipped, reads at most 512 tokens, the question and its options first, so 443 to 465 tokens of each page, and the page did not try Laya's long-document recipe. Not a gate drop, and recorded here so the ledger adds up."
   },
   {
    "kind": "digest",
    "reason": "hand-edit",
    "item": null,
    "item_chars": 233,
    "text_sha": "ef18a5c8aa04",
    "section": "what-it-takes-to-run",
    "withheld": "one section digest removed from this pack by hand",
    "removed_utc": "2026-10-01",
    "why": "it says small deciders were tested on a 24 GB RTX 3090, where the section says every small decider we put on a graphics card ran on one, and OpenDecider nano and deem-0.8-v1 ran only on a laptop's processor; and it sets Von 1.3 and Laya apart as running on a desktop's eight-core processor, where the section times both on card B, a used RTX 3090, as well as on that processor",
    "detail": "removed from this pack by hand on 2026-10-01 (UTC), after the writer ran and after the gates had counted: it says small deciders were tested on a 24 GB RTX 3090, where the section says every small decider we put on a graphics card ran on one, and OpenDecider nano and deem-0.8-v1 ran only on a laptop's processor; and it sets Von 1.3 and Laya apart as running on a desktop's eight-core processor, where the section times both on card B, a used RTX 3090, as well as on that processor. Not a gate drop, and recorded here so the ledger adds up."
   },
   {
    "kind": "digest",
    "reason": "hand-edit",
    "item": null,
    "item_chars": 261,
    "text_sha": "5123c56258b6",
    "section": "what-we-did-not-measure",
    "withheld": "one section digest removed from this pack by hand",
    "removed_utc": "2026-10-01",
    "why": "it says energy per decision was not measured, where the section, headed what we did not measure or did not report, says several runs recorded it and none recounted it: measured, and not reported",
    "detail": "removed from this pack by hand on 2026-10-01 (UTC), after the writer ran and after the gates had counted: it says energy per decision was not measured, where the section, headed what we did not measure or did not report, says several runs recorded it and none recounted it: measured, and not reported. Not a gate drop, and recorded here so the ledger adds up."
   },
   {
    "kind": "digest",
    "reason": "hand-edit",
    "item": null,
    "item_chars": 174,
    "text_sha": "c4f2dcf9474b",
    "section": "where-each-figure-comes-from",
    "withheld": "one section digest removed from this pack by hand",
    "removed_utc": "2026-10-01",
    "why": "it says the section gives the machine names, where the section says the paths in the bench's record are printed with machine names rewritten; and it says the kit holds the files needed to re-derive the tables, where the page says that for the held-back sets the kit has figures, not files",
    "detail": "removed from this pack by hand on 2026-10-01 (UTC), after the writer ran and after the gates had counted: it says the section gives the machine names, where the section says the paths in the bench's record are printed with machine names rewritten; and it says the kit holds the files needed to re-derive the tables, where the page says that for the held-back sets the kit has figures, not files. Not a gate drop, and recorded here so the ledger adds up."
   }
  ]
 },
 "generated_utc": "2026-10-01T17:31:25Z"
}
