# Small open deciders, measured

*A decider is a language model that reads a question and its options and returns a probability for each. We asked fourteen open models five questions from our own products. On the hardest, the top five small deciders labelled Apache-2.0, given no examples, could not be told apart from a word counter trained on our own lines. One of them, APUS-OpenJev-v1 9B, passed a pre-registered screening test that the model then guarding our exhibit missed by three lines, letting through only kinds of line its rules did not name.*

*Published 2026-09-30 (UTC) · A small (human) team and a fleet of AI agents.*

**the short version:** A decider reads a question and its options and returns a probability for each: nothing to parse, and a number to set a hand-off threshold on, once that threshold is checked on your own cases. For a yes-or-no call about a short passage, the eight models we asked could not be told apart: each scored 39 or 40 of 40 on one such question, so choose by licence and speed. For questions about long pages, Lev-4B took all 53 we sent and got 52 right. For screening what visitors type, APUS-OpenJev-v1 9B alone of the Apache-2.0 models we asked passed our pre-registered test, refusing 32 of 36 lines written to be refused and 0 of 12 harmless ones. For sorting short texts by how they are written, train a bag of words first: ours got 68 of 108, and the top five small deciders labelled Apache-2.0, at 66 to 72, could not be told apart from it. OpenJev 27B scored highest, 101, and its 4-bit build sits whole on one used RTX 3090, but its weights are licensed for non-commercial use, and commercial use starts with a discussion on its repository.

7,028 words · about 32 minutes (at 220 words/min) · 6 tables · data kit: yes

https://research.strata2signal.com/small-open-deciders-measured/

---

*[strata→signal](https://strata2signal.com) is a small workshop that runs its own machines and writes up what it measures. We ran everything here from 2026-09-13 to 2026-09-28 (UTC) and recounted every count from the runs' own files, except the few named in [how to check our work](#how-to-check-our-work); figures quoted from others say so. We never ran TypeSafe's hosted Jev: every model here is open weights, and OpenJev's and Lev's cards say they are not affiliated with TypeSafe.*

## What a decider is {#what-a-decider-is}

Some jobs inside a product are choices, not writing: should the table hear this line? Does this page answer this question? For a choice, a paragraph is a liability: something has to parse it, and it carries no honest number for how sure the model was.

A decider reads the question and its labelled options and turns one score per option into probabilities that sum to one. Nothing is generated, so a product can act above a threshold and hand the unsure cases to a person, once the threshold is checked on the product's own cases, as the Lev, imajev and APUS-OpenJev-v1 cards advise. This page measures no calibration; [Reading the answer instead of writing it](https://research.strata2signal.com/reading-the-answer/) measured it for OpenJev, Kev and Gemma 4, and tells how this began: TypeSafe announced Jev, a hosted decider, on 2026-09-15, and open models that answer the same way followed.

## Which one for which job {#which-one-for-which-job}

Each pick rests on one or two of our questions, asked zero-shot of each model as shipped, at the revision we benched, and scored on its top choice: a start for your own measuring, not a verdict on any model. One model's count on 108 questions is uncertain by about nine percentage points either way, and a gap between two is less certain still, so read the paired tests, not the gap.

- **A yes-or-no call about a short passage: any of the eight we asked.** Each scored 39 or 40 of 40 on whether a rulebook excerpt answers a question: the set's ceiling, so it cannot rank them, and 40 of 40 still allows about one miss in eleven. Choose by [licence](#the-licences) and [speed](#what-it-takes-to-run). See [the easier questions](#the-easier-questions).
- **A three-way call: none stood out.** On whether an answer is grounded in its page, scored strict, the small ones got 89 to 95 of 119 and OpenJev 98 (another layout; we ran no test on these gaps). Check any of them on your own lines.
- **Questions about long pages: Lev-4B.** It took all 53 we sent, on pages up to 14,590 tokens, and got 52 right. Kev and APUS-OpenJev-v1 got right all they took, but refuse, rather than truncate, anything over about 8,000 tokens.
- **Screening what visitors type, under Apache-2.0: APUS-OpenJev-v1 9B with `effort` at `high`**, the only Apache-2.0 model to pass our pre-registered test: 32 of 36 lines written to be refused, 0 of 12 harmless ones, a median 0.251 s a line. It is one test of 48 lines, 0 of 12 still allows about one harmless line in four, its authors warn that its probabilities "are not calibrated confidence", and our review of its training data is not done. It is built on Qwen3.5, not on OpenJev 27B. Not the 4B, which refused 9 of the 12 harmless lines. We asked no dedicated screening model. See [the doorman](#the-doorman).
- **Sorting short texts by how they are written: train a bag of words on your own labels first** (our six-way: which of six characters said a line, a matter of style). Ours, from 227 lines, got 68 of 108; the top five small deciders labelled Apache-2.0 got 66 to 72 (65.3 to 70.0 over six option orders, for the four we shuffled), none told apart from it by a paired test (p 0.60 to 1.0). The chat model Gemma 4 26B-A4B got 88 but wrote 51 of the lines; on the other 57 it got 49 to the word counter's 36 (uncorrected p 0.002, in a split made after the rows were seen). See [a word counter kept up](#a-word-counter-kept-up).
- **Options in no natural order: Lev-4B or Kev-9B.** Over the six orders Kev's server generates, their answers moved on 27 and 33 of 108 lines, too close for a paired test to separate; APUS-OpenJev-v1 9B's moved on 41, which a paired test does separate from Lev's. Ask your own questions in several orders and count how often the answer moves.
- **Hardware:** each small decider we put on a graphics card ran in bf16 on one used 24 GB RTX 3090, and we tried no smaller card. The 4B ones peaked at 8,596 to 22,736 MiB; APUS-OpenJev-v1 9B filled the card on an 8,187-token page.
- **No graphics card:** on a laptop's processor, Lev-4B took a median 8.465 s a line and imajev-4B 18.134 s; Lev's card says it "Needs a GPU for real-time use", imajev's names "MLX on a Mac or PyTorch on one GPU". Kev and APUS-OpenJev-v1 were not tried there.
- **Not for a choice like our six-way, as shipped:** OpenDecider small (46, level with its own base), OpenDecider nano (13; it refused all 48 screening lines) and deem-0.8-v1 (5, VOID by the instrument as registered; valid only under R-1), all under Kev-9B's 72 of 108, which our pre-registrations for them named as a reference, not a pass mark.

OpenJev 27B's 101 of 108 (99 as a 4-bit build that sits whole on one used RTX 3090) is not a pick: its weights are CC BY-NC 4.0, "free for research and other non-commercial use, with attribution", and for commercial use its README asks for a discussion; ours had no reply at our last read.

## The questions we asked {#the-questions-we-asked}

**The models**, fourteen:

- **Nine small deciders:** Kev-9B and Kev-4B, Lev-4B, APUS-OpenJev-v1 9B and 4B, imajev-4B, OpenDecider small and nano, deem-0.8-v1.
- **OpenJev 27B**, at three precisions.
- **Jevify**, a third-party fine-tune of Gemma 4 26B-A4B (a LoRA merged into the weights) that answers typed questions through a Jev-compatible API, with no teacher model, its card says.
- **Two chat models our products run**, as yardsticks: Gemma 4 26B-A4B and Mistral Small 3.2 24B.
- **A control**, Qwen3-4B-Instruct-2507, OpenDecider small's base, asked through OpenDecider's prompt without the adapter.

**How each was asked.** Kev, Lev, APUS-OpenJev-v1 and imajev take one request shape (a context the format calls the "state", a question, its options), and wherever Kev-9B also answered, their rows record its request hash. Kev reads it through a small pointer head, Lev through a short code per option in two orders, imajev through a 256-code readout over four, APUS-OpenJev-v1 by scoring each option's label. OpenJev, jevify, Gemma 4 and Mistral Small 3.2 got our first Jev bench's layout (OpenJev's own helper's, on its older untargeted path), OpenDecider and Deem their own. A *readout* reads each option letter's probability where the answer would begin; a *written letter* is parsed from what the model writes.

**The six-way.** Our exhibit, the long table, has language models play six long-dead guests at a dinner and saves every line with its speaker. We froze 108 lines and ask of each: which guest said it? No line names its speaker, and chance is 18. The guests share a table and a topic, so the answer is in how each one talks: our closest question to sorting short texts by how they are written. The lines and answers have been public since 2026-09-23 in [the six-way file of the Reading the answer kit](https://research.strata2signal.com/reading-the-answer/data/kit/task_c.json). The long table's two chair models wrote them: 51 by Gemma 4, the other 57 by Mistral Small 3.2.

**Forty-eight held-out lines** from the same public wall, never in any kit and never printed, control for a model that learned the kit, though not for the wall. They run eight per guest rather than eighteen, 28 by Mistral Small 3.2 and 20 by Gemma 4.

**The word counter** is a multinomial naive Bayes written in standard-library Python, trained on 227 other lines our long table wrote, with textbook settings (single words, add-one smoothing), none chosen on the six-way. It read 68 of 108; with one more rule (cross off any guest the line names), 73, a figure we print and never gated on.

**None was built for our six-way.** At the revisions we benched, their cards name jobs from "browser action selection" (APUS-OpenJev-v1) to "routing, moderation, intent detection, triage, grading, and checking LLM output" (Lev) and judging "whether an answer is grounded" (OpenJev). Moderation and grounding are close to our doorman and judge questions, but none of the evaluation sets the cards name asks who wrote a text.

## The six-way scoreboard {#the-six-way-scoreboard}

The six-way and the held-out lines, for every model that read them.

| Model (how it was asked) | Six-way, right of 108 | Answer moved over six orders, of 108 | Held-out, right of 48 |
|---|---:|---:|---:|
| OpenJev 27B, 16-bit (Jev layout, readout) | 101 | — | — |
| OpenJev 27B, 8-bit (Jev layout, readout) | 101 | 15 | — |
| OpenJev 27B, 4-bit GGUF (Jev layout, readout) | 99 | 14 | 44 |
| Jevify, a Gemma 4 26B-A4B fine-tune, 4-bit (Jev layout, readout) | 90 | — | — |
| Gemma 4 26B-A4B, 4-bit (Jev layout, written letter) | 88 | — | — |
| Mistral Small 3.2 24B, 4-bit (Jev layout, written letter) | 73 | — | — |
| Kev-9B (shared request) | 72 | 33 | — |
| Word counter, ours (no prompt) | 68 | — | 27 |
| imajev-4B (shared request) | 68 | — | 30 |
| APUS-OpenJev-v1 4B, high (shared request) | 67 | 47 | 29 |
| APUS-OpenJev-v1 9B, high (shared request) | 66 | 41 | 31 |
| Lev-4B (shared request) | 66 | 27 | — |
| Kev-4B (shared request) | 60 | 33 | — |
| OpenDecider small (OpenDecider's prompt) | 46 | 64 | 25 |
| Qwen3-4B-Instruct-2507, no adapter (OpenDecider's prompt) | 46 | 71 | 26 |
| APUS-OpenJev-v1 9B, low (shared request) | 30 | — | 15 |
| APUS-OpenJev-v1 4B, low (shared request) | 25 | — | 12 |
| OpenDecider nano (OpenDecider's prompt) | 13 | 96 | 5 |
| deem-0.8-v1, as shipped (Deem's prompt; VOID by the instrument as registered; valid only under R-1) | 5 | — | 4 |

*How to read it.*

- **Rows with the same words in brackets share a prompt layout;** across layouts, read neighbours, not a ranking. A dash is [not measured](#what-we-did-not-measure). Wilson 95 per cent intervals are wide: 87.2 to 96.8 per cent for OpenJev's 101, 57.3 to 74.8 for Kev-9B's 72, 53.6 to 71.5 for the word counter's 68. They, and every paired test here, treat the 108 lines as independent, though they come from 16 dinners; resampling whole dinners changes no verdict against the word counter but widens some intervals.
- **Scored strict:** a written reply with no readable letter is wrong. Gemma 4's 88 is 89 under our first Jev bench's lenient rule and 87 as a readout (floored: on 101 of 108 lines some letter fell outside the 20 scores returned). Mistral Small 3.2 read 74 in another run, on our workstation card.
- **Six orders.** The six-way column is the kit's option order, the registered figure; the moved count has no bar. Kev, Lev and APUS-OpenJev-v1 used the six orders Kev's server generates, the rest the kit's plus five from a fixed seed: compare within one set. Means over the six, computed after the rows were seen and not registered: OpenJev 98.7 at 8 bits and 98.2 at 4; Kev-9B and APUS-OpenJev-v1 9B 70.0; APUS-OpenJev-v1 4B 65.7; Lev-4B 65.3; Kev-4B 61.0; Qwen3-4B 53.8; OpenDecider small 48.7 and nano 11.2. The kit's order is APUS-OpenJev-v1 9B's lowest and one of Kev-9B's highest.
- **Exact ties.** APUS-OpenJev-v1 4B's 67 is 65 to 68 over four exactly tied lines; the 9B's 66 is 65 to 67, and its held-out 31 is 31 to 32. Kev's server returns probabilities to two decimals, so its ties are not read.
- **VOID by the instrument as registered; valid only under R-1.** Our registered script for deem-0.8-v1 had no answer for an exact tie; R-1, Deem's own rule (the first option wins), was ratified after the rows were seen and is labelled so. The six-way figure moves under no tie rule; the held-out one reads 4 to 5.
- **APUS-OpenJev-v1 at `low`** stops halfway through the network, a setting its authors offer "for a smaller compute budget" (their 4B: 82.50 per cent at `high`, 76.25 at `low`, on their 80 questions); ours drops in the same direction.

### Did anyone see the answers? {#did-anyone-see-the-answers}

The six-way was on a public wall from 2026-09-19 and in the published kit from 2026-09-23; the held-out lines test the kit, not the wall. For every model that read both, the six-way share minus the held-out share (positive where the kit read better) has a Newcombe 95 per cent interval spanning zero, from +1.6 points (APUS-OpenJev-v1 4B at `high` and OpenDecider nano) to Qwen3-4B's −11.6 (−27.5 to +5.2). The word counter, which cannot have learned the kit, reads +6.7 (−9.4 to +23.1). Intervals about 32 points wide, as for the small deciders near its level, would show only a kit advantage of roughly 15 points or more: "no signal at this size", not "clean".

Gemma 4 and Mistral Small 3.2 were released before these lines existed (2026-03-31 and 2025-06-20), though they wrote them. The jevify, Kev and OpenJev weights we benched (2026-09-20 and 09-21) have no held-out reading but predate the kit, and OpenJev's 4-bit build of those weights read +0.0 (−8.4 to +11.9); Lev (2026-09-24) postdates the kit, with no such check. All but Gemma 4, Mistral Small 3.2 and the Qwen3-4B control (2025-09-17) postdate the wall, which no check of ours covers.

## A word counter kept up {#a-word-counter-kept-up}

Kev-9B's 72 against our word counter's 68 is four lines net: Kev-9B got 18 right that the word counter missed, and the word counter 14 that Kev-9B missed, well within chance by an exact McNemar test (p 0.60). That does not prove a tie: the paired 95 per cent interval for Kev-9B's edge runs from about 7 lines behind to 15 ahead. Nor is it an artefact of the option order: in each of the six orders Kev's server generates, none of the four top-five deciders we shuffled is told apart from the word counter (p 0.27 at the closest).

| Model | Right only here | Right only in the word counter | Exact p |
|---|---:|---:|---:|
| OpenJev 27B, 16-bit, readout | 37 | 4 | 1e-07 |
| OpenJev 27B, 4-bit GGUF, written letter | 36 | 5 | 7.8e-07 |
| Jevify, 4-bit, readout | 29 | 7 | 0.00031 |
| Gemma 4 26B-A4B, 4-bit, written letter | 28 | 8 | 0.0012 |
| Mistral Small 3.2 24B, 4-bit, written letter | 21 | 16 | 0.51 |
| Kev-9B | 18 | 14 | 0.60 |
| imajev-4B | 21 | 21 | 1.0 |
| APUS-OpenJev-v1 4B, high | 18 | 19 | 1.0 |
| APUS-OpenJev-v1 9B, high | 18 | 20 | 0.87 |
| Lev-4B, on a used RTX 3090 | 18 | 20 | 0.87 |
| Kev-4B | 17 | 25 | 0.28 |
| OpenDecider small | 13 | 35 | 0.0021 |
| Qwen3-4B-Instruct-2507, no adapter | 13 | 35 | 0.0021 |
| APUS-OpenJev-v1 9B, low | 15 | 53 | 4.1e-06 |
| APUS-OpenJev-v1 4B, low | 12 | 55 | 1e-07 |
| OpenDecider nano | 1 | 56 | 8e-16 |

*Two-sided exact McNemar on the lines where a model and the word counter (a naive Bayes) disagree. deem-0.8-v1 reads 2 against 65 (p 3.1e-17): VOID by the instrument as registered; valid only under R-1.*

Everything at Gemma 4's level or above is ahead of the word counter by more than chance explains, and everything below 60 behind it, even after correcting for all seventeen tests. Between them, Kev-9B, imajev-4B, both APUS-OpenJev-v1 models at `high`, Lev-4B, Kev-4B and the chat model Mistral Small 3.2 cannot be told apart from 227 labelled lines and a word count.

Who wrote a line matters too, in a split we made after the rows were seen and did not register. Gemma 4's lead over Kev-9B is larger on the 51 lines its own family wrote, 39 against 27 (p 0.0018), than on the other 57, 49 against 45 (p 0.29), though whether the two leads differ is not settled at this size. Mistral Small 3.2 shows no such edge on its own lines (32 of Gemma 4's 51, 41 of its own 57).

The top five small deciders read 24 to 28 of the Gemma-written lines and 38 to 45 of the Mistral-written ones, the word counter 32 and 36, so being level on all 108 is two gaps of opposite sign: the word counter ahead of each by 4 to 8 lines on the first, behind each by 2 to 9 on the second. None of the split's 56 uncorrected paired tests survives a correction for their number, and we do not know why it falls this way. If your texts come from one writer, measure the bar on them.

None of this rates the models in general: our answer lives in how six characters talk, the home ground of a bag of words. The point is the bar: if your decision looks like ours, a few hundred labelled examples and a textbook classifier are what a zero-shot decider must clear, and no small one we tried cleared it by more than chance explains. The word counter cannot screen a line or read a page, though.

## The doorman: screening what visitors type {#the-doorman}

A visitor to the long table may type one line for the guests. It joins the dinner's record and can reach the public wall, so a model stands at the door. Ours, when we tested it, was an 8-bit Gemma 4 26B-A4B, with written rules that name four kinds of line to refuse (contempt for a group, sexual content, a private person named, lines that break the dinner's frame) and let everything else through.

On 2026-09-13 we planted 36 lines, six of each of six kinds: the four named, plus self-harm and instructions to a living person, which the rules do not name. Beside them went 12 harmless lines of the shapes a careless filter refuses: lovers in the abstract, a dead public figure named, a trial in the guest's own year. The planted lines are hostile by design and stay private.

The gate, written before the first call: **at least 30 of the 36 planted lines refused, and at most 2 of the 12 harmless ones.** Misses in the two unnamed kinds, it said in advance, would be a gap in the rules, not a fault in the model.

**The live door failed**, 27 and 0, refusing all 24 named lines and 3 of the 12 unnamed: it did what it was told. Every decider then got the same rules word for word, the line, and two options (the table hears it, or not) in its own layout. The rules were written for a chat model that writes a verdict; a question shaped for a decider might fare differently, which we did not test.

| Model | Should-refuse lines refused, of 36 | Harmless lines refused, of 12 | The 09-13 gate | Named kinds refused, of 24 | Unnamed kinds refused, of 12 |
|---|---:|---:|---|---:|---:|
| Gemma 4 26B-A4B, 8-bit, the live doorman, a written verdict | 27 | 0 | FAIL | 24 | 3 |
| OpenJev 27B, 8-bit, readout | 33 | 0 | Pass by the rule, not registered | 24 | 9 |
| OpenJev 27B, 4-bit GGUF, readout | 32 | 0 | PASS | 24 | 8 |
| APUS-OpenJev-v1 9B, high | 32 | 0 | PASS | 23 | 9 |
| APUS-OpenJev-v1 9B, low | 32 | 1 | Pass by the rule, not registered | 23 | 9 |
| Gemma 4 26B-A4B, 8-bit, readout | 23 | 0 | Fail by the rule, not registered | 23 | 0 |
| APUS-OpenJev-v1 4B, high | 27 | 9 | FAIL | 17 | 10 |
| APUS-OpenJev-v1 4B, low | 24 | 0 | Fail by the rule, not registered | 17 | 7 |
| imajev-4B | 21 | 0 | FAIL | 20 | 1 |
| Lev-4B | 19 | 0 | FAIL | 16 | 3 |
| Kev-9B | 15 | 0 | Fail by the rule, not registered | 15 | 0 |
| Qwen3-4B-Instruct-2507, no adapter | 15 | 0 | FAIL | 11 | 4 |
| Kev-4B | 4 | 0 | Fail by the rule, not registered | 4 | 0 |
| OpenDecider small | 4 | 0 | FAIL | 4 | 0 |
| OpenDecider nano | 36 | 12 | FAIL | 24 | 12 |

*PASS or FAIL where the run's own pre-registration named the 09-13 gate before its first call; "by the rule, not registered" where we applied it afterwards. The split by kind was computed after the rows were seen and is not registered; every total equals the registered count. Breaking exact ties the other way changes no gate. deem-0.8-v1 was not asked.*

Most that missed refused too few planted lines and no harmless one, from 4 (Kev-4B, OpenDecider small) to 27 (the live door): partly the rules' gap (imajev refused 20 of the 24 named lines), partly not (Kev-9B refused 15). Two refused too much: OpenDecider nano refused all 48, its probability of refusing 0.57 to 0.78 on every line; APUS-OpenJev-v1 4B at `high` turned away 9 of the 12 harmless lines, all four trial lines among them.

**APUS-OpenJev-v1 9B at `high` is the one model here that passed and whose weights carry the Apache-2.0 text:** 32 and 0, registered, with 23 of the 24 named lines and 9 of the 12 unnamed. OpenJev's 4-bit build passed too, registered; its 8-bit build (33 and 0) and the 9B at `low` (32 and 1) pass by the rule, not registered.

The 9B's margin is thin: one planted line sits a 16-bit rounding step from a tie, and a FAIL would take three flips. Against the live door on the same 36 lines, it refused 6 the door let through, all of unnamed kinds, and let through 1 named line the door refused, a split that a paired test cannot call (p 0.13). As the one registered Apache-2.0 pass among many rows, it is the likeliest to read lower on a fresh planted set. An independent check recomputed every probability from the raw scores and traced by hand how the runtime maps yes and no, finding no inversion.

Two things come before any door of ours: our review of the 9B's training data and a test shaped like the product. The 4B is no guide to it, or the reverse: the two checkpoints were not trained for the same number of steps.

## The easier questions {#the-easier-questions}

Three questions separated the models less, or only by how much text each accepts: does a rulebook excerpt answer a question (the field exam, 40); is an answer grounded in its page, not grounded, or off it (the judge question, 119); does a page of ours answer a question (long pages)?

| Model | Field exam, right of 40 | Judge question, strict, of 119 | Judge question, any listed answer, of 119 | Long-page questions reached, of 53 | Right, of those reached |
|---|---:|---:|---:|---:|---:|
| OpenJev 27B, 8-bit, readout | 40 | 98 | 118 | 53 | 53 |
| Gemma 4 26B-A4B, 8-bit, readout | 40 | — | — | 53 | 53 |
| Kev-9B | 40 | 95 | 115 | 38 | 38 |
| APUS-OpenJev-v1 9B, high | 40 | 93 | 113 | 37 | 37 |
| imajev-4B | 40 | 92 | 112 | — | — |
| Kev-4B | 40 | 92 | 112 | 38 | 38 |
| APUS-OpenJev-v1 4B, high | 40 | 91 | 111 | 37 | 37 |
| Lev-4B | 39 | 89 | 108 | 53 | 52 |

*Strict counts the one labelled answer; the next column accepts any answer the set lists, so name the scoring when you compare. Of 63 long-page questions, ten were over our first run's 16,384-token window, the prompt limit OpenJev's card states, and left out of every main run. Kev's server refuses over 8,192 tokens, APUS-OpenJev-v1's runtime stops at 8,191, and imajev's card sets 4,096, which the pages of 48 of the 53 exceed, so it was sent none of them; anything over about 14,600 tokens is untested with Lev's server on a 24 GB card. At a 32,768-token window OpenJev's 8-bit build reached 6 of those 10 questions, all 6 right, past its card's limit.*

## The licences {#the-licences}

OpenJev's README says: "OpenJev weights are released under CC BY-NC 4.0: free for research and other non-commercial use, with attribution. For commercial use, open a discussion on this repository." We opened discussion #1 on 2026-09-22 at 01:19:18 UTC; at our last read, 2026-09-30 at 01:12:00 UTC, it had no reply.

What each model's repositories carry, as read on 2026-09-28 between 07:48 and 15:57 UTC, deem-0.8-v1's the day before:

- **OpenJev 27B, 16-bit and 8-bit.** The full CC BY-NC 4.0 text in `openjev/openjev`'s `LICENSE` at `0b6bb6e5` (we benched `5ec9e5fd`; only the README differs), and CC BY-NC 4.0 on both cards; a second file, Apache-2.0, covers the `helper/` and `serve/` code and, its `NOTICE` says, the Qwen3.8-27B base. No licence file in the 8-bit repository.
- **OpenJev 27B, 4-bit GGUF.** No licence file at `208220bc`; Apache-2.0 in the card's metadata, while the OpenJev README puts the weights under CC BY-NC 4.0. We follow the README until its authors say which applies, and treat the build as bench-only, like its source.
- **APUS-OpenJev-v1 9B and 4B.** Apache-2.0 word for word in each `LICENSE`, byte-identical to the Qwen3.5 bases'.
- **Lev-4B.** Code: Apache-2.0, word for word. Adapter: Apache-2.0, "the same license as the base model".
- **imajev-4B.** Code: Apache-2.0, sections 1 to 9 word for word. Adapter: no licence file; Apache-2.0 in its card metadata.
- **OpenDecider nano and small.** One `LICENSE`, the same bytes in the code and all six model repositories: an Apache-2.0 text whose section 6 lacks "reasonable and customary use in" and section 9 "choose to offer, and charge a fee for, acceptance of". GitHub reads it as unidentified; the cards say Apache-2.0.
- **Kev-9B and Kev-4B.** Apache-2.0 in the GitHub API and the card metadata; the text not compared word for word.
- **deem-0.8-v1.** Apache-2.0 on the card; the Apache-2.0 text in the code repository's `LICENSE`.
- **Gemma 4 26B-A4B and Mistral Small 3.2 24B.** Apache-2.0 in the card metadata; the text not compared.
- **Jevify.** `license: gemma` on both its cards, over a base whose card says Apache-2.0; not resolved here.

*We print what each repository carries and lacks, and do not say what any difference means in law.*

**Four of the many things called OpenJev:** the non-commercial `openjev/openjev`; an Apache-licensed fine-tuning framework on GitHub, `S1LV3RJ1NX/openjev`, by another author; the OpenJev organisation's GGUF builds; and APUS-OpenJev-v1, from APUS AI-LAB, whose stated comparator is TypeSafe's hosted Jev and whose bases are Qwen3.5: none of its 294 text files mentions `openjev/openjev`, CC BY-NC or "non-commercial". Name the repository when you cite one.

**What Apache-2.0 on the weights does not settle** is the training data, which we read for four models only. Lev's card warns that some of its training datasets carry "non-commercial licenses" and says to "Review them before commercial use"; four of its 26 carry CC BY-NC tags. imajev's documentation records the rights to some of its image sources as unverified or not granted. APUS-OpenJev-v1's training notes say "separate provenance and redistribution review remains necessary", and one or two of its small sources are not described: 80 records labelled "Local counterfactual", and an `oracle_only` source seen only in its quantisation sample, perhaps the same records. OpenDecider's notice lists no non-commercial set; its training pool is unpublished. Kev's card says its "datasets carry their own licenses"; we did not review those, or the other cards' sources.

## What it takes to run {#what-it-takes-to-run}

Every small decider we put on a graphics card ran in bf16 through its own Python server or package on one used 24 GB RTX 3090, never at 4 or 8 bits and never through ollama, llama.cpp or vLLM. Two different used RTX 3090s carry the rows: card A at 250 W and card B at 300 W.

| Model | Grew at load, MiB | Highest, MiB of 24,576 | Median per six-way line, s | Precision and runtime | Card |
|---|---:|---:|---:|---|---|
| OpenJev 27B, 4-bit GGUF, readout | 16,727 | 16,730 | 1.131 | Q4_K_M, ollama 0.32.13 | Card B, 300 W |
| APUS-OpenJev-v1 9B, high | 18,471 | 24,126 | 0.183 | BF16, its vendor's runtime | Card B, 300 W |
| APUS-OpenJev-v1 4B, high | 9,235 | 16,684 | 0.171 | BF16, its vendor's runtime | Card B, 300 W |
| Lev-4B, two orders per answer | 9,125 | 22,736 | 0.217 | BF16, Lev's server 0.1.1 | Card B, 300 W |
| imajev-4B, four orders per answer, timed on the held-out lines | 9,798 | 11,256 | 0.516 | BF16, its own server | Card B, 300 W |
| OpenDecider small | 8,595 | 8,596 | 0.063 | BF16, transformers | Card B, 300 W |
| Kev-9B | Not recorded | 20,278 | 0.106 | BF16, Kev's own server | Card A, 250 W |
| Kev-4B | Not recorded | 13,358 | 0.070 | BF16, Kev's own server | Card A, 250 W |

*Grew at load: used memory after the warm-ups less the reading before the load, from each run's own check; highest: the run's top telemetry reading. Both, and imajev's times (on the 48 held-out lines; its six-way ran only on a laptop), were computed after the rows were seen.*

*APUS-OpenJev-v1 9B peaked on an 8,187-token page, 450 MiB under the card's 24,576 but level with the 24,126 MiB PyTorch reports as its total, and logged no out-of-memory line; Lev on a 14,590-token page, after its allocator hit the limit once and recovered. Kev's runs recorded no reading before the load; loaded and idle, the card read 15,866 MiB with Kev-9B and 8,776 with Kev-4B. Lev, APUS-OpenJev-v1 and imajev ran without the flash-linear-attention kernels Kev's server had. Each row is its own runtime and card: divide none by another.*

**OpenJev on one card.** Its 8-bit build needed two used RTX 3090s; the Q4_K_M GGUF its organisation published on 2026-09-24, from the revision we had measured, sits whole on one at 300 W through ollama 0.32.13, with 7,846 MiB of headroom. It read 99 of 108, the same guest as the 16-bit model on 106: the registered bar, at least 97 right and 100 the same, CARRIES as a written letter and as a readout.

**Elsewhere,** at the median per six-way line: at 4 bits through ollama on card A, jevify took 1.017 s and Gemma 4 1.025 s as readouts, and Mistral Small 3.2 0.619 s writing a letter (at 300 W). On a laptop's processor, OpenDecider nano took 0.199 s and deem-0.8-v1 0.421 s.

## Three teams, one base {#three-teams-one-base}

Lev-4B, imajev-4B and APUS-OpenJev-v1 4B come from three teams, all on `Qwen/Qwen3.5-4B` at revision `851bf6e8`. On the six-way they land within two lines (66, 68 and 67); at the door they part, letting 17, 15 and 9 planted lines through and refusing 0, 0 and 9 harmless ones. OpenDecider small, on Qwen3-4B, matched its own base's 46 (6 lines each way, p 1.0); its card reports distillation lifting that base from 0.700 to 0.735 on its own 200 decisions and cutting its calibration error from 0.289 to 0.087, which we did not measure. Without the bare Qwen3.5-4B, we cannot say what each team's changes add.

## A board's rank, and what ours adds {#a-boards-rank-and-what-ours-adds}

JevBench, Benchmark Heaven's public leaderboard for Jev-style deciders, ranked imajev-4B first in the board's v1.4.2.2 release when we read it (2026-09-28, 15:48:19 UTC), on a composite of 67.4 blending four axes (intelligence, calibration, speed and cost) over 534 public and 308 sealed questions; imajev's card quotes the same rank and score. The board's accuracy rows for it, quoted and not checked, read 86.1 per cent public and 37.0 sealed. Another entry, Hopper, has the same shape (82.3 and 34.1), which fits the sealed questions simply being harder; two entries of 91 cannot say more.

imajev's card says an 8-gram check kept JevBench items out of its training, and on our control it scored about the same on lines never in a kit (62.5 per cent) as on the kit's (63.0). The board ran it with one option order, its note says; we ran its card's four.

A board's rank blends accuracy with speed and cost over questions that are not yours. On ours, never on that board, imajev-4B read 68 of 108 on the six-way, level with our word counter (21 lines each way, p 1.0), 30 of 48 held-out lines, and 21 and 0 at the door, a FAIL.

## Our fine-tune that voided itself {#our-fine-tune-that-voided-itself}

deem-0.8-v1, a 0.8-billion-parameter decider, read 5 of 108 on the six-way as shipped, below chance (VOID by the instrument as registered; valid only under R-1). Our run's notes, written after the fact, find that it most often picks a guest the line itself names, which this set rules out by construction. Small, fast on a processor and Apache-licensed, it looked like the right model to teach, so we pre-registered T1: tune it on the 227 lines our word counter learned from, with the recipe fixed before any result.

The design carried its own alarm: three controls trained the same way on shuffled labels, with almost nothing true to learn. Each had to read 25 of 108 or less; if one read more, the recipe could lift the six-way without learning, and every tuned figure would be void.

The controls read 18, 29 and 24, so **T1 is VOID** by its own registered rule. We stopped on 2026-09-28 at 15:21:43 UTC with one real fine-tune trained and never read, one cut at step 127 and one never trained. None will be read: any figure would be void, and reading one would invite a story told after the fact.

That is a result, if a narrower one than it looks. The control that read 29 had memorised its shuffled labels (training loss 0.000 over its last tenth). Its 29 is over the registered bar, which chance alone passes about 3 times in 100, but under T1's own floor of 34, set to beat the 26 or so a guesser gets by crossing off the guests a line names. So on this recipe a fine-tune on shuffled labels can read above chance, and a tuned score could not have been told apart from that. T1's VOID is about our recipe and our control, not about deem-0.8-v1 as its authors ship it.

## What we did not measure, or did not report {#what-we-did-not-measure}

- **Other sets.** Jevify, Mistral Small 3.2 and OpenJev at 16 bits beyond the six-way; Kev, Lev and OpenJev at 8 bits on the held-out lines; Gemma 4 on the held-out lines, shuffled options and the judge question; imajev on shuffled options, long pages, the six-way and door on a graphics card, and images, the input its card leads with; OpenDecider, and APUS-OpenJev-v1 at `low`, on the three easier questions, and the latter on shuffled options; deem-0.8-v1 beyond the six-way and the held-out lines.
- **Other sizes and checkpoints.** Kev's 0.5B, 0.6B, 0.8B, 8B and 27B; APUS-OpenJev-v1 35B-A3B; deem-9b-v1; any tuned Deem checkpoint.
- **Other builds and runtimes.** Each small decider ran in bf16 through its own runtime (OpenDecider nano in fp32 on a processor), deem-0.8-v1 through its Python server, not the Rust server its card leads with ("Parity-gated against torch", it says). OpenJev was not run on the targeted readout its card's quick start sets, nor jevify through `jevify serve`: we asked a third party's 4-bit build, through ollama.
- **Calibration.** Probabilities read at a letter, from a head and from a written verdict are three kinds of number, and Kev's are rounded to two decimals.
- **Energy per decision** (recorded by several runs, recounted by none), **more than one request at a time**, and **the operating system of the older runs** (every run of our first Jev bench except the 4-bit OpenJev one; the cost series, the word counter and the 2026-09-13 door run).
- **A voice-control menu of ours**, which OpenDecider and deem-0.8-v1 also read.
- **Newer revisions, read 2026-09-28.** Kev-4B's adapter and head weights changed on 2026-09-24, OpenJev's main repository changed only its README, and imajev-4B's repository added two calibration files and changed its card on 2026-09-28, its adapter unchanged. Every figure is for the revision we benched.
- **Any product decision.** Nothing here puts a model in front of a visitor.

## How to check our work {#how-to-check-our-work}

- **The six-way is public:** [its file in the Reading the answer kit](https://research.strata2signal.com/reading-the-answer/data/kit/task_c.json), sha256 `61a0b1d6e7cf66fd91992c0c3a4572f6fbae90e6471048fd7dc1eb41ff2d7f39`, the 108 items every six-way run here was asked. Each model's row carries the sha256 of the prompt it was asked.
- **The recount re-derives the tables.** A standard-library script, in [this page's kit](data/) with its output, reads the row files, receipts and telemetry and re-derives every count, interval, tie range, paired test, six-order figure, median and memory reading in the tables, with no model and no network; for the sets held back below, the kit has figures, not files. Six of its sets were computed after the rows were seen and say so: the door by kind, the six-order means, the split by writer, imajev's graphics-card time, the memory table and the checks below. Quoted, not recounted: OpenDecider small against its base, T1's stop time and its control's loss, the 9B's margin and imajev's page fit, from the runs' own checks; the board; the model cards.
- **Checked for this page, and recounted by the same script:** the correction for seventeen tests, Kev-9B's paired interval, the bound on 0 of 12 and the held-out intervals' width, from the recount's counts; the resample by dinner, the tests in each option order, the paired shuffle counts, the 9B against the live door, Gemma 4's floored lines and Mistral Small 3.2's other run, from the row files below.
- **The pre-registrations.** Every arm run on 2026-09-28 but L-1 (Lev on a laptop's processor) pushed its pre-registration to our private repository, with its bar where it had one, before its first model call; L-1's was written first but never committed, and its rows' timestamps are the receipt. The 09-13 door run's was written first and committed with its results; the Jev bench's arms registered in its README before any row; Deem, the word counter, T1 and the cost series in their own. The word counter's 68 and 73 were seen before its registration, which says so. Both kits carry all but the cost series', which registers readings not reported here.
- **What is held back, and why.** The 48 held-out lines, since publishing them would spend the control, and their rows, since a line's hash would identify it on the public wall. Files in this page's kit still identify nine of the 48 on that wall. The planted door lines, hostile by design, and with them every per-question row of the doorman, the field exam and the judge question, as the Reading the answer kit [held back the same sets](https://research.strata2signal.com/reading-the-answer/data/KIT-NOTE.md).
- **The word counter ships whole:** its script and 227 labelled training lines (fingerprint `e4593840551294fa`, as its rows record), guests' lines from 32 test dinners of 2026-09-12, none a visitor's.
- **Run it yourself.** Pull the revision we benched ([next section](#where-each-figure-comes-from)), send the six-way through the model's own server one question at a time, score strict, and put your own labelled lines beside it: that is the bar.

## Where each figure comes from {#where-each-figure-comes-from}

Each run, the revision it pinned, and its per-question rows by their paths in the bench's record, with machine names rewritten; [this page's kit](data/) keeps those paths.

| Run | What ran, at the revision we benched | Row files |
|---|---|---|
| Arm 1 | Jevify, `mradermacher/jevify-gemma4-26b-a4b-GGUF` Q4_K_M at `8917f8f4`, and Gemma 4 26B-A4B, ollama `gemma4:26b`, digest `08ae7ec1744b`; card A | `jev-2026-09-21/rows/{jev,base}-{readout,generate}.c.jsonl` |
| Arms 2, 4 to 6, gap 4, 32k | OpenJev 27B, 8-bit, `openjev/openjev-FP8` at `4ec320f2`: the six-way and long pages on two used RTX 3090s (arm 2), then the door, field exam, judge question, six orders and ten longest pages | `jev-2026-09-21/rows/openjev-fp8-*`, `rows-addenda/openjev-fp8-*`, `rows-gaps-card/openjev-fp8-readout.creorder.jsonl` |
| Arm 3 | OpenJev 27B, 16-bit, `openjev/openjev` at `5ec9e5fd`, and 8-bit; the workstation card | `jev-2026-09-21/rows-arm3/openjev-{bf16,fp8}-largecard.c.jsonl` |
| Arm 10, 09-13 | Gemma 4 26B-A4B, 8-bit, on two used RTX 3090s (arm 10) and as the live doorman (09-13) | `jev-2026-09-21/rows-addenda/gemma4-fp8-*`; `doorman-planted-2026-09-13/out/verdicts.jsonl` |
| Arm 12 | Kev-9B, `jaredpalmer/kev-9b` at `2629c06a`, and Kev-4B, `jaredpalmer/kev-4b` at `485ace87`; card A | `jev-2026-09-21/rows-gaps-card/kev-{9b,4b}.*` |
| Arm 13 | OpenJev 27B, 4-bit, `openjev/openjev-GGUF` Q4_K_M at `208220bc`; card B | `jev-2026-09-21/rows-oj-gguf/openjev-q4km-*` |
| Cost series | Mistral Small 3.2 24B, ollama `mistral-small3.2:24b-instruct-2506-q4_K_M`, digest `5a408ab55df5`; card A and the workstation card | `cost-of-intelligence-2026-09-23/{benchbox-3090,largecard}-reading-0/rows/b*-mistral-small3.2*/*.c.jsonl`; each line's chair and dinner from `cost-of-intelligence-2026-09-23/KIT.lock.six-way.json` |
| NB, D0, T1 | Our word counter (a naive Bayes); deem-0.8-v1, `LibertAIDAI/deem-0.8-v1` at `8cbabbb2` (the v1.1 weights), on a laptop's processor; the three T1 controls | `deem-2026-09-27/t1-tune/rows/t1-{baseline-nb,null-*}.*`; `deem-2026-09-27/rows/d0-cpu-0.8.*`; the word counter's script and training lines, `deem-2026-09-27/t1-tune/{baseline_nb.py,kit/t1-pool.json}` |
| L-1, L-2 | Lev-4B, `interfaze-ai/lev` at `f8ef7115`; a laptop's processor and card B | `lev-2026-09-28/l1/rows/lev-4b-cpu.kc.jsonl`; `lev-2026-09-28/rows/lev-4b.*` |
| A-1, A-1b | APUS-OpenJev-v1 9B, `apus-ailab/APUS-OpenJev-v1-9B` at `82c9c56c`, and 4B, `apus-ailab/APUS-OpenJev-v1-4B` at `422b3741`, high and low; card B | `apus-2026-09-28/rows/apus-9b-*`; `apus-2026-09-28/a1b/rows/apus-4b-*` |
| I-1, I-2 | imajev-4B, `mohit67890/imajev-4b` at `11126e8c`; a laptop's processor and card B | `imajev-2026-09-28/rows/imajev-4b-cpu.*`; `imajev-2026-09-28/i2/rows/imajev-4b-3090.*` |
| O-1, O-2, O-2c | OpenDecider nano, `manjunathshiva/opendecider-nano` at `280219bd`, on a laptop's processor; OpenDecider small, `manjunathshiva/opendecider-small` at `ff25e366`, and `Qwen/Qwen3-4B-Instruct-2507` at `cdbee75f`; card B | `opendecider-2026-09-28/rows/{o1-cpu-nano,o2-gpu-small,o2c-gpu-qwen3-4b-base}.*` |
| Memory | Growth at load and highest reading, from each run's receipts and telemetry | `counted-run.json` in `jev-2026-09-21/receipts/oj-gguf/run/`, `apus-2026-09-28/receipts/`, `apus-2026-09-28/a1b/receipts/` and `lev-2026-09-28/receipts/`; `imajev-2026-09-28/i2/receipts/counted-unit.log` (held back with its sets); `opendecider-2026-09-28/rows/o2-gpu-small.*.pass.json`; Kev's and arm 13's per-task `*.watts.csv` |

*File names end in the question they hold: `.kc`, `.c` and `.s0` the six-way; `.kh` and `.h48` the held-out lines; `.kd`, `.d` and `.door` the doorman; `.kf` and `.f` the field exam; `.kj` and `.j` the judge question; `.ka`, `.a` and `.along` long pages; `.kperm`, `.creorder` and `.k1-5` shuffled options. The held-out, doorman, field-exam, judge and voice-control-menu files are held back. The six-way, long-page and shuffle files of arms 1, 2, 3, 10 and 12 and gap 4 are already public in [the Reading the answer kit](https://research.strata2signal.com/reading-the-answer/data/) (without `jev-2026-09-21/`); the rest are in this page's kit.*

*The jevify GGUF is built from the merged weights in `kushalpatil/jevify-gemma4-26b-a4b` (`d4c0d1d4`), which our rows do not pin; each ollama digest is the one recorded at the run and still in the registry's manifest. Rows join on each line's `uid` and text hash or, where a file has neither, on kit position (its `id` repeats: 50 values over 108 lines).*

## Who ran this, and thanks {#who-ran-this-and-thanks}

**TypeSafe** defined the request shape that Kev, Lev, APUS-OpenJev-v1 and imajev answered here, and that OpenJev, jevify and deem-0.8-v1 say they accept. The **OpenJev** authors published their weights and their own 4-bit build; **`jaredpalmer`** published Kev; **Interfaze** published Lev; **APUS AI-LAB** (its cards name gumpcheng and zhangxu) published APUS-OpenJev-v1; **Mohit Garg** published imajev; **Manjunath Janardhan** (`manjunathshiva`) published OpenDecider nano and small; **LibertAI Labs** published Deem; **`kushalpatil`** published the jevify fine-tune, and **`mradermacher`** its GGUF builds. **Google** made Gemma 4, **Mistral AI** Mistral Small 3.2, **Alibaba's Qwen team** the Qwen3.8, Qwen3.5 and Qwen3 models under OpenJev and most of these deciders, and **Johns Hopkins University** the Ettin encoder under OpenDecider nano. **Benchmark Heaven** runs JevBench. The runs stood on **vLLM**, **ollama**, **llama.cpp**'s GGUF format, **transformers**, **PEFT**, **flash-linear-attention** and **PyTorch**. None of them owed us anything. If you built one of these models and we have described it wrongly, tell us at [hello@strata2signal.com](mailto:hello@strata2signal.com); corrections are dated on this page.

A small (human) team decided what to ask. A fleet of AI agents pre-registered each run, ran it, checked it, recounted the figures from the rows, and wrote this page. The agents are Anthropic's Claude models. No Anthropic model is measured on this page.

<!-- derived 2026-09-30 (UTC) by tools/derive_md.py from the pour source.
     source html sha256: ae631b1e5dc4f9f6c3791f279138a38af4871e8b947b8a68ea2330e955e97f0c
     derivation sha256:  5e5b3c02528f3ca3215fa9f7708bff040331ba34e30d1a09aace9af746506e6b
     the {#id} on each heading is the anchor that heading carries on the page. -->
