## RESULTS — per candidate x per instrument

Every cell is copied from that leg's own scorer JSON (named in §sources); this view computes nothing. Blanks are never bare: **NOT-RUN (spill)** is a doctrine decision from a measured fit verdict, **NOT-RUN (disk STOP)** is a safety trigger, **NOT-RUN (roster trim)** is an operator time-box — three different reasons a cell can be empty, kept apart on purpose.

### muse-glimmer:30b  ·  _re-probe_

| instrument | result |
|---|---|
| fit (24 GB) | FITS 100% (15831/15831 MiB) |
| C0 polarity | STANDARD TRAP · **SEAT-BLOCKED** |
| C1 judge | **medium** kill 11/12 · pres 7/9 · did NOT clear · EXPLORATORY<br>**high** kill 12/12 · pres 6/9 · did NOT clear · EXPLORATORY |
| C2 asst | 16/20 (exact 13/13, proxy 3/7) · DESCRIPTIVE |
| C3 tools | 13/19 · UNMEASURABLE |
| C5 tok/s | 1k **41.1** · 8k **NOT-CARRIED** (10 null) · 32k **39.52** tok/s · cold 15.52s<br><sub>recovered from raw — summary pass crashed</sub> |
| C7 filing | recall 18/18 · abstain 18/18 · fab 0 · DESCRIPTIVE |
| field | 55/60 · 91.7% |
| seat-43 | **UNMEASURABLE** — resp-fail over the 10% ceiling; counts not scored<br><sub>kills 27/27 (PASS) · pres 6/16 (FAIL) · resp-fail 18/129</sub> |

### gemma4:31b  ·  _new kid_

| instrument | result |
|---|---|
| fit (24 GB) | **SPILLS** 92% (−1541 MiB→CPU) |
| C0 polarity | SCHEMA HELD BOTH WAYS |
| C1 judge | **think:false** kill 0/12 · pres 0/9 · did NOT clear · NOT-CARRIED |
| C2 asst | 16/20 (exact 13/13, proxy 3/7) · DESCRIPTIVE |
| C3 tools | 12/19 · RANKED |
| C5 tok/s | NOT-RUN (spill) |
| C7 filing | NOT-RUN (spill) |
| field | 58/60 · 96.7% |
| seat-43 | **FAIL** · kills 27/27 (PASS) · pres 9/16 (FAIL) · resp-fail 12/129 |

### laguna-xs-2.1:latest  ·  _new kid_

| instrument | result |
|---|---|
| fit (24 GB) | FITS 100% (19440/19440 MiB) |
| C0 polarity | INVERTED TRAP |
| C1 judge | NOT-RUN (disk STOP) → NOT-RUN (roster trim) |
| C2 asst | NOT-RUN (roster trim) |
| C3 tools | NOT-RUN (roster trim) |
| C5 tok/s | NOT-RUN (roster trim) |
| C7 filing | NOT-RUN (roster trim) |
| field | NOT-RUN (roster trim) |
| seat-43 | NOT-RUN (roster trim) |

### qwen3.6:27b  ·  _repro row_

| instrument | result |
|---|---|
| fit (24 GB) | FITS 100% (16469/16469 MiB) |
| C0 polarity | SCHEMA HELD BOTH WAYS |
| C1 judge | **think:false** kill 11/12 · pres 8/9 · CLEARED · RANKED |
| C2 asst | 15/20 (exact 11/13, proxy 4/7) · DESCRIPTIVE |
| C3 tools | 14/19 · RANKED |
| C5 tok/s | 1k **62.02** · 8k **60.75** · 32k **59.14** tok/s · cold 11.84s<br><sub>recovered from raw — summary pass crashed</sub> |
| C7 filing | recall 16/18 · abstain 18/18 · fab 0 · DESCRIPTIVE |
| field | 58/60 · 96.7% |
| seat-43 | **UNMEASURABLE** — resp-fail over the 10% ceiling; counts not scored<br><sub>kills 7/27 (FAIL) · pres 5/16 (FAIL) · resp-fail 93/129</sub> |

### nemotron-3.5-lightning:30b  ·  _repro row_

| instrument | result |
|---|---|
| fit (24 GB) | **SPILLS** 83% (−4136 MiB→CPU) |
| C0 polarity | NOT-RUN (roster trim) |
| C1 judge | NOT-RUN (roster trim) |
| C2 asst | NOT-RUN (roster trim) |
| C3 tools | NOT-RUN (roster trim) |
| C5 tok/s | NOT-RUN (roster trim) |
| C7 filing | NOT-RUN (roster trim) |
| field | NOT-RUN (roster trim) |
| seat-43 | NOT-RUN (roster trim) |

### sources
`fit-<slug>.json` · `probes-<slug>.json` · `summary-C1-<slug>.json` · `c2-scores.json` · `c3-scores.json` · `summary-C5-<slug>.json` · `c7-<slug>.json` · `field/<model>.json` · `seat43-summary.json` · `c5-recovered.json`.

---

### glm-5.3-flash:cloud  ·  _REFERENCE ARM · cloud · dated 2026-08-28_

**Not a candidate, and not comparable to the rows above by ranking.** A reference
arm is a dated ceiling reading from a hosted model — printed beside the candidate
rows, ranked against nothing, never eligible for a seat, and never the basis of a
threshold. Class defined in PLAN.md, "ROSTER ADDENDUM 2026-08-28" (operator ruling
D-20260828-18). It does not overturn the roster's `glm-5.2` size-check rejection:
cloud-only models remain permanently ineligible as candidates.

**Measured 2026-08-28T16:37:59Z → 2026-08-28T16:49:43Z** (UTC, both ends), from
the bench box against the local ollama daemon at `the bench endpoint`, v0.32.14,
which forwards `:cloud` tags upstream. No estate seat port was addressed and no
inference or production host of ours was contacted at any point. The 2026-08-28 morning
attempt recorded here as NOT-RUN (cloud path unauthorised) was unblocked by the
operator running `ollama signin` on the bench box; that 401 history is kept verbatim in
the gate receipt as provenance.

**Posture: `think:true`, and the reason is measured, not stylistic.** GLM-5.3-Flash
reasons unconditionally, so `think` does not switch reasoning off on this path — it
only decides whether ollama *parses* the reasoning out of the reply. Under
`think:false` there is no `thinking` field at all and the raw chain-of-thought lands
in `message.content` (3 of 4 probes closed by a literal `</think>`, 1 of 4 not
delimited at all). Every leg here scores `message.content`, so `think:false` would
score the model's private deliberation as its answer. **`think:false` is therefore a
MEASUREMENT FAILURE on this path, not a low score, and is not run** — a fifth reason
a cell can be empty, alongside the four below.

A fourth reason a cell can be empty joins the three in the header: **NOT-RUN (cloud
path unauthorised)** is a transport state — the request was rejected at HTTP 401
before reaching the model — and it is a statement about this box's ollama.com
sign-in, not about the model. It no longer applies to this row; it is retained
because the definition is what makes the row's history readable.

| instrument | result |
|---|---|
| fit (24 GB) | NOT-APPLICABLE (hosted; no local weights, nothing to fit) |
| C0 polarity | NOT-RUN (reference arm; only C3 and the field exam are carried) |
| C1 judge | NOT-RUN (cloud path invalid) — verdict binds on a `format` schema; **confirmed on this date**: `format` sent with a prose-inviting prompt returned prose, so the grammar is accepted and ignored |
| C2 asst | NOT-RUN (cloud path invalid) — same `format` defect |
| C3 tools | **16/19** · Wilson95 [62.4%, 94.5%] · think_true · DESCRIPTIVE<br><sub>single 5/6 · chains 5/5 · distractor 2/2 · honesty 3/5 · error-recovery 1/1</sub> |
| C5 tok/s | NOT-RUN (not our silicon) — decode rate over the public internet is unattributable |
| C7 filing | NOT-RUN (not our silicon) |
| field | **39/40 valid** · Wilson95 [87.1%, 99.6%] · T1 20/20 · T2 19/20<br><sub>T3 schema-extract (20 of 60) EXCLUDED — MEASUREMENT FAILURE (cloud: `format` unenforced). The model did emit conforming JSON 20/20 unaided, but the schema-enforced instrument was never applied, so that is a probe and not a score. **The 60-item headline is not quotable on this path.**</sub> |
| seat-43 | NOT-RUN (reference arm; seat exams do not apply to a model that can never hold a seat) |

**Beside the seat — printed for scale, ordered by nothing.** The scorer assigns
outcome states mechanically and stamped this row `RANKED`; the class label overrides
it, and the arm is **ranked against nothing**. Neither column is a threshold.

| | gemma4:26b — the seat (2026-08-12) | glm-5.3-flash:cloud (2026-08-28) |
|---|---|---|
| C3 task success | 14/19 think_false · 15/19 think_true | **16/19** think_true |
| C3 grounded task success | 12/19 · 12/19 | 12/19 |
| C3 honesty traps | 2/5 · 2/5 | 3/5 |
| C3 chain completion | 2/5 · 2/5 | 1/5 |
| C3 argument fidelity | 58/66 (87.9%) · 59/66 (89.4%) | 49/66 (74.2%) |
| C3 spurious calls | 15/63 (23.8%) · 25/86 (29.1%) | 39/88 (44.3%) · Wilson95 [34.4%, 54.7%] |
| C3 error recovery | recovered ×2 (8 calls) | recovered ×2 (6 calls) |
| field, the comparable 40 (T1+T2) | **40/40** | 39/40 |

Two dates are two different instruments and the house publishes no delta across
that line; the seat's C3 rows were taken on the shared production daemon on
2026-08-12, and this arm on the cloud path on 2026-08-28.

**Cost signals for the whole arm** (ollama exposes no per-call price; these are the
response's own counters, summed verbatim): C3 — 100 scored rounds + 4 pre-probe
responses, `eval_count` **8,115**, `prompt_eval_count` **195,532**,
`total_duration` **155.4 s**. Field exam — 60 items, `eval_count` **4,166**,
`prompt_eval_count` **14,424**, wall **67.6 s**. Arm total: `eval_count`
**12,281**, `prompt_eval_count` **209,956**. Note the shape of that pair — the
prompt side is ~17× the completion side, because C3 resends the ten tool schemas
and the whole transcript on every round, so a toolbench is a prompt-token cost
before it is a completion-token cost. Library page places the tag in the
"Medium Usage" tier against the operators' ollama.com plan; no currency figure is quoted,
because none was taken.

Narrative + receipts: `REFERENCE-glm-5.3-flash-2026-08-28.md` ·
`REFERENCE-glm-5.3-flash-2026-08-28.gate.json` ·
`raw/c3/glm-5.3-flash_cloud/` · `raw/c3/manifest-glm-5.3-flash-cloud-2026-08-28.json` ·
`field/glm-5.3-flash-cloud.json`.
