## ROSTER ADDENDUM 2026-08-28 — the REFERENCE ARM class, and its first member (glm-5.3-flash:cloud)

**What this section is, for a reader who scrolled straight to it.** This is a
dated addition to the Chair Trials roster (the plan above was frozen 2026-08-12
and amended 2026-08-17). It does two separate things: it *defines a new class of
row* — the reference arm — and it *admits one model into that class*,
`glm-5.3-flash:cloud` (Z.ai GLM-5.3-Flash; 320B total / 18B active MoE, 1M
context, MIT licence, published to ollama's cloud 2026-08-26, read off the
library page 2026-08-28T16:2xZ). Operator ruling D-20260828-18, 2026-08-28.
Nothing above this line is amended, revised, or overturned by it.

### The reference-arm class — what it is, and what it can never be

A **reference arm** is a dated ceiling reading. It is measured on our own frozen
instruments, at our own dated moment, and published beside the candidate rows so
a reader can see how far the local seat sits from a very large hosted model on
the same nineteen tasks and the same sixty items. That is the whole of its job.

Four fences, all binding:

1. **Never a candidate.** A reference arm is not eligible for any chair, seat, or
   product decision this house makes, no matter what it scores. It cannot clear a
   floor, because no floor is applied to it.
2. **Ranked against nothing.** Its row carries no ordering, no margin, no
   "beats"/"loses to" language, and never appears in a ranked table. It is
   printed beside the rows; it does not sit in them.
3. **Dated, and only ever dated.** A hosted model behind a tag is not a fixed
   object — the weights, the router, and the serving stack behind
   `glm-5.3-flash:cloud` can change under the same spelling with no notice and no
   digest we can pin. So a reference-arm reading is a statement about **one day**,
   stamped in UTC at both ends, and is never carried forward as "the ceiling"
   after that day.
4. **No thresholds are ever derived from it.** A reference arm does not set,
   move, or justify any floor, gate, or seat criterion. If a future rung wants a
   threshold, it pre-registers one on measured local evidence, as always.

**This does NOT overturn the cloud-only rejection above.** The roster section
records `glm-5.2` as a *size-check rejection* — "cloud-only, zero local weights"
— and that rejection stands exactly as written, for exactly the reason written.
The estate ethos forbids cloud AI inside live products; a model with no local
weights can never hold a seat here, and `glm-5.3-flash:cloud` never will either
(it is cloud-only on ollama by construction — there is no local blob to pull).
The reference-arm class does not make cloud models candidates. It creates a
*second, separate kind of row* for a model that is disqualified as a candidate
but is still informative as a dated ceiling. Precedent: kimi-k3's C1 row was
published EXPLORATORY on the cloud protocol (RULINGS-0812, night rulings) —
published, never ranked. The reference arm is that posture, named and fenced.

### The two legs a reference arm runs, and why only two

- **C3 — the toolbench (hero), 19 tasks exact**, in the five frozen categories:
  single-call 6 · chains 5 · distractor 2 · honesty traps 5 · error-recovery 1.
  Same `tasks.json`, same ten tools, same 2 repeats, same 8-round cap, same
  scorer. Native `tools`-array emission, so no schema grammar is involved.
- **The field exam, 60 items** (T1 grounded-qa 20 · T2 stay-grounded 20 ·
  T3 schema-extract 20), the estate kit at its own frozen settings.

Both are **DESCRIPTIVE** legs under the verdict discipline above — counts of N,
Wilson intervals beside every rate, no floors, no winner, no ordering. That is
not a concession made for the reference arm; it is what C3 and the field exam
already are.

Nothing else runs. C1 (judge), C2 (assistant), C5 (speed), and C7 (filing
cabinet) are **NOT-RUN**, each for a stated reason rather than a blank:

| leg | disposition | reason |
|---|---|---|
| C1 judge | NOT-RUN (cloud path invalid) | the verdict is bound by a `format` JSON-schema grammar; ollama's cloud path **accepts `format` and silently ignores it** (docs + probed 2026-07-31; ollama/ollama#12362 open). A schema-bound leg run where the schema is not enforced is a different instrument wearing the published name. |
| C2 assistant | NOT-RUN (cloud path invalid) | same reason — its checked items bind on enforced JSON. |
| C5 speed | NOT-RUN (not our silicon) | a decode rate measured over the public internet against someone else's fleet, under someone else's batching, is not a number this house has any way to attribute. Publishing it beside our own tok/s rows would invite exactly the comparison it cannot support. |
| C7 filing | NOT-RUN (not our silicon) | same; the long-context recall leg's value here is bound to a VRAM-resident model on a card we can name. |
| field exam T3 (schema-extract, 20 of the 60) | reported as **MEASUREMENT FAILURE (cloud: `format` unenforced)**, never as a score | T3 is the one field-exam task that sends `format`. T1+T2 (40 items) send none and are valid on this path. The 60-item headline is therefore **not** quotable for a cloud arm; the arm's field row reads as 40 valid items plus a named 20-item exclusion. |

### The gate this arm must clear before a single scored call

**Tool emission on the cloud path, proven, not assumed.** Before any C3 scoring,
one `/api/chat` round trip carrying a `tools` array must come back with a
populated `tool_calls`. If it does not, the arm is **NOT-CARRIED** with its raw
emission shown, and no score is recorded — the same rule every local candidate
faces at the C3 pre-probe. Two known cloud-path traps make this a real gate and
not a formality: `format` is unenforced (above), and GLM-5.3-Flash is documented
as *reasoning-always-on with a per-request effort dial*, so the estate's explicit
`think: false` may be refused or ignored where a local model would honour it.
The posture that actually rode the wire is read back off the response and
recorded, as for every other row.

### Provenance, cost, and what the row must carry

The tag is hosted, not pulled: there is no digest, no blob, no `ollama show
--license` receipt to take locally. The licence (MIT) and the architecture
figures (320B total / 18B active, 1M context) are **vendor-stated, read off the
ollama library page on the date above** and labelled as vendor figures in the
row — the same convention the roster already uses for active-params cells. Cost
is ollama-cloud usage against the operators' account (library page: "Medium Usage"
tier); whatever usage signal the API exposes is recorded verbatim in the run
manifest, and any figure that was not actually taken is not quoted.

Every published cell for this arm is labelled **REFERENCE ARM · cloud · dated**.
