{
 "event": "BUDGET \u2014 gpt-5.5 C-NONE leg re-sized, not run to the cap",
 "at": "2026-08-16T14:58:36Z",
 "arm": "openai-gpt-5.5",
 "condition": "C-NONE",
 "cells_collected": 5,
 "cells_not_collected": 19,
 "output_tokens": {
  "values": [
   4151,
   5187,
   3633,
   1155,
   1284
  ],
  "median": 3633,
  "registered_limit": 900,
  "over_limit_factor": 4.04
 },
 "prompt_tokens": {
  "total": 372
 },
 "spend_usd_uncached": 0.4642,
 "sub_cap_usd": 2.6,
 "projected_full_leg_usd": 2.228,
 "why_the_precondition_9_probe_missed_it": "the registered probe ran on a C-MAP-shaped prompt, per the plan's own wording, and measured a 58-token output median. The blow-up is CONDITION-SPECIFIC: with no site content in context the model has nothing to ground on and reasons at length before answering \"the site doesn't say\". A probe that only samples the grounded condition is blind to the ungrounded one, and this bench's own probe was.",
 "disposition": "PLAN-THIRTEEN registers it: 'if the measured median exceeds 900 output tokens, the leg is re-sized or that arm publishes NOT-RUN \u2014 BUDGET before it starts'. The leg is RE-SIZED: the 5 collected cells stand as data, the remaining 19 publish NOT-COLLECTED \u2014 BUDGET, and this arm's C-NONE denominator re-prints as 5/24 wherever it appears.",
 "consequence_for_the_filter": "the global exclusion rule counts arms answering correctly closed-book; on the 19 items without a gpt-5.5 C-NONE cell this arm is simply not a voter. The rule still requires >=2 arms from >=2 families, and seven arms voted on every item.",
 "headline_is_unaffected": "C-NONE is the contamination meter, not a competitor. The headline conditions for this arm measured 58 output tokens in the pre-leg probe and are projected at ~$0.85 for all 48 cells, comfortably inside the $2.13 remaining on its sub-cap."
}
