# Arm D + TT, the re-run: results (Ollama vs vLLM, release 1)

*Cold scroll: what this file is.* The readings of arm D + TT's re-run, the first chain of release 1's day. Arm D + TT
is the bench's whole path run once on the versions the article stamps: vLLM 0.30.0, Ollama 0.34.4 with 16 slots, and
llama-server from Ollama 0.34.4's own bundle (b11081), on one EVGA GeForce RTX 3090 24 GB at 300 W with persistence
on. TT is the first-token decomposition with its registered hypothesis (`../PREREG.md` Part 1 §2.7 and Amendment 2).
The first D, on 2026-09-28 (`../dtt/`), was REFUSED on vLLM's BOS. Amendment 3 registered the fix (#196), and D
re-ran whole on the fixed kit. This run is `ovv-20260930T200002Z-D`. Its window is 2026-09-30T20:00:02Z to 20:07:18Z
(UTC; `rows/D.json` `utc_start` and `utc_end`), and its first recorded run is at 20:01:26Z. It ran on the kit at
estate `ae950129` (kit tree `1d89dacb`, `KIT_VERSION` v2), under Amendment 4, whose PREREG sha256 the rows carry
(`prereg.prereg_sha256`, `withheld-2026-10-07`). The box: an AMD Ryzen 7 3700X, Ubuntu 26.04 LTS, kernel 7.0.0-31-generic,
NVIDIA driver 595.84 (`host`, `os`). This folder is the results PR the day's runbook calls PR-A (§2.7). Arms C, P,
Q, A and G land in folders of their own.

*Where the figures come from.* Every figure here is read from a row file, with its gates. The re-run's figures come
from this folder's `rows/`, the kit's outward files, written through its pen. The first D's figures, set beside them,
come from `../dtt/rows/D.json` (landed by #191). Each figure is cited by its JSON field, and the provenance table at
the end gives the file, the field and the line of every figure a page may carry. No figure is taken from a receipt
line alone. Three facts come from the runs' private halves, in words only and with no figure: the compile line kinds
in the vLLM serve log, the saved compile it loaded (the same directory name in both D runs' serve logs), and Ollama's
`cloud disabled` line. No row file holds a `voided_by_gate` field. The gates read here are the rows' own:
`gates[].ok`; each engine's `refused`, `residency`, `gone`, `unread` and `servelines_mismatches`; each cell's
`voided`, `voids` and `flags`; and the run's `halted`, `refused`, `held` and `withheld`.

## The line

**D's re-run read clean on every gate that stops D (Amendment 2 item 5), with one VOID level: TT road (d)'s, as
registered.** TT reads **CONFIRMED**. It is a replication, and it reads the verdict of record, so arms C and P carry
the reading of record and their first-token rows are not held (Amendment 3 item 6). P-BOOT and C-2 read **CONFIRMED**
on this boot, which loaded a saved compile: a warm compile cache. The first D's boot compiled cold, and it read both
**REFUTED**. Both readings stay on the record, and each prints with its condition (the lead's ruling OF-RECORD,
D-20261001-067).

```
RECEIPT ovv D ovv-20260930T200002Z-D frame=ok render=differs:llama-server raw_parity=368/368 nonce_counts=ok eligible=series-p512,long-write,long-prefill tt=CONFIRMED replan_c=re-planned calibration=ok egress=ready void_levels=1 halted=no
RECEIPT ovv chain ovv-20260930T200002Z-D arms=1/1 void_levels=1 halted=no
```

- `void_levels=1` is road (d)'s registered VOID: hidden reasoning, because Ollama's `/v1/completions` is not raw.
  Road (c)'s reload is its registered reading (`replan_c=re-planned`), and it never counts as a void level
  (Amendment 2 item 4).
- `halted`, `refused` and `held` read null, and `withheld` and `shm_swept` are empty.

## P-BOOT and C-2: two readings, two compile states

The registered words (`verdicts[1].words`, `verdicts[2].words`):
- **P-BOOT:** "vLLM's boot line reports a maximum concurrency ≥ 16 at 5,120 tokens with 512 batched tokens".
- **C-2:** "vLLM reserves ≥ 90 % of the card at rest at its default".

Each run reads both from vLLM 0.30.0's headline boot, once.

| | The first D, 2026-09-28: a cold compile | This re-run, 2026-09-30: a warm compile cache | Field (in each file) |
|---|---|---|---|
| Run and file | `ovv-20260928T172031Z-D`, `../dtt/rows/D.json` | `ovv-20260930T200002Z-D`, `rows/D.json` | `run_id` |
| Launch to ready | 144.35 s | 53.79 s | `engines[0].boot_s` |
| Peak activation in vLLM's memory profile | 2.47 GiB | 0.85 GiB | `engines[0].readback.boot.peak_activation_gib` |
| KV pool | 10.47 GiB | 12.09 GiB | `engines[0].readback.boot.kv_pool_gib` |
| KV blocks | 10,718 | 12,377 | `engines[0].readback.cache_config.num_gpu_blocks` |
| KV tokens | 75,691 | 87,407 | `engines[0].readback.kv_cache_tokens` |
| The boot line's maximum concurrency at 5,120 tokens | 14.78× | 17.07× | `engines[0].readback.max_concurrency_x` |
| **P-BOOT** (≥ 16) | **REFUTED**, "14.78 < 16" | **CONFIRMED**, "17.07 >= 16" | `verdicts[1]` |
| Grown at rest | 20,847 MiB | 22,505 MiB | `engines[0].memory.at_rest_mib` |
| Share of the card's 24,576 MiB | 0.8483 (84.83 %) | 0.9157 (91.57 %) | `verdicts[2].value`; R-1's `memory.total` |
| **C-2** (≥ 0.90) | **REFUTED**, "0.8483 < 0.9" | **CONFIRMED**, "0.9157 >= 0.9" | `verdicts[2]` |
| Compile-cache files before the boot | 18,270 | 19,259 | `engines[0].internals.compile_cache_files_before` |
| The kit's `cold_compile` | false | false | `engines[0].internals.cold_compile` |

- **What the condition is.** On a configuration's first boot, vLLM compiles the model's graph and saves the result.
  A later boot loads it.
  - The first D's boot was the bench box's first vLLM 0.30.0 boot. Its serve log printed a `Compiling a graph for
    compile range` line (quoted in `../dtt/RESULTS.md`). The first D refused at R-F1, for a cause (vLLM's BOS) that
    does not touch the boot line; its P-BOOT and C-2 stand as the kit's readings (D-20260929-002).
  - This boot's serve log has a `Directly load AOT compilation` line and no compile-range line. That log is in the
    private half, and it was read by line kind only. The directory it loaded is the one the first D's boot saved:
    the same directory name appears in both runs' private serve logs.
  - The loading boot profiled a peak 1.62 GiB lower, and vLLM gave its KV cache 1.62 GiB more (2.47 − 0.85 and
    12.09 − 10.47; arithmetic).
- **The kit's flag cannot carry the label.** `cold_compile` reads false on both boots, because the shared compile
  cache already held files before each one: 18,270 and 19,259. Amendment 2 item 7 registered this. The label is the
  day's runbook's (§2.8): the cold case stays the first D's 0.30.0 boot, n = 1, and the day's boots load from a warm
  compile cache.
- **Each bar sits between the two readings.** P-BOOT's 16× lies between 14.78× and 17.07×, and C-2's 0.90 between
  0.8483 and 0.9157. The compile state, not run-to-run noise, separates them.
- **How the two readings print.** Every P-BOOT and C-2 line prints both readings, each with its condition
  (OF-RECORD, D-20261001-067):
  - cold: 14.78× and 84.83 %, REFUTED;
  - warm: 17.07× and 91.57 %, CONFIRMED.

  `cold_compile: false` is never printed as the condition, and the refuted reading stays on the record.

## TT, as registered: a replication beside the reading of record

**The hypothesis** (Part 1 §2.7, M-8): with `num_ctx` absent from the request and the context set at the server,
Ollama's level-1 `load_duration` on `/api/generate` is ≤ 50 ms.

**The roads:** road (a) is that request. The roads ran in the order a, b, d, c, e, each with one warm-up and 10 scored
runs, at level 1 on `series-p512` (Amendment 2 item 4).

**CONFIRMED.** Road (a)'s 10 scored runs read `load_duration` from 1.31 ms to 3.119 ms, every one ≤ 50 ms
(`verdicts[0]`: "max 3.119 <= 50").
- **The reading of record** is the first D's: run `ovv-20260928T172031Z-D`, "max 3.111 <= 50" (`../dtt/rows/D.json`
  `verdicts[0]`). The kit's `prereg.TT_OF_RECORD` names that file and its sha256, `withheld-2026-10-07`.
- **This replication reads the same verdict.** So arms C and P carry the reading of record, with this one beside it,
  and their first-token rows are not held (Amendment 3 item 6).

| Road | What | Runs after the warm-up | First token, median (min–max), ms | `load_duration`, min–max, ms | State |
|---|---|---:|---|---|---|
| a | `/api/generate` raw, no `num_ctx` | 10 | 445.128 (440.358–580.697) | 1.31–3.119 | clean |
| b | the same, with `num_ctx` 5120 | 10 | 486.951 (445.728–581.771) | 1.24–2.622 | clean |
| d | `/v1/completions`, the same string | 10 | VOID(hidden_reasoning) | none (no server clock) | VOID: 256 hidden tokens on every run |
| c | as (a), with `num_ctx` 8192 | 10 | VOID(reload) | VOID(reload) | VOID: the reload, its registered reading |
| e | llama-server alone, `/completion` | 10 | 441.351 (438.671–571.757) | none (no such field) | clean |

- **A VOID level is printed `VOID(<kinds>)` and never as a number** (Amendment 2 item 5). The kit keeps roads (c)
  and (d)'s runs in `rows/D.json`, with their medians moved to `unscored_medians`. None of those runs, and no
  warm-up, is a figure here.
- **Road (d)'s VOID** is the caveat the kit printed on it (`cells[10].tt_caveat`): Ollama's `/v1/completions` is not
  raw. It renders the string as a user turn, with thinking at the model's default.
  - Every run read 563 prompt tokens, where road (a) reads 548.
  - Every run carried 256 hidden tokens, and every visible completion was empty (`cells[10].runs[].streams[0]`).
  - No logprob replay ran on it: `voids[0].replay` is empty, and `confirmed` is null. So it reads
    `VOID(hidden_reasoning)` with no replay verdict beside it.

  It stays as its registered reading (Amendment 3 item 6).
- **Road (c)'s reading is the reload itself.** Its `num_ctx` of 8192 differs from the server's 5,120 per slot.
  - The runner reloaded once, in the road's warm-up (`cells[11].all_layers.reloads`; `cells[11].tt_reading`), and it
    still offloaded 49 of 49 layers.
  - The re-plan hypothesis read "re-planned" (`d.tt.replan_c`).

  The reload does not count in `void_levels`, and none of the road's runs is a figure.
- **These are D's dry-path readings, at level 1 on `series-p512`, one request at a time.** They are not C's sweep,
  and the article's first-token rows come from C.

## Every gate

| Gate | Result | Reading (`rows/D.json`) |
|---|---|---|
| The staged tree | pass | estate `ae950129`, kit tree `1d89dacb`, 112 files, clean (`gates[0]`) |
| No other bench unit · disk floor · UPS | pass | No matching unit active. 98.34 GiB free on each store, against a 15 GiB floor. The UPS on line at 5.0 % and 57.0 W, for the whole box (`gates[1]` to `gates[3]`) |
| R-1, the card · its cap | pass | One card, EVGA GeForce RTX 3090 24 GB, with 24,576 MiB. 300.0 W against a 350.0 W default, within 1.0 W. Persistence Enabled (`gates[4]`, `gates[5]`) |
| R-2, the weights | pass | vLLM's eight checkpoint files read the bytes and sha256 that `MODEL.lock.json` pins. The Ollama tag's manifest `38044be4…`, 4 layers (`gates[6]`, `gates[7]`) |
| R-4 · R-5 | pass | The bench's three loopback ports were free, and no tenant held the card (`gates[8]`, `gates[9]`) |
| S-1, vLLM 0.30.0 read back | pass | Version 0.30.0; `max_seq_len` 5120; `max_num_batched_tokens` 512; prefix caching on; `gpu_memory_utilization` 0.92; the concurrency line present. Nothing UNREAD and no mismatch (`engines[0]`) |
| S-1, Ollama 0.34.4 16 slots read back | pass | Runner argv `-c 81920 -np 16 -b 512 -ub 512`, from `/proc` and from the log, agreeing. Nothing UNREAD (`engines[1]`) |
| S-1, llama-server read back | pass | Ollama's runner argv, with its port replaced and `--device CUDA0` added. Nothing UNREAD (`engines[2]`) |
| R-DEV | pass | One CUDA device (`CUDA0`) and no Vulkan device (`engines[2].readback.devices`) |
| All layers | pass | `offloaded 49/49 layers to GPU` on Ollama's boot and on llama-server's |
| Residency by growth | pass | vLLM grew 22,505 MiB (floor 7,847), Ollama 16,389 (floor 6,478) and llama-server 16,377 (floor 6,478) |
| R-5-held | pass | The board read 1.0 MiB after each engine (baseline 1.0), with no process on it |
| R-F1, the frame is the template | pass | vLLM's chat render equals the composed string on all four prompts: 548, 2,612, 2,553 and 90 IDs (`d.frame`) |
| R-P2, raw-path parity | pass, 368/368 | Every raw-path string of the arms the kit carries (`d.strings_checked`) reads the same IDs on every engine, BOS first exactly once (`d.raw_parity`) |
| R-N, nonce counts | pass | Each prompt's nonce strings share one count on every engine (`d.nonce_counts`) |
| Hidden-gate calibration | pass, 12/12 | Exact 0 and re-tokenised 0 on every engine × prompt. `calibration_blocks_C` is empty (`d.calibration`) |
| Hidden-token gate | 1 VOID | TT road (d) only. Every other stream reads 0 hidden tokens |
| The egress preflight, for arm G | ready | bwrap 0.11.1 and strace 6.19 present; the device check ok. The positive control was seen: its lookup and its connect were refused as unreachable, and neither succeeded (`d.egress_gates`) |
| Stop guard | no stop | `halted` is null. The rows' 67 per-run UPS records (140 reads, none blind) peak at 40.0 % load and 408.0 W, for the whole box, against an 80 % stop, and none tripped. The card's core peaks at 70.0 °C over `D.power.csv`'s 478 samples, against an 83 °C stop (Amendment 2 item 10) |

## The boots (n = 1 each, page cache warm)

| Engine · posture | Launch to ready, s | Grown at rest, MiB | The first D's, 2026-09-28 (`../dtt/rows/D.json`) |
|---|---:|---:|---|
| vLLM 0.30.0, headline | 53.79 (a warm compile cache) | 22,505 | 144.35 s (a cold compile); 20,847 MiB |
| Ollama 0.34.4, 16 slots | 2.86 | 16,389 | 4.06 s; 16,387 MiB |
| llama-server b11081, the bundle | 2.95 | 16,377 | 2.92 s; 16,377 MiB |

- **Ollama 0.34.4's second start on the bench box.** Its loading request read `load_duration` 8,235.897 ms by
  Ollama's own timer, and 8.62 s from start to the first answer (`engines[1].load`). The rows label it "page cache
  warm (PLAN-v2 §2.7 M/K): start to first answer, n = 1". The first D's, under the same label, read 38,899.497 ms and
  68.438 s (`../dtt/rows/D.json` `engines[1].load`). These rows do not say why the second start was faster.
- **"Grown at rest"** is the growth of the card's used memory after each boot and its load: the median of three reads
  (`engines[].memory`), over a 1.0 MiB baseline. It is not memory under load.

## The other readings

- **Render parity (R-P1, reported, never a refusal).** Three renders give the same IDs on all four prompts: vLLM's
  chat render (`/tokenize` with messages), Ollama's `/api/chat` with `think: false`, and Ollama's
  `/v1/chat/completions` with `reasoning_effort: "none"`. On `series-p512` that is 548 IDs, sha `6981bc03…`.
  llama-server's `/apply-template` renders ChatML, because Ollama's runner argv carries `--no-jinja --chat-template
  chatml`. It reads 558 IDs on `series-p512` and differs at index 1 on every prompt (`d.render_parity`;
  `engines[2].readback.chat_template_note`). That is the receipt's `render=differs:llama-server`.
- **Finish reasons.** These are one run per engine × prompt, with no warm-up: the first requests after each boot, and
  never a speed row.
  - `series-p512`, `long-write` and `long-prefill` ended on `length` at 256 tokens on all three engines.
  - `short-chat` stopped early on all three, at 85, 91 and 91 tokens, and each such run carries its finish flag. So
    `short-chat` carries no headline field (`d.finish`, `d.headline_eligible`).
- **The resolved samplers (N-9).**
  - vLLM's default-sampling line names the model's `generation_config.json` (`temperature` 1.0, `top_k` 64, `top_p`
    0.95). Every request overrides it with temperature 0 and seed 0.
  - Ollama's runner `/slots`, after a D request, reads temperature 0.0, seed 0, `top_k` 64 and `top_p`
    0.949999988079071 (`d.samplers`).

  Temperature 0 is greedy on both.
- **The 16-slot runner, read back** (`engines[1].readback.internals.llama_context`): 16 sequences, a context of 81,920
  tokens (5,120 per slot), batch and micro-batch 512, flash attention `auto`, and `kv_unified` false.
- **llama-server's cache defaults at b11081 (A-13)**, from the startup lines, `--help` and `/props` in the rows
  (`engines[2].readback.cache_defaults`):
  - "prompt cache is enabled, size limit: 8192 MiB";
  - "context checkpoints enabled, max = 32, min spacing = 8192";
  - "idle slots will be saved to prompt cache upon starting a new task";
  - "KV cache shifting is not supported for this context";
  - `--swa-full` false, and 16 slots of 5,120 tokens.
- **The phone-home switches, read back from each engine's `/proc`** (`engines[].readback.env_serving`):
  - vLLM: `HF_HUB_OFFLINE`, `TRANSFORMERS_OFFLINE`, `VLLM_NO_USAGE_STATS`, `ORT_DISABLE_TELEMETRY`,
    `HF_HUB_DISABLE_TELEMETRY` and `DO_NOT_TRACK`, all 1.
  - Ollama: `OLLAMA_NO_CLOUD`, `ORT_DISABLE_TELEMETRY`, `HF_HUB_DISABLE_TELEMETRY` and `DO_NOT_TRACK`, all 1. Its
    serve log reads `Ollama cloud disabled: true` (the private half).

  The kit's clean environment dropped the unit's own switches from both engines (`inherited_env_dropped`), as
  Amendment 2 item 8 registered.

## What this does not say

D is the whole path run once, not a sweep. Its finish reads are one run per engine × prompt, and never a speed row.
TT reads level 1 on `series-p512` only. The boot readings (launch to ready, memory at rest and Ollama's start) are
n = 1 per engine, and they are not a comparison. P-BOOT and C-2 are one boot line and one memory read per run, under
two compile states, and the article prints both. No request figure here compares vLLM with Ollama. Nothing here says
why a warm compile cache lowers vLLM's profiled peak, or why Ollama's second start was faster.

## Files

| File | Bytes | sha256 |
|---|---:|---|
| `rows/D.json` | 1,066,094 | `withheld-2026-10-07` |
| `rows/D.power.csv` | 80,758 | `dbcd47d2e5c934ac922fedb59a625dd7d14185f092aa22c5eaf40d1df5d7ad55` |
| `rows/parity.json` | 58,376 | `3b0b9b436aaff34c3e76ee1f05040d26c799058b4a8bc81dbfeb3c24df6dcd9a` |
| `rows/servelines.json` | 61,817 | `withheld-2026-10-07` |
| `rows/RECEIPT` | 313 | `3cdfb4a59db31dbdc78b88c7feb30f231f48cd7c89e97aa73a6ba6aeeb84ca62` |

- `rows/` is the kit's outward output, every byte through its pen: a card label, never a UUID or a bus id; box-free
  run ids; UTC. `withheld` is empty in every JSON file.
- Each file is byte-equal to the run's archived copy, which read sha256-equal to the bench box's at the day's close.
- The pen's own final scan (`pen.check`, with the card's UUID loaded) reads 0 hits on all five files.
- The run's private half stays off the repo, archived on the operator's laptop (Amendment 4 item 10). It holds the
  UUID, the raw serve logs, `/proc` argv and environments, the stop guard's own record and the unscrubbed twins.

## Provenance: the figures a page may carry

Each line cites the value itself, in the file as committed. The sha256 of each file is in Files above; `../dtt/rows/D.json`
reads `withheld-2026-10-07`, as the kit's `prereg.TT_OF_RECORD` registers it.

| Figure | Value | File | Field | Line |
|---|---|---|---|---:|
| The window opens | `2026-09-30T20:00:02Z` | `rows/D.json` | `utc_start` | l.14 |
| The window closes | `2026-09-30T20:07:18Z` | `rows/D.json` | `utc_end` | l.15 |
| The first recorded run | `2026-09-30T20:01:26Z` | `rows/D.json` | `cells[0].runs[0].utc` | l.1574 |
| The receipt line | `RECEIPT ovv D ovv-20260930T200002Z-D …` (the whole line is above) | `rows/D.json` | `receipt` | l.29215 |
| TT, the replication | `CONFIRMED` | `rows/D.json` | `verdicts[0].verdict` | l.28844 |
| TT, road (a)'s `load_duration`, min, ms | `1.31` | `rows/D.json` | `verdicts[0].min` | l.28866 |
| TT, road (a)'s `load_duration`, max, ms | `3.119` | `rows/D.json` | `verdicts[0].max` | l.28867 |
| TT, the reading of record | `max 3.111 <= 50` | `../dtt/rows/D.json` | `verdicts[0].why` | l.28623 |
| TT, road (a)'s first token, median, ms | `445.128` | `rows/D.json` | `cells[8].medians.ttft_ms_p50.median` | l.8813 |
| TT, road (b)'s first token, median, ms | `486.951` | `rows/D.json` | `cells[9].medians.ttft_ms_p50.median` | l.12249 |
| TT, road (e)'s first token, median, ms | `441.351` | `rows/D.json` | `cells[16].medians.ttft_ms_p50.median` | l.28738 |
| TT, road (d): VOID, its kind | `hidden_reasoning` | `rows/D.json` | `cells[10].voids[0].fail_kind` | l.15608 |
| TT, road (c): VOID, its kind | `reload` | `rows/D.json` | `cells[11].voids[0].fail_kind` | l.19084 |
| TT, road (c)'s reading | `re-planned` | `rows/D.json` | `d.tt.replan_c` | l.30716 |
| P-BOOT, warm: the verdict | `CONFIRMED` | `rows/D.json` | `verdicts[1].verdict` | l.28871 |
| P-BOOT, warm: the boot line's concurrency, x | `17.07` | `rows/D.json` | `verdicts[1].value` | l.28880 |
| P-BOOT, cold: the verdict | `REFUTED` | `../dtt/rows/D.json` | `verdicts[1].verdict` | l.28649 |
| P-BOOT, cold: the boot line's concurrency, x | `14.78` | `../dtt/rows/D.json` | `verdicts[1].value` | l.28658 |
| C-2, warm: the verdict | `CONFIRMED` | `rows/D.json` | `verdicts[2].verdict` | l.28885 |
| C-2, warm: the share of the card | `0.9157` | `rows/D.json` | `verdicts[2].value` | l.28894 |
| C-2, warm: vLLM grown at rest, MiB | `22505.0` | `rows/D.json` | `engines[0].memory.at_rest_mib` | l.665 |
| C-2, cold: the verdict | `REFUTED` | `../dtt/rows/D.json` | `verdicts[2].verdict` | l.28663 |
| C-2, cold: the share of the card | `0.8483` | `../dtt/rows/D.json` | `verdicts[2].value` | l.28672 |
| C-2, cold: vLLM grown at rest, MiB | `20847.0` | `../dtt/rows/D.json` | `engines[0].memory.at_rest_mib` | l.714 |
| The card's memory, MiB (R-1) | `24576.0` | `rows/D.json` | `gates[4].reading.boards["EVGA GeForce RTX 3090 24 GB"]["memory.total"]` | l.170 |
| vLLM, warm: the KV pool, GiB | `12.09` | `rows/D.json` | `engines[0].readback.boot.kv_pool_gib` | l.595 |
| vLLM, warm: KV tokens | `87407` | `rows/D.json` | `engines[0].readback.kv_cache_tokens` | l.420 |
| vLLM, warm: KV blocks | `12377` | `rows/D.json` | `engines[0].readback.cache_config.num_gpu_blocks` | l.443 |
| vLLM, warm: the profile's peak activation, GiB | `0.85` | `rows/D.json` | `engines[0].readback.boot.peak_activation_gib` | l.589 |
| vLLM, cold: the KV pool, GiB | `10.47` | `../dtt/rows/D.json` | `engines[0].readback.boot.kv_pool_gib` | l.616 |
| vLLM, cold: KV tokens | `75691` | `../dtt/rows/D.json` | `engines[0].readback.kv_cache_tokens` | l.442 |
| vLLM, cold: KV blocks | `10718` | `../dtt/rows/D.json` | `engines[0].readback.cache_config.num_gpu_blocks` | l.465 |
| vLLM, cold: the profile's peak activation, GiB | `2.47` | `../dtt/rows/D.json` | `engines[0].readback.boot.peak_activation_gib` | l.610 |
| vLLM, warm: launch to ready, s | `53.79` | `rows/D.json` | `engines[0].boot_s` | l.356 |
| vLLM, cold: launch to ready, s | `144.35` | `../dtt/rows/D.json` | `engines[0].boot_s` | l.350 |
| Ollama 0.34.4, second start: launch to ready, s | `2.86` | `rows/D.json` | `engines[1].boot_s` | l.753 |
| Ollama 0.34.4, second start: `load_duration`, ms | `8235.897` | `rows/D.json` | `engines[1].load.load_duration_ms` | l.1181 |
| Ollama 0.34.4, second start: start to the first answer, s | `8.62` | `rows/D.json` | `engines[1].load.wall_to_first_answer_s` | l.1180 |
| Ollama 0.34.4, second start: its label | `page cache warm (PLAN-v2 §2.7 M/K): start to first answer, n = 1` | `rows/D.json` | `engines[1].load.label` | l.1184 |
| Ollama 0.34.4, first start: `load_duration`, ms | `38899.497` | `../dtt/rows/D.json` | `engines[1].load.load_duration_ms` | l.1227 |
| Ollama 0.34.4, first start: start to the first answer, s | `68.438` | `../dtt/rows/D.json` | `engines[1].load.wall_to_first_answer_s` | l.1226 |
| Ollama 16 slots, grown at rest, MiB | `16389.0` | `rows/D.json` | `engines[1].memory.at_rest_mib` | l.1187 |
| llama-server, grown at rest, MiB | `16377.0` | `rows/D.json` | `engines[2].memory.at_rest_mib` | l.1547 |
