# EmbeddingGemma 2's shared space for text, images and audio, measured on what we publish

*EmbeddingGemma 2 is Google's open embedding model, released 2026-10-06: it turns text, images and audio into one kind of vector, a list of numbers that places similar things near each other. Between 22:42Z on 2026-10-06 and 04:56Z on 2026-10-07 (UTC) we ran ten tests on it, on one laptop's CPU and no graphics card, every rule and pass mark written down before the first score it governs, using our own published photographs, paintings, screenshots, code and music and one public image set, against the Nomic AI pair, the text model our stores (the search indexes our products keep) are built on and the image model one product's retrieval code is wired to. On images EmbeddingGemma 2 is ahead: on 1,000 public captioned photographs it puts the right one first for 729 captions against the Nomic pair's 497, a net 232, clearing the 100-in-1,000 pass mark we set for replacing a model, on a set both models have most likely trained on; on our own 14 photographs it leads by more, too few to count as more than exploratory. On code, text only, it is ahead by 50 of 316 queries with the code-retrieval prefix, a gain that depends on that prefix. On audio (exploratory, with no incumbent to beat) it tracks how strongly a music model's adapter, its add-on weights, was applied, and a loudness change barely moves it, but it cannot match a clip to the words that requested it. The Nomic image model ran on an older library release (transformers 4.57.6), because our own product, [RuleSage](https://rulesage.strata2signal.com/), pins a release that refuses to load it: that is our product's defect, recorded for its own fix. Nothing in our stores changes; we estimate what switching would cost in CPU time per item, and print no total because the largest store's size in tokens was not read.*

*Published 2026-10-09 (UTC) · A small (human) team and a fleet of AI agents.*

**the short version:** EmbeddingGemma 2, written EG2 in this page's verdict words, turns text, code, images, video and audio (per its model card) into one 768-number space. Our stores are built on Nomic AI's text model, and one product's retrieval code is wired to its image partner; the companion page, "EmbeddingGemma 2, measured on release day against the embedder we already run", found NO MATERIAL DIFFERENCE between EmbeddingGemma 2 and nomic's text model on our own text, so this bench tested images, audio and code.

12,013 words · about 55 minutes (at 220 words/min) · 11 tables · data kit: no

https://research.strata2signal.com/embeddinggemma-2-shared-space/

---

- **Public photographs.** On 1,000 from COCO, a widely used public set of captioned photographs, EmbeddingGemma 2 puts the right one first for 729 captions, the nomic pair (on transformers 4.57.6) for 497: **EG2 AHEAD**, with **CLEARS THE HOUSE BAR** (our 100-in-1,000 pass mark for replacing a model). Both models have most likely trained on pictures of this kind (COCO has been public since 2014), so this is weaker evidence than a lead on our own.
- **Our 14 photographs.** The right one first for 33 of 40 published descriptions against the pair's 13 (nomic on transformers 4.57.6; 79 % against 27 % with each of the 12 scenes weighted equally, as registered; chance 8 %): exploratory-only, about these photographs only.
- **The nomic image model** loads neither under our product's pinned library (RuleSage, transformers 5.12.1) nor under any 5.x release we tried; it ran on transformers 4.57.6, and every nomic image figure here says so.
- **One list for images and text** works for EmbeddingGemma 2 beside unrelated text, for neither model beside related text, and never for the nomic pair, whose text always ranks first: if you keep nomic's text and image vectors in one index, its images never surface.
- **Code, text only.** The right function first for 173 of 316 docstring queries against nomic's 123 (a docstring is the description written at the top of a function; its first paragraph is the query): **EG2 AHEAD ON CODE**, a gain that depends on the code-retrieval prefix (a short fixed string the model expects in front of a text).
- **Audio** (exploratory; nomic has no audio model): a loudness change alone (each mastered copy is the same render with one gain change) barely moves the vector, and applying the music model's adapter (its add-on weights) more strongly moves it in order, more steadily than the root-mean-square difference between two clips' sound samples (the published ruler this test was registered against), but it cannot tell which of 12 requests made a clip (1 of 12, chance; each label is the request that made it).

We are not switching: no material difference on our text, no code index of ours for the code win, no working image search of ours for the image win, and a provenance guard (a check that ties every stored vector to the exact model files that made it) comes first.

*If you're new here: [strata→signal](https://strata2signal.com) is a small workshop that builds on its own machines and writes up what it measures. Every bench figure below was read from the result file named beside it, as two independent verifies left them on 2026-10-07 (UTC); the store counts come from a dated survey. Machines are named by their role, never their hostname; a number we could not source is a dash with its reason. Each test's result is a phrase in capitals (EG2 AHEAD, NEAR CHANCE, ...) fixed before the test ran; we print the one the rule selects exactly as registered, even where it reads oddly, with the measurement beside it.*

*Words this page leans on.* An **embedding model** turns an input into a list of numbers (a **vector**; every model here writes 768), so that inputs about the same thing get similar lists; a **shared multimodal space** is one model doing that for text, images and audio alike. **recall@1** is the share of queries whose correct answer ranks first; **macro over groups** counts each scene once however many descriptions it has; **MRR** (mean reciprocal rank) averages one over the rank of the first correct answer. **Cosine** measures how closely two vectors point the same way (1 is identical, near 0 unrelated; "1 − cosine" is a distance). A **bootstrap interval** shows how far a difference could move if the test's items (images, or scenes) were redrawn; a **sign test** counts the scenes that went up against those that went down. **McNemar's exact test** compares two models on the same questions, using only those one finds and the other misses. **N\*** is the better-scoring of the incumbent's two recipes, the one the challenger must beat; **Δ\*** is EmbeddingGemma 2's recall@1 minus N\*'s. A **pre-registration** is the test, the thresholds and the words for each outcome, committed before the first score; a **gate** is one of its pass marks. A **prefix** is a short fixed string these models expect in front of a text; images and audio take none. The **incumbent** is the model our stores use now, or, for images, the one our product's code is wired to; **transformers** is the Python library that loads these models. **EXPLORATORY** marks a number, or a word, that decides nothing: a large-looking exploratory result is a reason for a new registration, never a finding of this one.

## Where this page starts from {#where-this-page-starts}

EmbeddingGemma 2's claim is one checkpoint with three towers (sub-networks, one per kind of input: a 270 M-parameter text model, 170 M for vision, 300 M for audio) and one 768-number output, so that a sentence and the photograph it describes can be searched in one list. Our stores are built on the text half of the alternative, a *pair* of models trained to share a space, nomic-embed-text-v1.5 and nomic-embed-vision-v1.5 from Nomic AI, with no audio model; one product's retrieval code is wired to its image half. [The companion page](https://research.strata2signal.com/embeddinggemma-2-on-release-day/) measured the text model on release day: 68 questions over 2,569 passages of our own, nomic-embed-text 44 hits in the top 8, EmbeddingGemma 2 45 against a pass mark of 51, the registered verdict NO MATERIAL DIFFERENCE, read as "no effect of at least 10 points was detected at n = 68". It also found Google's prefixes worth nine questions in an exploratory row, the 8-bit GGUF file on llama.cpp at parity with Google's code, and the Ollama release on our bench laptop, 0.32.14, unable to pull the model. This page is the follow-up, registered as bench E-3, with the same rules and one sentence more: **E-3 is exploratory-only by registration; its strongest outcome licenses drafting a larger confirmatory bench, nothing more.**

## What we tested, and why these tests {#what-we-tested}

**The core, frozen first.** By the freeze (2026-10-06 22:42Z), the workshop had released 14 photograph files on four pages ([GPUs and orchids](https://research.strata2signal.com/gpus-and-orchids/), [What it costs to run a high-end AI rig at home](https://research.strata2signal.com/what-it-costs-to-run/), [One 3090 out, one 3090 Ti in](https://research.strata2signal.com/one-3090-out-one-3090-ti-in/) and [The same eighteen pictures, card by card](https://research.strata2signal.com/the-same-eighteen-pictures-card-by-card/)): 12 distinct scenes and 40 distinct published texts (16 alt texts, 15 captions, 9 thumbnail descriptions), every one checked live on the public site at the freeze. The set is hard on purpose: the same graphics cards recur across scenes, and all seven orchid photographs share one orchid, one pot and one wall. Twelve scenes are too few for a confirmatory gate, so the core registers exploratory-only verdict words, and a power clause printed beside every verdict:

> E-3 is exploratory-only by registration: 14 photographs in 12 scenes, 40 published texts. At 12 groups an exact sign test reaches p ≤ 0.05 only when at least 6 scenes differ and all in one direction. Its words describe these photographs, never images in general, and nothing is adopted from them.

**The incumbent, as the product ships it and as its model cards write it.** The pair's text and image models are the ones [RuleSage](https://rulesage-live.strata2signal.com/)'s retrieval code (RuleSage is our rules-question assistant) is wired to use for text and for pictures, and the bench imported that product's own loading and prefix code unchanged. Two recipes were registered, because the product and the model cards disagree: the product embeds a text query for image search with the *document* prefix (`search_document: `) and no layer norm; the model cards say `search_query: ` with a layer norm (a rescaling step applied to the model's output). The challenger had to beat whichever scored higher (N\*).

**The challenger.** EmbeddingGemma 2 through Google's own code (sentence-transformers 6.1.0 over transformers 5.19.0), 32-bit floats, text and vision towers loaded, the model card's default budget of 280 image tokens (the model card allows 70 to 1,120); text queries carry the model card's search prefix (`task: search result | query: `), images take none. Exploratory variants: no prefix, the document prefix on queries, 70 and 1,120 image tokens, the vector cut to 256 numbers, and the GGUF files (the single-file model format of llama.cpp, an open-source model runner) on a llama.cpp server with the vision projector (the file that turns an image into the model's input), the text file in 16-bit (BF16) and 8-bit (Q8_0) form, behind a parity gate.

**Nine further tests, registered as an addendum before any of their scores,** chosen by a design panel of two AI agents and ranked by information per CPU-minute, 240 minutes allowed:

| # | Test | What it asks | Set | Class |
|---|---|---|---|---|
| T1 | One index | In one list of images *and* texts, does a text query still reach its image? | The core; 1,000 COCO images; our 2,569 text passages | Words per model; gated comparison |
| T2 | Hard negatives | H1: does a text beat photographs of the same graphics card in another scene? H2: does a photograph beat a copy of its text with another graphics card's name swapped in? | The core; 169 swapped variants | H1 no words; H2 EXPLORATORY words |
| T3 | A public set | Image-text retrieval at a size with statistical power | COCO 2014, one 1,000-image fold of its standard (Karpathy) test split | Gated, margin 0.050 at n = 1,000 |
| T4 | Three product readings | Screenshots; "did the painting take the world's name?"; an alt-text lint | 23 released non-photograph images | EXPLORATORY words |
| T5 | Code | Docstring → function retrieval on our released kit code | 746 functions, 316 queries | Gated at +0.100, TEXT-ONLY |
| T6 | Audio | Does the audio tower track effect strength, unmoved by loudness, better than the published root-mean-square ruler; does a clip find its request? | 133 released music clips | EXPLORATORY, no incumbent |
| T7 | Speed and memory | CPU seconds per image, per 1,000 text tokens, per 10 s of audio | 30 items per configuration (10 at 1,120 tokens) | Measurement, no words |
| T8 | Projector quantisation | The 8-bit vision projector against the 16-bit one, on llama.cpp | 237 images | EXPLORATORY words |
| T9 | The text quant ladder | Three smaller GGUF files against the reference, then against the 68-question bench | 268 texts; then the full sample | EXPLORATORY words, TEXT-ONLY |

Among those dropped by name: the Flickr30k test split (its terms allow non-commercial research and education only), a permissive-licence-only COCO subset, any video, and any Ollama arm (0.32.14 refuses the model).

**Two pass marks appear below, and they are not the same.** A test's registered *word* (EG2 AHEAD, say) needs a lead the test can resolve: on the 1,000-image public set that is +0.050 of recall@1, one query in twenty, which n = 1,000 can detect. The *house bar*, +0.100, is the margin we set before any score for replacing a model; on the public set it prints beside the word, never instead of it, and the code test is gated at the house bar alone. Our own photographs use the +0.100 margin for their exploratory words and carry no gate: at 12 scenes they are exploratory-only by registration.

**The timeline, from the commit and push logs.** The core registration was committed at **2026-10-06 22:42:12Z**; its first vector at 23:33:14Z, its first score at 23:33:53Z. The further-tests addendum was committed at **23:52:00Z**; its first vector at 00:12:03Z on 2026-10-07, its scorer pinned at 00:32:42Z, its first score at 00:32:55Z, its results at 02:05:06Z. The nomic image arms and the swap test waited on two decisions, recorded at **03:42:48Z**: an addendum quoting them was committed at 04:10:44Z, the first second-round vector came at 04:11:09Z, its results at 04:26:40Z. The body of the registration never changed after the freeze; six dated addenda append the further tests, receipts, the two decisions and the pins. The further tests used 119.47 of their 240 unit-minutes (a unit is one isolated job on the laptop; the budget counts each one's minutes).

**The machine.** One laptop's CPU, an Intel Core Ultra 9 290HX Plus (24 logical processors), 8 threads per path. The laptop has a GPU and the bench never used it: both PyTorch builds are CPU-only wheels, the llama.cpp build links no GPU library, every unit's environment hid the GPU, and the GPU's process list was read empty at the start and end of every step. Every unit that executed the incumbent's downloaded model code ran with its network taken away (trap 7).

## The incumbent's image model: which library, and why {#which-library}

**What failed.** Before the freeze, the preparation step tried to load nomic-embed-vision-v1.5, unpatched, under the library release our own product pins (RuleSage, transformers 5.12.1). It refused with `StrictDataclassFieldValidationError`: the model's configuration carries `n_inner` as `2048.0`, and 5.x's strict validation wants an integer. The registration allowed six steps down the release ladder (5.11.0 to 5.6.2); every step refused the same way, and every 5.x release down to 5.0.0 refused too (5.0 to 5.3 with a different error, `'NomicVisionModel' object has no attribute 'all_tied_weights_keys'`). One older release, **transformers 4.57.6, loads both nomic towers unpatched**.

**What ran instead, and when that was decided.** The core's registered verdict reads **NOT EVALUATED: NO INCUMBENT** and stays so: by the registration's own rule, no addendum can change the core's words once its arms are declared NOT RUN. But the further-tests addendum, committed at 23:52:00Z on 2026-10-06, after the core's first EmbeddingGemma 2 score (23:33:53Z) and before any score of its own tests or any nomic image vector, had registered a pair of arms for exactly this case: the two nomic recipes on transformers 4.57.6, as the nomic side of each test's registered words (gated where the test is gated, exploratory on our own photographs), every number labelled "on transformers 4.57.6" because the product's own library pin cannot load this model. The decision to admit those arms was recorded at 03:42:48Z on 2026-10-07: after the first round's EmbeddingGemma 2 scores existed, and before any nomic image vector did. They then ran as registered.

**What the library change does and does not change.** A check first: 16 synthetic texts, each through both recipes, went through the product's text loader under 5.12.1 and under 4.57.6 and gave 32 of 32 vectors bit-identical, so on the text side the library choice changes nothing. Whether it changes the image vectors cannot be checked, because no release the product pins loads the image tower. So **this page does compare EmbeddingGemma 2's images with nomic's**: on the public set with gated words, and on our photographs as an exploratory reading beside a frozen NOT EVALUATED.

**The defect is ours.** The pair's image path, as RuleSage pins it, does not load as of 2026-10-06, and as of 2026-10-07 no product of ours sends it a picture. That is a defect in our product, recorded for its own repair and being routed to it; it is separate from this bench and, as of 2026-10-08, not yet fixed. One more thing bears on every nomic image figure below: the image processor set in nomic-embed-vision-v1.5's own files, which the product's loader uses unchanged, resizes every image straight to 224 × 224, dropping its proportions, so every image reaches the model as a stretched 224 × 224 square (the second verify, described below, fed it pre-squashed copies and got bit-identical vectors on 8 of 8 images). What the squash costs nomic was not measured; a padded variant is registered and did not run.

## What the images show {#what-the-images-show}

**The 14 photographs (the core; `e3-core-scores.json`, `r2-CORE.json`).** When this test was registered, the incumbent's image model could not be loaded under our product's own library pin, so its registered words are **NOT EVALUATED: NO INCUMBENT** for text → photograph and **NOT EVALUATED** for photograph → text, and stay that way. A backup pair of arms, registered after EmbeddingGemma 2's core scores and before any nomic score, ran nomic on the older library (transformers 4.57.6, above); read with the registered rule, EmbeddingGemma 2 is ahead on these 14 photographs, an exploratory reading only, under the power clause above: 12 scenes, these photographs only. Chance is 0.0833 in both directions.

| Arm | Direction | recall@1, macro over 12 scenes | recall@5 | MRR | recall@1 over all queries |
|---|---|---:|---:|---:|---:|
| **EmbeddingGemma 2, 280 tokens** | Text → photograph (40 texts rank 14 images) | **0.791667** | 0.972222 | 0.847917 | 0.825 (33 of 40) |
| | Photograph → text (14 images rank 40 texts) | **0.750000** | 1.000000 | 0.861111 | 0.785714 (11 of 14) |
| **nomic, the model-card recipe (N\*)†** | Text → photograph | **0.272222** | 0.538889 | 0.411013 | 0.325 (13 of 40) |
| | Photograph → text | **0.416667** | 0.666667 | 0.566604 | 0.500 (7 of 14) |
| `nomic`, the product's recipe† | Text → photograph | 0.208333 | 0.372222 | 0.354912 | 0.275 (11 of 40) |
| | Photograph → text | 0.416667 | 0.666667 | 0.545378 | 0.500 (7 of 14) |

† nomic on transformers 4.57.6, a library release our own product does not pin, because the release it pins cannot load nomic's image model (see "The incumbent's image model" above). Every nomic image figure on this page carries this mark.

**The exploratory reading.** Beside the frozen words, the registered rule read on the 4.57.6 arms gives **EG2 AHEAD ON THIS SET** in both directions: text → photograph Δ\* +0.519444 (cluster bootstrap over the 12 scenes 0.305556 to 0.736111; scenes up 9, down 0, 3 ties, sign test p = 0.003906); photograph → text Δ\* +0.333333 (4 scenes up, 0 down, p = 0.125). The model-card recipe over the product's is +0.063889 (2 scenes up, 1 down, p = 1.0).

**By kind of text** (text → photograph, macro over scenes; EmbeddingGemma 2's count of texts ranked first beside it; the counts are per text, so they do not divide to the macro figures):

| Kind of text | EmbeddingGemma 2, 280 tokens | `nomic`, the model-card recipe† | `nomic`, the product's recipe† |
|---|---:|---:|---:|
| Alt texts (16) | 0.916667 (15 of 16 ranked first) | 0.333333 | 0.250000 |
| Captions (15) | 0.666667 (11 of 15) | 0.270833 | 0.187500 |
| Thumbnail descriptions (9) | 0.750000 (7 of 9) | 0.062500 | 0.125000 |

**The misses.** EmbeddingGemma 2's seven: the photograph of two 3080 Ti graphics cards edge-on puts the neighbouring photograph of two cards fan-side first for all three of its texts; two orchid captions rank their photograph 4th and 12th, an orchid thumbnail description 2nd, and one graphics-card caption 5th. Photograph → text misses three times, all orchids, among seven near-identical shots of one plant; both nomic recipes miss the same three orchid photographs, and four of the seven photographs that are not orchids besides (7 of 14 in all).

The exploratory variants of EmbeddingGemma 2 (`arms/*.score.json`), each against the reference, deciding nothing:

| Variant | Text → photograph, macro recall@1 | Photograph → text | Δ vs reference (T2I) | Scenes up / down (p) |
|---|---:|---:|---:|---|
| Reference, 280 tokens | 0.791667 | 0.750000 | (itself) | (no comparison) |
| No prefix on either side | 0.763889 | **0.916667** | −0.027778 | 1 / 2 (1.0) |
| Document prefix on queries | 0.727778 | 0.750000 | −0.063889 | 0 / 2 (0.5) |
| 70 image tokens | 0.527778 | 0.666667 | **−0.263889** | 2 / 5 (0.453) |
| 1,120 image tokens | 0.766667 | 0.750000 | −0.025000 | 1 / 2 (1.0) |
| 256 numbers (cut, re-normalised) | 0.791667 | 0.666667 | 0.000000 | 0 / 0 (1.0) |
| llama.cpp, BF16 or Q8_0 text + BF16 projector | VOID by the parity gate | | | |

Three readings, none a verdict. Cutting the budget to 70 tokens cuts macro recall@1 by 0.264 (5 scenes down, 2 up); 1,120 buys nothing here (prediction Q-4, that more tokens would read graphics-card lettering in captions better, was REFUTED: captions only, 0.645833 against 0.666667). The 256-number cut is free in one direction and costs one scene in the other. A bare `encode()` with no prefix did *better* photograph → text (11 of 12 scenes against 9): a 12-scene reading whose interval (0.000 to +0.417) touches zero.

**The llama.cpp image path is VOID.** The server, built at the commit that merged the model, does embed images with the vision projector. Its text vectors agree with Google's code at cosine ≥ 0.999993 (BF16) and ≥ 0.999717 (Q8_0) on all 40 texts; its image vectors do not: minimum 0.978505 with the 16-bit text file and 0.978162 with the 8-bit one (both through the 16-bit projector), 13 of the 14 photographs below the 0.999 floor. The token count is not the cause; the likeliest is the server's image preprocessing (a different patch grid; not traced). A VOID is a runtime fault, never a loss for the model. The registered prediction that the server's image side would agree less than its text side (Q-5) reads UNTESTED, because the registered rule makes a VOID arm's prediction untested; the first round's scorer had printed HELD on these readings, and a decision on 2026-10-07 corrected it to the rule.

**1,000 public images (T3; `a2-X.json`, `r2-X.json`).** COCO 2014's standard test split (Karpathy's: 5,000 images in five 1,000-image folds), the first fold from a public mirror (the Hugging Face dataset `undefined443/coco-karpathy-wds`, where `undefined443` is the uploading account's name as written, revision `cfebcb4f…`, shard sha256 `2c0f3695…`; one fold, and one caption per image where published results use all five, so not comparable with published averages), each image with its first caption as the query. COCO's terms were read first-hand (annotations CC BY 4.0; the images under the Flickr terms, for evaluation only, never redistributed). Chance 0.001.

| Arm | Caption → image recall@1 | @5 | @10 | MRR@10 | Image → caption recall@1 | @5 | @10 | MRR@10 | Median image tokens |
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|
| **EmbeddingGemma 2, 280 tokens** | **0.729** | 0.920 | 0.967 | 0.811660 | **0.698** | 0.905 | 0.961 | 0.786328 | 264 |
| **nomic, the model-card recipe (N\*)†** | **0.497** | 0.780 | 0.889 | 0.618181 | **0.507** | 0.784 | 0.876 | 0.624562 | (224 × 224) |
| `nomic`, the product's recipe† | 0.464 | 0.742 | 0.852 | 0.580957 | 0.450 | 0.757 | 0.853 | 0.577644 | (224 × 224) |
| EmbeddingGemma 2, 70 tokens (EXPLORATORY) | 0.699 | 0.905 | 0.958 | 0.788366 | 0.674 | 0.907 | 0.956 | 0.771783 | 63 |

† nomic on transformers 4.57.6, outside our product's pin (see "The incumbent's image model").

**Words: EG2 AHEAD**, primary and secondary, with **CLEARS THE HOUSE BAR** printed beside the primary word, never instead of it. Caption → image: Δ\* = +0.232, 232 of 1,000 captions, against the +0.050 the word needs and the +0.100 the house bar needs; 276 captions only EmbeddingGemma 2 finds against 44 only nomic finds, McNemar exact two-sided p = 3.2 × 10⁻⁴² (the file rounds it to 0.0; the verify's exact binomial), image bootstrap 0.200 to 0.264. Image → caption: Δ\* = +0.191, 250 against 59, p = 3.9 × 10⁻²⁹, bootstrap 0.158 to 0.223. Against the product's own recipe the gap is wider (306 against 41, +0.265). Prediction X-1 (that EmbeddingGemma 2 would read EG2 AHEAD here) HELD. The 70-token budget costs 0.030 of recall@1 (45 captions only it finds, 75 only the reference finds, McNemar p = 0.0078), within the 0.050 prediction X-2 allowed, so X-2 HELD. COCO's images have been public since 2014, so they most likely sit inside both models' training distributions (we have not confirmed this from either model's documentation), and, as the registration notes, EmbeddingGemma 2's model card reports results on MMEB v2, a benchmark whose training split includes COCO retrieval (whether EmbeddingGemma 2 trained on that split is not stated, and we did not verify it). So a lead here is weaker evidence than a lead on our own photographs, and this set alone cannot tell a better model from one that has seen more pictures like these.

**One index for images and text (T1; `a2-U.json`, `r2-U.json`).** The 1,000 COCO images and our 2,569 text passages (about a historical town, unrelated to any photograph) go into one list, and each caption asks whether its image still comes first.

| Variant | Pool | Mixed recall@1 | Image-only recall@1 | Crowding loss | Top item is an image | Words |
|---|---|---:|---:|---:|---:|---|
| **EmbeddingGemma 2, off-topic text** | 1,000 images + 2,569 passages | **0.721** | 0.729 | **0.008** | 0.975 | **ONE SPACE HOLDS** |
| EmbeddingGemma 2, on-topic text | + the other 999 captions as documents | 0.139 | 0.729 | 0.590 | 0.139 | (no words) |
| **nomic, the model-card recipe, off-topic†** | 1,000 images + 2,569 passages | **0.000** | 0.497 | 0.497 | 0.000 | **TEXT CROWDS OUT IMAGES** |
| **nomic, the product's recipe, off-topic†** | The same | **0.000** | 0.464 | 0.464 | 0.000 | **TEXT CROWDS OUT IMAGES** |
| EmbeddingGemma 2, the core, off-topic | 37 images + 2,569 passages; 40 texts, macro over 12 scenes | 0.791667 | 0.791667 | 0.000 | 1.000 | (no words) |
| EmbeddingGemma 2, the core, on-topic | + the other scenes' texts and 39 alt texts and captions of the other released images | 0.211111 | 0.791667 | 0.580556 | 0.200 | (no words) |

† nomic on transformers 4.57.6, outside our product's pin (see "The incumbent's image model").

Read plainly: beside text about other things, EmbeddingGemma 2's images keep their place (8 captions in 1,000 lose their image to a passage); nomic's never surface, beside off-topic and on-topic text alike (its on-topic rows also read 0.000, so the table prints only its off-topic ones: adding text to a pool cannot lift an image that already ranks below every passage). For nomic, every passage outranks every image for every one of the 1,000 captions, because its image and text vectors share a direction but sit far apart: a caption's cosine to its own image averages 0.066 (product recipe) or 0.086 (model-card recipe), and to a random passage about 0.49. The gated comparison is **EG2 AHEAD IN ONE INDEX** (Δ\* +0.721, 721 captions against 0, p = 1.8 × 10⁻²¹⁷); predictions U-1 (that text would crowd out nomic's images) and U-2 (that EmbeddingGemma 2 would lose less to crowding than nomic) HELD. Beside text about the same things, EmbeddingGemma 2's caption usually finds another caption nearest, and only 139 of 1,000 still put their image first; the core shows the same shape. Part of the reason is the modality gap: on average a caption sits at cosine 0.743 to its own image, 0.535 to a random other image and 0.546 to a random passage (an average, not a per-query bar; how close captions sit to one another was not measured). Image-to-anything is worse: an image's own caption is rarely its nearest item (0.134 on COCO and 0.083 on the core for EmbeddingGemma 2; 0.000 for nomic). An exploratory per-modality z-score lifts nomic's off-topic figure to 0.323 and 0.443 and moves EmbeddingGemma 2's on-topic figure from 0.139 to 0.405 (off-topic 0.721 to 0.712); on our own 37-image pools the same z-score takes EmbeddingGemma 2's figure from 0.791667 to 0.000 off-topic and from 0.211111 to 0.000 on-topic, so it is no general fix. **So one model does not mean one ranked list.** A unified store would still filter by modality, or rank each modality and merge. The pools hold no duplicate images, and the metric would count a duplicate ranked first as a miss; a real store serves one photograph on several pages and would de-duplicate by content first.

**Hard negatives (T2; `a2-H1.json`, `r2-H.json`).** Every nomic figure here is on transformers 4.57.6 (†, above). H1: for each of the 40 texts, the photographs of other scenes whose texts name the same graphics card are its hard negatives. Share of (text, hard negative) pairs where the gold photograph scores higher, macro over the 12 scenes, chance 0.5, no words: EmbeddingGemma 2 **0.895833** at 280 tokens, the same at 1,120; nomic 0.596759 (model-card recipe) and 0.442130 (product recipe: below chance, so more often than not it ranks another scene's photograph of the same card above the right one). EmbeddingGemma 2's weakest scenes are the two 3080 Ti cards edge-on (0.416667) and the ZOTAC card beside the orchid (0.666667). H2, the swap test: 169 copies of the published texts, each with one graphics card's full name replaced by another's; it covers the nine scenes that have a text naming exactly one card by its full name (the other three scenes have no such text, so nothing to swap). A pair is won when a photograph scores its true text strictly above the swapped copy, averaged over variants and then over the 9 scenes. Words: EmbeddingGemma 2 **READS THE CARD** (macro win 0.844356; all 9 scenes above 0.5, one-sided sign test p = 0.001953; 0.829145 at 1,120 tokens); nomic **NEAR CHANCE** on both recipes (0.597443 with the product's recipe, 0.0026 under the 0.600 line where PARTIAL begins; 0.563492 with the model-card recipe). EmbeddingGemma 2 is weakest on two scenes from the orchid page, the ZOTAC card beside the orchid and the two cards fan-side (0.571429 each); nomic's product recipe wins 1 variant in 21 on one graphics-card scene. Predictions H-1 (that the better nomic recipe would read NEAR CHANCE) HELD; H-2 (that 1,120 image tokens would win at least as often as 280) REFUTED.

**Three product readings (T4; `a2-R.json`, `r2-R.json`; EXPLORATORY).** Twenty-three released images that are not photographs: seven interface screenshots ([How the long table works](https://research.strata2signal.com/how-the-long-table-works/), [Introducing the playground](https://research.strata2signal.com/introducing-the-playground/)), fifteen generated paintings ([Six Worlds for the Print Lab](https://research.strata2signal.com/six-worlds-painted-large/), [Nine Worlds, One Dog](https://research.strata2signal.com/nine-worlds-one-dog/)) and one illustration ([A Dinner Party for the Dead](https://research.strata2signal.com/a-dinner-party-for-the-dead/)), with 39 published texts, and the 15 world names (nine and six) on their own, cut from those texts, as a third kind of query.

| Reading | What it asks | EmbeddingGemma 2, 280 tokens | 70 | 1,120 | `nomic`, product recipe† | `nomic`, model-card recipe† | Words |
|---|---|---|---|---|---|---|---|
| Screenshots (7) | Alt text → its screenshot, among the 7 screenshots | 5 of 7 | 5 | **2** | 3 | 3 | None at n = 7 |
| | Alt text → its screenshot, among all 37 images (these 23 and the 14 photographs) | 5 of 7 | 5 | **2** | 3 | 3 | |
| | Screenshot → its alt text, among the 7 alt texts | 5 of 7 | 4 | 4 | 4 | 3 | |
| Nine worlds | Nine paintings of one subject, their alt texts differing only in the world's name: alt → painting / painting → alt (SKIN READ iff ≥ 8 of 9 both ways) | 4 of 9 / 7 of 9 | 3 / 5 | 4 / 5 | 2 / 4 | 1 / 3 | **SKIN NOT READ** (all five) |
| Nine worlds, by the world's name alone | The world's name as the whole query → painting | 6 of 9 | 4 | 5 | 5 | 5 | |
| Nine worlds, by caption | Captions that describe each scene: caption → painting / painting → caption | 9 / 9 | 9 / 9 | 9 / 9 | 6 / 6 | 7 / 6 | |
| Six worlds, by alt | Alt texts that describe each scene: alt → painting / painting → alt | 6 / 6 | 6 / 6 | 6 / 6 | 5 / 6 | 6 / 6 | |
| Alt-text lint | Each core photograph ranks the alt texts of its page; a scene is flagged when a photograph's top alt is not its own; 36 planted swaps of two scenes' alt texts; the 13 scenes on the two photograph pages (7 and 6; one scene appears on both) can each raise a false alarm (LINT READY iff 36 of 36 caught and 0 false alarms) | 33 of 36 caught; 3 false alarms of 13 | 32; 4 | 33; 3 | 28; 8 | 29; 7 | **LINT NOISY** (all five) |

† nomic on transformers 4.57.6, outside our product's pin (see "The incumbent's image model").

The nine-worlds reading is the sharpest. The nine alt texts read "the founder's dog asleep under a tree in …" with only the world's name changing, and given the nine alts EmbeddingGemma 2 finds the right painting for 4 of 9 (given a painting, it picks the right alt for 7 of 9; SKIN READ needed 8 of 9 both ways); given only the world's name it finds 6 of 9; given the captions, which describe each scene, it places all nine both ways. nomic reads the world's name back less often still. EmbeddingGemma 2's three lint misses are the orchid photographs that confuse the core. And 1,120 image tokens did *worse* on screenshots than 280 (2 of 7 against 5); predictions R-2 (more tokens help in-image text) and R-1 (the world's name is read) are REFUTED.

**The 8-bit vision projector (T8; `a2-MQ.json`).** Two llama.cpp servers, the same 16-bit text file, one with the 16-bit vision projector (982,074,880 bytes, sha256 `995aaa56…`) and one with the 8-bit projector (554,821,120 bytes, `90e7b023…`), embedded the same 237 images. What it measured: the 8-bit projector moved image vectors by at most 0.0008 of cosine (agreement between the two projectors' vectors: minimum **0.999208**, median 0.999859), and of the 200 COCO captions' top-1 results **198 of 200** were identical, the two that changed sitting on near ties (top-1 accuracy 0.895 with 16-bit, 0.900 with 8-bit, from the run's notes). The registered word, read literally, is **QUANT-HARMFUL**, and it never prints without this sentence beside it: TRANSPARENT required every top-1 identical, VISIBLE required a minimum below 0.999 with at least 95 % of top-1s unchanged, and this case falls between them into the third word (the word was registered before the score; a decision on 2026-10-07 ruled that it never prints without this sentence). Prediction MQ-1 (TRANSPARENT) is REFUTED. Both projectors still miss Google's code on the photographs (0.9781 to 0.9998), which points at the server's preprocessing again (not traced), not at the projector's precision.

## What the audio shows {#what-the-audio-shows}

From the music a small music model made for [Listen for Yourself](https://research.strata2signal.com/listen-for-yourself/), [Teaching a Music Model Chopin in Five Minutes](https://research.strata2signal.com/chopin-in-five-minutes/) and [Half an Hour with Dead Composers](https://research.strata2signal.com/half-an-hour-with-dead-composers/), all released pages, the test took 133 clips, about 85 minutes (48 kHz stereo, 5,100 seconds): among them 63 raw-and-mastered pairs (a mastered copy here is the same render with one constant gain applied and nothing else altered: 32 of the 63 were brought to −16 LUFS; the other 31, whose peaks already sat at a −1 dBTP ceiling, were left within half a decibel of their raw level, 8 of them unchanged), a five-step dial on one Chopin request (how strongly the model's adapter, a small set of add-on weights that steer its output, was applied: 0, 0.25, 0.5, 0.75, 1.0), 15 series at rising adapter strengths, and 12 requests each with a nothing-added render (the adapter at 0). Those pages also published a signal-level ruler (root-mean-square levels of clips and differences between clips, 12 printed values). The nomic pair has no audio model, so there is no incumbent and every audio word prints with that line beside it; the test was registered as EXPLORATORY with the published ruler as the comparator. (`a2-A.json`; the full 744 M-parameter checkpoint, each clip cut into 10-second windows at 16 kHz mono. Loading the audio tower moved no text vector.)

**First, the ruler reproduces.** Three candidate formulas were fixed in advance; the first to match all 12 published values at their printed precision would be the comparator. Formula F1 (RMS over the interleaved samples) matched 12 of 12 (F2 0, F3 4): **R-0 PASS**, and the RMS rows below are recomputed from the published WAV bytes.

| Reading | EmbeddingGemma 2 (1 − cosine) | RMS ruler |
|---|---:|---:|
| Largest distance from a raw clip to its mastered copy, over 63 pairs | **0.002262** | 0.049619 |
| Median of the same | 0.000054 | 0.005705 |
| Spearman rank correlation between the Chopin dial's five strengths and their distance from strength 0 | **1.0** | 1.0 |
| Series whose distance from step 0 rises strictly with strength, of 15 | **14** | 12 |
| Effect ÷ loudness: median, over the 15 series, of the distance from step 0 to the strongest step ÷ median raw-to-mastered distance | **416.5** | 18.4 |

Words: **STEADIER RULER** (no incumbent: the nomic pair has no audio model), every registered clause met (largest loudness-only distance ≤ 0.010; the dial perfectly ordered; at least 12 series in order; an effect-to-loudness ratio at least twice the RMS ruler's). Plainly: a gain change alone (the mastered copy) barely moves the audio embedding, while applying the adapter more strongly moves it in order in 14 of 15 series and about 400 times further than the gain change does (median against median; at least half of the 63 copies differ from their render by a gain of 0.43 dB or less). Prediction A-1 (that mastering would move the embedding by at most 0.010 of 1 − cosine) HELD.

**Then the request probe, which is at chance.** Each of 12 nothing-added renders ranks the 12 published requests that made them: **1 of 12** puts its own request first (the registered words needed 4 or more hits for AUDIO FINDS ITS WORDS, which a random assignment reaches with probability 0.019, and 3 for WEAK SIGNAL, 0.080; 2 or fewer reads AT CHANCE); request → clip, 2 of 12; the 36 other renders of the same requests, used as extra clip queries, put their own request first for 0.083 macro over requests. Words: **AT CHANCE**; prediction A-2 (that audio would find its words) REFUTED. Two caveats, as registered: each label is the request that made the clip, so this measures the music model's prompt-following as much as the embedder's hearing; and there is no incumbent. The audio tower can place clips by how strongly the adapter was applied; it cannot, on this material, connect a clip to a genre prompt like "hypnotic rolling groove".

## What the code shows {#what-the-code-shows}

The model card's largest claimed gain over its predecessor is on code (on MTEB, a public embedding benchmark: code 78.68 against 68.76). This bench built a set from code the workshop has already published: the Python files in the data kits of ten released pages (among them [The Ceiling Is Not the Corpus](https://research.strata2signal.com/the-ceiling-is-not-the-corpus/)), 61 files. Each function's docstring's first paragraph is the query (at least five words, 316 of them, from 46 files; 9 name their own function), and the function's source without its docstring is the document (746 functions; 17 withheld because their text carried an address or a path). Gold is the query's own function. **TEXT-ONLY** everywhere, gated at the house's +0.100 margin with a paired exact McNemar test against N\*. (`a2-C.json`.)

| Arm | Class | recall@1 (of 316) | recall@5 | MRR@10 | recall@1, the 307 queries not naming their function |
|---|---|---:|---:|---:|---:|
| **EmbeddingGemma 2, code prefixes** (`task: code retrieval \| query: ` / `title: none \| text: `) | Gated | **0.547468** (173) | 0.791139 | 0.650948 | 0.550489 |
| **nomic, the product's recipe** (`search_query: ` / `search_document: `, unit length) | Gated, N\* | **0.389241** (123) | 0.613924 | 0.488717 | 0.387622 |
| `nomic`, the model-card recipe (layer norm, then unit length) | Gated | 0.389241 (123) | 0.613924 | 0.488717 | 0.387622 |
| BM25 keyword search (the standard keyword-ranking formula; identifiers pre-split at underscores and case changes) | Reference | 0.275316 (87) | 0.525316 | 0.379236 | 0.273616 |
| EmbeddingGemma 2 with `title: <file name>` | EXPLORATORY | 0.544304 | 0.803797 | 0.658674 | 0.540717 |
| RRF of BM25 and EmbeddingGemma 2, k = 60 (RRF merges two ranked lists by summing one over (60 + rank)) | EXPLORATORY | 0.408228 | 0.664557 | 0.511332 | 0.407166 |
| RRF of BM25 and nomic, k = 60 | EXPLORATORY | 0.389241 | 0.642405 | 0.492057 | 0.390879 |

**Words: EG2 AHEAD ON CODE** (text only). Δ\* = +0.158228, 50 of 316 queries, against the +0.100 the gate needs; 64 queries only EmbeddingGemma 2 finds against 14 only nomic finds, McNemar exact two-sided p = 8.6 × 10⁻⁹; the registered bootstrap over the 46 query files gives 0.105 to 0.211, wholly above +0.100. The two nomic recipes tie on every metric: their vectors differ, but sit at cosine 0.99994 or more to each other, so the layer norm does not change nomic's text rankings here (a check read from the two caches after the score; it decides nothing). Prediction C-1, that keyword search would beat every dense model (an embedding model, as both models here are) as on the earlier bench (the companion page's), is REFUTED: BM25 trails every dense arm, by 0.272 against EmbeddingGemma 2 (p = 1.3 × 10⁻¹⁵) and 0.114 against nomic (p = 0.0005). C-2 (a gain of at least 0.050) HELD.

**The nomic numbers stand for the copy our stores run.** The product loads nomic in its own process from Nomic's Hugging Face release; our other stores reach a converted copy of the same release through Ollama. A registered check (N-PAR) embedded 200 of the earlier bench's passages through the product's loader and compared them with the Ollama copy's cached vectors: minimum cosine 0.999999332, PASS. One reach limit: 20 of the 746 code documents exceed nomic's 2,048-token context (4 of them gold); if the Ollama copy shortened those differently, at most 12 of nomic's first places could change (the 4 queries whose own function is long, all misses now, and 8 whose nomic top item is a long non-gold document), and Δ\* would still be at least 38 of 316 (0.120), so the word cannot change.

**The gain rides on the prefix, and that matters for any store.** The audit (an independent pass over the closed runs, described below) re-embedded the 316 queries three more ways, real vectors, post hoc, deciding nothing: with EmbeddingGemma 2's *search* prefix (`task: search result | query: `) in place of the code one, recall@1 falls to 0.496835 (Δ\* +0.107595, 55 against 21); with no prefix, 0.401899 (+0.012658, inside the gate's noise); with nomic's `search_query: ` string, 0.335443 (−0.053797, below nomic). A store that sends every query with one prefix keeps only part of this gain, and one that sends the wrong model's prefix loses it.

**What this set cannot say.** Docstrings written beside their code share its identifiers (that favours keyword search, which still lost); this workshop writes long "why" docstrings (that favours dense models); queries from one file are not independent (hence the file bootstrap); nothing about other languages or other people's code; the kits were published after the January 2025 data cutoff the registration records from the model card.

**The text quant ladder (T9; `a2-Q.json`, `a2-Q2.json`; TEXT-ONLY).** Three smaller GGUF files from the same repository (`unsloth/embeddinggemma-2-GGUF`, revision `ba388827…`) went through the earlier bench's 268-text parity check, and the two that did not break through the full 68-question bench:

| File | Bytes | Stage 1: minimum cosine to the reference (268 texts) | Words | Stage 2: hits at 8, of 68 (the reference's path: 45) | Words |
|---|---:|---:|---|---:|---|
| UD-Q4_K_XL | 175,673,856 | 0.966964 | **PARITY BREAKS** | Not run (stage 1 broke) | |
| UD-Q5_K_XL | 210,055,680 | 0.991992 | **PARITY DRIFTS** | 42 (loses 3 questions, gains none; p = 0.25) | **DEGRADES** |
| UD-Q6_K_XL | 248,812,032 | 0.997544 | **PARITY DRIFTS** | 44 (loses 1, gains none; p = 1.0) | **SERVING-EQUIVALENT** |
| Q8_0 (the earlier bench) | 309,855,520 | 0.999507 | (the earlier bench's parity gate: PASS) | 45 | |

The prediction that the 6-bit file would hold parity (Q-1) is REFUTED: it drifts, and "SERVING-EQUIVALENT" at n = 68 says nothing about small losses, as the registration states. The 8-bit file is the one with a measured parity; the 6-bit one is 61 MB smaller for one question.

## Speed and memory: measured, contended, not compared {#speed}

Every configuration ran with other jobs on the laptop (the 1-minute load average at each start was between 6.2 and 8.1), so by the registered rule every row is CONTENDED and no cross-runtime comparison is printed. The rows are the price lines in the section "What moving from nomic to EmbeddingGemma 2 would take" and, within one runtime, the image-budget ratios in trap 3. (`a2-S.json`; medians with a 95 % bootstrap interval, 8 threads. CPU seconds are the process's CPU time summed over all its threads (8 compute threads per path), so they run several times the wall-clock beside them: 5.3 to 9.5 times in these rows.)

| Configuration | Items | Median CPU seconds per item (95 %), summed over all threads | Median wall-clock per item | Per 10,000 images, or per million text tokens | Resident after load / peak |
|---|---:|---:|---:|---:|---|
| EmbeddingGemma 2, Google's code, 70 image tokens | 30 images | 3.145 (3.138 to 3.151) | 0.395 s | 8.74 CPU-hours | 2.36 GiB / 3.23 GiB: one process across all three budgets and the text row |
| **the same, 280 tokens (the reference)** | 30 images | **14.80** (14.77 to 14.84) | **1.97 s** | **41.1 CPU-hours** | |
| The same, 1,120 tokens | 10 images | 97.11 (96.59 to 99.76) | 13.29 s | 269.8 CPU-hours | |
| llama.cpp, BF16 text + BF16 projector | 30 images | 5.20 (5.12 to 5.52) | 0.783 s | 14.4 CPU-hours | 1.58 / 2.09 GiB (the server) |
| llama.cpp, BF16 text + Q8_0 projector | 30 images | 4.625 (4.49 to 4.865) | 0.875 s | 12.8 CPU-hours | 1.18 / 1.69 GiB |
| EmbeddingGemma 2 text, Google's code | 30 captions (median 19.5 tokens) | 0.1755 (0.1736 to 0.1824) | 23.7 ms | 9.12 CPU-seconds per 1,000 tokens; 2.53 CPU-hours per million | (as above) |
| `nomic` text, the product's loader | 30 captions (median 17 tokens) | 0.0963 (0.0961 to 0.0999) | 12.3 ms | 5.71 per 1,000 tokens; 1.59 per million | 0.98 / 1.56 GiB |
| EmbeddingGemma 2 audio, full checkpoint | 30 windows of 10 s | 3.560 (3.505 to 3.638) | 0.374 s | 9.89 CPU-hours per 10,000 windows | 3.51 / 4.60 GiB |

nomic's image path has no row: it was not timed in the speed block. The comparison run's own receipt (nomic on transformers 4.57.6†) read 46.6 ms median wall-clock per COCO image on a quiet machine; every row above was timed in a different run on a loaded one, and two runs at different loads do not divide into a ratio, so none is printed.

Three cautions. The text rows come from short captions, so a call's fixed cost dominates them. The llama.cpp image rows price a different output (their vectors miss the reference). And a per-unit memory peak the registration asked for was not captured for the Python rows (the driver read it after each unit had been removed); the memory column is each process's own resident and high-water figures, and the memory peaks the system recorded for each llama.cpp server were captured in their own logs (2.18 GiB and 1.82 GiB for the projector test's servers; 1.58 GiB and 1.18 GiB here).

## The audit and the verifies {#the-audit-and-the-verify}

Three independent passes read the runs after they closed, each with its own code; none changed a number.

**The audit** (02:07Z to 02:34Z on 2026-10-07) recomputed every table figure of the first round from the raw vector caches with its own ranking code: **238 figures, 237 exact and 1 differing by one unit in the sixth decimal** (a delta printed from two rounded recalls). It confirmed the registration order from the commit and push logs, that no GPU was touched, that every image, clip, code file and published text came from the workshop's released pages and COCO from its public mirror under its recorded terms (the off-topic passages are the earlier bench's unpublished sample, as the section on checking our work says), and, by reading each GGUF file's tensor types, that no served file holds a float16 (F16) tensor (bfloat16, a different 16-bit format the model card allows, was used on purpose). Then it attacked the instrument through the run's own scorer on scratch copies. **Two images' captions swapped:** the scorer refused the tampered cache first; with the receipt faked, the COCO figures moved by exactly the two planted queries (0.729 → 0.727; 0.698 → 0.696). **A model's prefix mislabelled:** no gate catches a wrong prefix; the run is clean because its code reads the prefixes from the frozen file, and the audit's own string check confirms it (every EmbeddingGemma 2 row's checksum equals the registered prefix plus the set's text: 0 mismatches over 4,793 rows; nomic's first-round caches carry no string checksum, and their prefixes were read from the product's archived source; the second round's carry one, 0 mismatches). **A duplicated image in the one-index test:** one exact copy whose id sorts before the original costs one query; every image copied costs *everything* (0.000); the run's pools hold none. It filed no blocking finding, two important ones about prose (the prefix dependence and the uncaptured memory peak, both now stated above), and eight minor notes.

**The first verify** (02:37Z to 03:00Z on 2026-10-07), by a different agent, re-derived the main results table, every prediction and the core table from the 36 pinned caches and manifests, with its own code (the audio ruler recomputed from the published WAV bytes): **337 figures compared, 335 equal, 2 the same sixth-decimal rounding.** Every test's word matches its registered rule. It found one prediction word printed off the rule (the server-parity prediction, above) and folded the audit's text findings, which the second verify then checked.

**The second verify** (04:27Z to 04:56Z on 2026-10-07), by a third agent, read that fold and the second round. The fold held, figure by figure. For the round, it recomputed the **550 figures in the five result files from the raw caches: 549 exact, 1 the same rounding**; every word and prediction word follows its registered rule; the decision preceded the addendum, which preceded the first vector; the 4.57.6 environment matched its lock, package for package; both nomic text recipes, re-implemented from the model cards, matched the run's vectors to within 3 × 10⁻⁸. Two attacks: a two-caption swap moved the COCO figures by exactly the planted queries through its own scorer and the run's alike (the run's checksum guard refusing first), and the pre-squashed-image test above.

## How to run it yourself, and the traps {#how-to-run-it}

As of 2026-10-06: the llama.cpp and Python package releases as last read at 18:08Z, and the model files' revisions as read during this bench's preparation (22:27Z to 23:19Z). Versions move; the traps mostly do not.

**Images and audio through Google's code.** In a fresh Python environment, `pip install "sentence-transformers>=6.1.0" "transformers==5.19.0" torch pillow torchvision` (a CPU-only torch build is fine; the two image packages are required even for text). Then:

```python
import torch
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("google/embeddinggemma-2",
    config_kwargs={"audio_config": None},        # text + vision; 439 M parameters
    model_kwargs={"dtype": torch.float32})        # never float16
model.max_seq_length = 8192
q   = model.encode("a lavender pot on a plain wall", prompt="task: search result | query: ")
img = model.encode([{"image": pil_image}], normalize_embeddings=True)     # no prompt on images
full = SentenceTransformer("google/embeddinggemma-2", model_kwargs={"dtype": torch.float32})  # all three towers, 744 M parameters
aud  = full.encode([{"array": mono_float32, "sampling_rate": 16000}])  # 16 kHz mono float32, no prompt
```

The image token budget is `max_soft_tokens` on the processor (70, 140, 280, 560 or 1,120; 280 is the default); an image is resized keeping its shape, so its token count depends on its proportions (our 1,600-pixel photographs took 256 to 273). Text and vision loaded take 438,760,448 parameters; in the speed runs, one item per call, that process held 2.36 GiB after loading and peaked at 3.23 GiB across all three image budgets, and the full checkpoint (744,371,488 parameters), counted as loaded, held 3.51 GiB and peaked at 4.60 GiB; the preparation's smoke test, four images per call, peaked at 5.25 GiB with text and vision loaded and 6.40 GiB with all three towers.

**Images and audio through llama.cpp.** Build from source at commit `4fbc76d` or later (the merge of pull request #30054; at 18:08Z on release day no release tag contained it). Download `mmproj-BF16.gguf` (982,074,880 bytes) or `mmproj-Q8_0.gguf` (554,821,120) beside the text file, and serve: `llama-server -m embeddinggemma-2-BF16.gguf --mmproj mmproj-BF16.gguf --embeddings --pooling mean -c 8192 -b 8192 -ub 8192 -np 1 -t 8`. `/embedding` then returns a 768-number vector for an image or an audio clip sent as `multimodal_data`. Two things the server does that will cost you an afternoon. **The media marker is random per start:** the literal `<__media__>` from the documentation gets HTTP 500; either read `media_marker` from `/props` each start, or start the server with `LLAMA_MEDIA_MARKER='<__media__>'` in its environment. **Its image vectors are not Google's:** text through the server agrees with the reference at cosine ≥ 0.9995 (8-bit) or ≥ 0.99998 (16-bit) on the earlier bench's 268 texts; images do not (minimum 0.9785 with the 16-bit text file and 0.9782 with the 8-bit one, on our 14 photographs; 13 of 14 below 0.999), the likeliest cause being the server's image preprocessing (a different patch grid; not traced). So image vectors from the server and from Google's code are close but not interchangeable, and a store must not mix them. The server's audio path logs itself "experimental"; a synthetic chirp agreed with the reference at 0.997. The server has no option to cut the vector: cut to 256 and re-normalise yourself.

**The traps, each measured or sourced above.**

1. **Prefixes are for text, and the right task prefix matters on top of having one.** Images and audio take none. On text, the code-retrieval prefix (`task: code retrieval | query: `) against the search prefix was worth 0.051 of recall@1 on our code (the audit's re-embeds, post hoc, deciding nothing); against no prefix, 0.146; against nomic's string, 0.212. Nothing adds a prefix for you: the model's configuration sets no default prompt, and the server adds none.
2. **Never float16.** The model card forbids it (activations overflow; NaN or silently degraded vectors). None of the repository's F16 files was downloaded, and the audit read every served file's tensor types to prove no float16 (F16) tensor was loaded. The 8-bit text file holds parity (0.9995), the 6-bit one drifts (0.9975), the 4-bit one breaks (0.9670).
3. **The image budget is not "more is better".** 1,120 tokens cost 6.6 times the CPU of 280 (contended; one runtime, so a budget comparison the rule allows) and did worse on screenshots (2 of 7 against 5) and on captions; 70 tokens cost a fifth and cost 0.264 of macro recall@1 on our scenes (5 scenes down, 2 up) and 0.030 on COCO.
4. **One space is not one list.** Images beside unrelated text: fine for EmbeddingGemma 2 (8 in 1,000 lost), hopeless for the nomic pair on transformers 4.57.6 (every caption lost). Images beside text about the same things: the text wins (139 in 1,000 put the image first). Filter by modality, or rank per modality and merge.
5. **What llama.cpp's server can and cannot embed as of 2026-10-06.** It can: text at parity, images and audio with the projector. It cannot, at the flags we ran and on our photographs: produce the reference's image vectors (in preparation, with the server's image-token count forced to equal the reference's, four small synthetic images reached 0.99987 to 0.99996 with the 16-bit files, while photo-sized images at the registered flags still differed; not traced further), cut the vector, or take the documented media marker without the environment variable.
6. **The Ollama version gap.** The bench laptop's Ollama (0.32.14) refuses even the text-only tag ("requires a newer version of Ollama that may be in pre-release", HTTP 412). Six of our eight text stores embed through Ollama.
7. **Remote code, read before it runs.** The incumbent's image model needs `trust_remote_code`, and its loading code lives in a *third* repository (`nomic-ai/nomic-bert-2048`) that the model's configuration names with no revision, so an unpinned load executes whatever that repository's `main` holds that day. We pinned the code by its cache reference (`7710840…`) and read both files whole before the first execution: `configuration_hf_nomic_bert.py` (55 lines; it stores hyperparameters and does no I/O) and `modeling_hf_nomic_bert.py` (2,556 lines). The read found that the text loader ignores `revision` (its keyword arguments swallow it), tests a working-directory-relative path first and would load a pickled `pytorch_model.bin` from there with `torch.load` (line 440 at that revision; torch 2.6 and later restrict such a load to plain weights by default), and otherwise resolves weights by the cache's `main` rather than the pinned revision (lines 79 to 101). Every execution ran in a unit with the network removed, from an empty working directory. For systemd user units: `IPAddressDeny=any` alone did *not* take the network away in a user-level manager (a probe fetched a web page, HTTP 200), nor did `PrivateNetwork=yes`; adding `RestrictAddressFamilies=AF_UNIX` did.
8. **The incumbent's image path, as shipped, may not load, and squashes what it loads.** nomic-embed-vision-v1.5 refused to load under every transformers 5.x release we tried, unpatched (its configuration carries `n_inner` as `2048.0` where 5.x's strict validation wants an integer: `StrictDataclassFieldValidationError`); 4.57.6 loads it, with its text side bit-identical to 5.12.1's. Its processor resizes every image straight to 224 × 224, so a portrait photograph is squeezed, not cropped. And the model cards' query recipe (`search_query: ` with a layer norm) scored above the product's document-prefix recipe on both image sets (0.497 against 0.464 on COCO; 0.272 against 0.208 on our 12 scenes, exploratory, sign test p = 1.0). Pin the library version that loads it, and read the model card.
9. **Both models are 768 wide.** A store that mixed them would pass every dimension check and simply return bad neighbours; so would a store that mixed one model's two runtimes for images.

## What moving from nomic to EmbeddingGemma 2 would take, and why we are not doing it {#what-moving-would-take}

We were asked to say what it would take, and to say it without doing it. From a survey of every vector store this workshop knew of (2026-10-06, 18:08Z to 18:25Z) and the figures above:

**What would move.** As of 2026-10-06, Nomic's text release writes the vectors eight stores hold (the conversation archive, our searchable record of past working sessions, with its messages and its session summaries counted apart; six stores through an Ollama tag, two from Nomic's own checkpoint loaded in-process: RuleSage's retrieval index and a memory store beside it). RuleSage's retrieval code is also wired to Nomic's vision model for pictures; as of 2026-10-07 no product sends it a picture, and under the product's own pins it does not load (above), so how many picture rows its store holds, if any, was not read. One model would replace both. Every store would be re-embedded, and three ties set the order: two stores share one vector space across two databases and must switch together; this site's related-page vectors (843 rows, generated 2026-10-05) are one file rewritten whole; [RealKeep](https://realkeep.strata2signal.com)'s worlds (a game of ours) never re-embed on replay, so a switch there means a new world. Four of the eight stores carry a model name on every row.

**What it would cost, per item.** The one large store we counted is the conversation archive: 701,937 message vectors and 489 session-summary vectors at the survey (counted 2026-10-06 18:18Z). The other row counts (the glossary, RuleSage's text and image rows, the memory store, the claims store (the ledger of checked factual claims the companion page's questions were drawn from), the worlds) were not read. The price per unit, from the speed table, contended, on one laptop CPU: 9.12 CPU-seconds per 1,000 text tokens for EmbeddingGemma 2, and 5.71 for nomic through RuleSage's own loader (both on short captions and both contended, so they are price lines, not a comparison; the Ollama copy six of our stores use was not timed); 41.1 CPU-hours per 10,000 images at the default budget; 9.9 CPU-hours per 10,000 ten-second audio windows. The archive's token count was not read, so no total or duration is printed; a 700,000-row re-embed is its own project. Every store's prefixes change, and after the code result, a store would also have to know *which* task prefix its queries carry.

**What would have to serve it.** Six of eight text stores embed through Ollama, and the release on our bench laptop refused the model on 2026-10-06; a switch there needs an Ollama release that loads the tag and passes the 268-text parity gate (not tried in this bench), or a separate llama.cpp server for text with Google's code for images. The reference path is Python with transformers 5.19.0, which RuleSage's pinned 5.12.1 is not.

**What comes first regardless.** No store records which installed copy of a model wrote a row (some stamp the model's name, none the bytes that answered), every dimension check in the workshop is a check for 768, which EmbeddingGemma 2 passes, and only one reader filters by that name. A provenance guard, needed with or without a switch, ties every vector to the digest of the files that made it, has each writer prove which files answered before it writes, and has the databases refuse a write without that proof. Once it is in, a switch follows a fixed road: a new column beside the old, re-embedded under the guard, readers flipped in one transaction, the old column retired after a waiting period.

**The measured reasons not to.** On our own text, NO MATERIAL DIFFERENCE (45 against 44 of 68, short of the pass mark of 51). On images, EmbeddingGemma 2 is ahead of the nomic pair: by 232 captions in 1,000 on a public set, a gated word that clears the house bar, and by 0.519 of macro recall@1 on our 14 photographs, exploratory at 12 scenes. But the public set most likely sits inside both models' training distributions; the pair ran on a library our product does not pin (transformers 4.57.6); the product's picture path is wired but cannot load its image model under the product's own pins (how many picture rows its store holds was not read), so no working image search of ours could take that win; and the registration says a lead on these photographs licenses only drafting a larger confirmatory bench. On code, a gated gain (+0.158 recall@1 on 316 docstring queries, p = 8.6 × 10⁻⁹) that depends on the code-retrieval prefix, and none of our eight stores is a code index. On audio, a capability our stores do not use. Against that stands the full re-embed, an Ollama gap, a re-pinned RuleSage, and a guard that is not built yet. The pass mark for displacing an incumbent was set by what a switch costs; the public image set and our code clear it, but neither is a store we run, and nothing measured on a store we run has cleared it.

**What would change our mind,** each a registered measurement, not an opinion: (1) image search heading into a product of ours, which opens the confirmatory photograph bench the registration names: drafted and registered once our released photographs reach 68 scenes (the count the registration set), it would need to clear a +0.100 gate with a scene-level sign test at p ≤ 0.05 against the nomic pair; (2) a store of ours whose queries look like code retrieval, benched with the same +0.100 gate on its own text. Either produces a measured proposal naming every store, every row count and the priced re-embed, with the guard landed first; a configuration edit is never an adoption. For the six stores that embed through Ollama, a switch would also need an Ollama release that pulls the tag and passes the parity gate.

## What this bench does not say {#what-this-bench-does-not-say}

- **Its image comparison is with the nomic pair on transformers 4.57.6**, a library release our product does not pin, because the pinned one cannot load the pair's image model; the core's registered verdict stays NOT EVALUATED: NO INCUMBENT, with an exploratory reading beside it.
- **Fourteen photographs of one subject**, with English texts written by this workshop; the words describe these photographs, never images in general. Whether either model saw them in training cannot be proven; they were first published from 2026-09-25 to 2026-10-05, after the January 2025 data cutoff the registration records from EmbeddingGemma 2's model card (not yet checked first-hand).
- **COCO is most likely inside both models' training distributions** (its images have been public since 2014), and, as the registration notes (not yet checked first-hand), EmbeddingGemma 2's model card reports a benchmark (MMEB v2) whose training split includes COCO retrieval; a lead there is weaker evidence than a lead on our own pictures.
- **What the 224 × 224 squash costs nomic was not measured**; a padded variant is registered and did not run.
- **The audio labels are the requests that made the clips**, so the request probe measures the music model as much as the embedder, and there is no incumbent.
- **The code set is our own**; nothing about other languages or other people's code.
- **The speed rows are contended** and compare nothing across runtimes.
- **The parity gates show runtimes agree with the reference; they cannot catch an error both share.** The code was under a day old.
- **The 8-bit projector's word is the registered word, read literally**, and the numbers beside it say what it measured.

## What to take with you {#what-to-take-with-you}

- **On images, EmbeddingGemma 2 is ahead of the nomic pair.** On 1,000 public COCO captions, 0.729 against 0.497 recall@1 (Δ\* +0.232, 276 captions against 44, p = 3.2 × 10⁻⁴²): EG2 AHEAD, CLEARS THE HOUSE BAR, on a set most likely inside both models' training data, so weaker evidence than a lead on our own photographs. On our 14 released photographs, 33 of 40 published descriptions put the right photograph first against the pair's 13 (79 % against 27 % with each of the 12 scenes weighted equally, the registered measure; chance 8 %), exploratory-only by registration (12 scenes, these photographs only; the frozen verdict NOT EVALUATED: NO INCUMBENT). Every nomic image figure is nomic on transformers 4.57.6, outside our product's pinned library, which cannot load the pair's image model.
- **One space is not one ranked list.** EmbeddingGemma 2's images hold their place beside off-topic text (8 in 1,000 lost) and lose it beside on-topic text (139 in 1,000 keep it); nomic's, on transformers 4.57.6, never surface beside any text (TEXT CROWDS OUT IMAGES; EG2 AHEAD IN ONE INDEX). A unified store still filters by modality.
- **On code, a registered win (TEXT-ONLY):** 0.547 against nomic's 0.389 recall@1 on 316 of our own docstring queries, Δ\* +0.158 (gate +0.100), p = 8.6 × 10⁻⁹, EG2 AHEAD ON CODE; the margin over nomic falls to +0.108 with the search prefix and +0.013 with none (the audit's re-embeds, post hoc).
- **Audio embeds and tracks the adapter's strength steadily** (a gain change alone moves it by at most 0.0023 of cosine; the adapter dial is perfectly ordered; no incumbent: the nomic pair has no audio model), but a clip cannot find the request that made it (1 of 12, chance; each label is the request that made the clip).
- **Nothing in our stores changes.** The text result found no material difference, the code win has no code index of ours to land in and the image win no working image search, eight stores would be recomputed, and the provenance guard comes first.

## How to check our work — and see it live {#how-to-check-our-work-and-see-it-live}

A data kit for this page ships separately, redacted for publication (internal names, paths and ports replaced, each replacement listed in the copy); a dated note here will say when it is up. **Almost everything this bench measured is already public**: the 14 photographs, the 23 other images and all their texts are on the pages linked above (as served at the freeze, 2026-10-06 22:42Z; the kit lists each file's checksum); the code set is the released kits' own files; the audio is the released clips. The kit carries: the registration and its six addenda as redacted copies; the set manifests (every image's public path and checksum, every text verbatim, every COCO id with its caption and attribution, every function's file and line, every clip's path and checksum, every swapped variant); every scorer output file named in the tables above, each with its gates; the per-query rank files; every unit's receipt with its GPU readings and contention list reduced to a count; and the checksum of every vector cache (the caches themselves are not shipped). The digests of any file the redaction touched (the registration, its addenda, and every scorer output, receipt or script the redaction changes) are withheld from every shipped file, because a digest of a file that was later redacted lets anyone confirm a guess at what was removed. Recompute any recall@1 from the rank files: count the queries whose first gold rank is 1, divide by the number of queries, and for the core average the 12 scene shares. The 2,569 text passages used as off-topic distractors are the earlier bench's unpublished sample; the kit carries their checksum.

To reproduce the model side: the weights are on Hugging Face at [google/embeddinggemma-2](https://huggingface.co/google/embeddinggemma-2) (revision `914f7f89…`) and [unsloth/embeddinggemma-2-GGUF](https://huggingface.co/unsloth/embeddinggemma-2-GGUF) (revision `ba388827…`; projectors `995aaa56…` and `90e7b023…`; the three small text files `ea905fd0…`, `a9d7a3b7…`, `dc84f042…`); the incumbent at [nomic-ai/nomic-embed-text-v1.5](https://huggingface.co/nomic-ai/nomic-embed-text-v1.5) (`e9b6763…`), [nomic-ai/nomic-embed-vision-v1.5](https://huggingface.co/nomic-ai/nomic-embed-vision-v1.5) (`e3a725b…`, Apache 2.0 at that revision) and its code at [nomic-ai/nomic-bert-2048](https://huggingface.co/nomic-ai/nomic-bert-2048) (`7710840…`); COCO's test shard from the mirror named above.

And the live door: [ask RuleSage a rules question](https://rulesage-live.strata2signal.com/), free and without an account, and open the wrench on the ruling (the panel behind every answer, which shows the search's own timing and the pages it cited). Its meaning-search runs the nomic text model measured on the code set here. Its retrieval code is wired for the nomic vision model too, and this bench could not load that model under the product's own library pins.

## The rest of the seminar {#the-rest-of-the-seminar}

- [EmbeddingGemma 2, measured on release day against the embedder we already run](https://research.strata2signal.com/embeddinggemma-2-on-release-day/) — the text half: 45 against 44 of 68, short of the pass mark of 51, NO MATERIAL DIFFERENCE, and the runtime traps this page builds on.
- [GPUs and orchids](https://research.strata2signal.com/gpus-and-orchids/) and [The same eighteen pictures, card by card](https://research.strata2signal.com/the-same-eighteen-pictures-card-by-card/) — the photographs this bench searched, with the words it searched them by.
- [Nine Worlds, One Dog](https://research.strata2signal.com/nine-worlds-one-dog/) — the nine paintings whose world names neither model reliably read back (4 of 9 one way, 7 of 9 the other, for EmbeddingGemma 2; nomic, on transformers 4.57.6, less often).
- [Listen for Yourself](https://research.strata2signal.com/listen-for-yourself/) — the clips and the published RMS ruler the audio tower was measured against.
- [A short history of Ollama](https://research.strata2signal.com/a-short-history-of-ollama/) — the server most of our stores embed through, and the story behind trap 6.

[The whole shelf](https://research.strata2signal.com/) holds the rest: the benches behind the claims we publish, failures included. If there is a measurement you want next, [say so](https://strata2signal.com/contact/); the suggestion box is read.

## Who ran this, and thanks {#who-ran-this-and-thanks}

The models: **EmbeddingGemma 2** (`google/embeddinggemma-2`, Google, Apache 2.0, revision `914f7f89…`), its GGUF conversions and projectors from **Unsloth** (`unsloth/embeddinggemma-2-GGUF`, revision `ba388827…`); **nomic-embed-text-v1.5**, **nomic-embed-vision-v1.5** and the `nomic-bert-2048` loading code from **Nomic AI** (Apache 2.0), at the revisions above. The data: this workshop's own released photographs, paintings, screenshots, kit code and music; **COCO** (the COCO Consortium; annotations CC BY 4.0; images under the Flickr terms, used for evaluation only). The runtimes: **llama.cpp** (ggml-org) at commit `4fbc76d`, **sentence-transformers** 6.1.0 over **transformers** 5.19.0 and **PyTorch** 2.14.1 (CPU build) for the challenger; **transformers** 5.12.1 (the product's pin; its text model) and 4.57.6 (the incumbent's image arms, outside the product's pin) over **PyTorch** 2.12.1 (CPU build) for the incumbent. The licences this shelf has read first-hand are on its [licences](https://research.strata2signal.com/licences/) page.

A small human team asked for this bench on the day the model's text half was measured, chose what it would and would not claim, and signed the numbers. A fleet of AI agents wrote the registration and its addenda, chose the nine further tests, built the sets from the published pages, ran the arms in isolated units, made the two decisions the second round waited on (under that team's standing instruction, each recorded with its time before the arms it admitted), audited and verified the runs with independent code, and drafted this page under that team's rulings. No visitor's data was involved.

## Changes {#changes}

- 2026-10-07 (UTC): draft v1, from the first round's verified results, with the two nomic image arms waiting on a decision; draft v2, a critic's round folded and the bench's second round of runs (the nomic image arms and the swap test, 2026-10-07 04:11Z to 04:26Z, verified fresh) folded into the text; draft v3, a fresh check's fixes; draft v4, a second fresh critic's round folded (1 blocking, 9 important, the small fixes as one batch); nothing released; the data kit ships separately.
- 2026-10-08 (UTC): draft v5, five terms (our stores, docstring, the audio ruler, the provenance guard, EG2) explained where the page's opening first uses them (exact edits; no figure or verdict word changed); draft v6, two fresh critics folded (a cold read and a laws-and-data read): the library story told as what failed, what ran instead and when that was decided; the core's frozen verdict and its exploratory reading separated in plain words; "model card" written in full everywhere; the speed table given a wall-clock column; a mastered copy defined as a gain change; the public-set training caveat hedged; no figure or verdict word changed; draft v7, a fresh check's fixes (the opening's "before the first score it governs" restored; when the backup nomic arms were registered, stated against the first core score; the core's two registered words printed exactly; CPU seconds defined as summed over all threads; the image model's load failure named as our product's in two more places; one scene counted on two pages), no figure changed; nothing released.
- 2026-10-09 (UTC): published; the byline dated.

Corrections and later measurements will be added below, each dated (UTC) with a window at both ends where one applies, each saying in plain words what it counts, and each comparing itself in one sentence to the reading it replaces.

<!-- derived 2026-10-09 (UTC) by tools/derive_md.py from the pour source.
     source html sha256: 5e46499076b1bfcaf6517104288c66ecc64cfe7ba1aadd625607d21d25355c69
     derivation sha256:  2a2ffe985c24bec2c68bada01d7de3154497f5bcf63a9b6c937167671aceac28
     the {#id} on each heading is the anchor that heading carries on the page. -->
