# EmbeddingGemma 2, measured on release day against the embedder we already run

*Google released a new open embedding model on 2026-10-06. By 20:26Z the same day (UTC) we had put its text model through our own 68-question retrieval bench against nomic-embed-text, the model our search stores already run, behind a pass mark of 51 of 68, written down and committed before any EmbeddingGemma score existed. The registered verdict is NO MATERIAL DIFFERENCE: 45 against 44 of 68 questions. Plain keyword search still beats both on this set, where a right answer is a passage that holds a quoted source passage word for word (keyword search's home ground), and every row of the results board is computed from one per-question file, which the data kit will carry.*

*Published 2026-10-09 (UTC) · A small (human) team and a fleet of AI agents.*

**the short version:** On 2026-10-06 Google released EmbeddingGemma 2, an open embedding model under the Apache 2.0 licence: it turns text (and images, video and audio) into lists of numbers that search can compare. The same day we wrote down a test, froze it in a commit, and ran the model's text part against nomic-embed-text, the embedder our search stores already run, on 68 questions over 2,569 passages of our own documents. nomic-embed-text finds a correct passage in the top 8 for 44 of the 68; EmbeddingGemma 2 finds one for 45, through Google's own code and through an 8-bit copy of its weights on llama.cpp (an open-source model runner) alike. The pass mark, written down before any score existed, was 51 of 68, so the registered verdict is NO MATERIAL DIFFERENCE: one question of difference is inside the noise, which is not the same as saying the two are equal. Plain keyword search (BM25) finds 58 of 68, and nothing we layered on it beat it by more than chance. Nothing in our stores changes.

9,788 words · about 44 minutes (at 220 words/min) · 9 tables · data kit: no

https://research.strata2signal.com/embeddinggemma-2-on-release-day/

---

*If you're new here: [strata→signal](https://strata2signal.com) is a small workshop that builds on its own machines and writes up what it measures. Every figure below was read on 2026-10-06 (UTC) from the result file, registration or log named beside it, or on the later date printed beside it. The data kit that holds those files ships separately; a dated note here will say when it is up. Machines are named by their role, never their hostname; a number we could not source is a dash with its reason.*

*Eight words this page leans on.* An **embedding model** turns a piece of text into a list of numbers (a **vector**; both models here write 768 of them per text), so that texts about the same thing end up with similar lists. **recall@8** is the share of questions with at least one correct passage in the top 8 results, the number our gate reads. **BM25** is the classic keyword-matching formula behind most full-text search, needing no model. A **pre-registration** is the test, the thresholds and the words for each outcome, committed before the first score, so nobody picks the metric after seeing the answer. A **gate** is one of its pass marks: a result clears it or does not, and the page prints which. A **prefix** is a short fixed string ("task: search result | query: ") these models expect in front of every input; forgetting it costs real accuracy, and this page measures how much. An **arm** is one configuration under test (a model, a runtime and a prefix choice), run over the same 2,569 passages and 68 questions; every row of the results board is one arm.

## What an embedding model is {#what-an-embedding-model-is}

Search with one is arithmetic: turn the question into its list, turn every passage into its list once and store them, and the passages whose lists point closest to the question's are the candidates.

That is the mechanism behind "semantic search" and part of the retrieval half of most retrieval-augmented assistants, [RuleSage](https://rulesage.strata2signal.com/), this workshop's board-game rules helper, included. The stored lists are the expensive part: computed once per passage with one model, they compare only against lists from that same model, so swapping the model means recomputing every one. That cost is why this page exists, and why its gate is set where it is.

The model that made our stored lists is **nomic-embed-text**, an open model from Nomic AI. This bench measured the copy we run through Ollama (tag `nomic-embed-text:latest`, manifest digest `0a109f422b47`; 274,302,450 bytes on disk, a 137 M-parameter model in 16-bit weights, as read from the Ollama store, with a 2,048-token context per the earlier bench's registration). The tag is the v1.5 release (Ollama's library lists `latest`, `v1.5` and `137m-v1.5-fp16` under that one ID, read on 2026-10-06 at 22:53Z). RuleSage's retrieval runs Nomic's own copy of the same release, `nomic-embed-text-v1.5` from Hugging Face, inside its own process: the same model in a different file and runtime, a copy this bench did not measure. Both write 768 numbers per text and expect `search_document: ` in front of a passage and `search_query: ` in front of a question.

## The question {#the-question}

Does EmbeddingGemma 2's text model find our own text enough better than nomic-embed-text to be worth recomputing every stored vector in the workshop?

That is narrower than "which model is better": it is the question an operator with a running system has, and its answer has a price, so the price sets the bar: the registered gate asks for an improvement of at least 10 points of recall@8 over the model we run (10 percentage points, which is 7 more of the 68 questions found), with a paired significance test at p ≤ 0.05. Anything less is recorded, in the registered words, as NO MATERIAL DIFFERENCE, and the incumbent keeps its place by measurement rather than by default.

## What Google released on 2026-10-06, and what we checked {#what-google-released}

**What it is, from Google's own pages.** [EmbeddingGemma 2](https://huggingface.co/google/embeddinggemma-2) is Google's second open embedding model. One checkpoint turns text, code, images, video and audio into the same 768-number vector. The text part is a 270 M-parameter model (a 130 M transformer plus a 140 M embedding table) and loads on its own; vision (170 M) and audio (300 M) are "selectively loadable". The vector can be cut to 512, 256 or 128 numbers, which the card calls "close to lossless down to 256". The context is 8,192 tokens. The card is explicit: "Run inference in `bfloat16` or `float32`. Do not use `float16`", because the activations exceed float16's range and give NaN or silently degraded vectors. The repository last changed on 2026-10-06 at 15:29Z; Google's [announcement](https://blog.google/innovation-and-ai/technology/developers-tools/embeddinggemma-2) is dated 2026-10-06. (Sources: the model card and `config.json`; the Hugging Face model API; the blog; [Google's EmbeddingGemma documentation](https://ai.google.dev/gemma/docs/embeddinggemma).)

**The announcement against the sources.** Unsloth, which published a [guide](https://unsloth.ai/docs/models/embeddinggemma-2) and [GGUF files](https://huggingface.co/unsloth/embeddinggemma-2-GGUF) for the model the same day (GGUF is the single-file model format llama.cpp loads), made four claims we checked:

| Claim (Unsloth's guide) | What Google's own pages say | Our reading |
|---|---|---|
| 740 M parameters = 270 M text + 170 M vision + 300 M audio | The same split; the API counts 744,371,512 | Right |
| "Runs on 0.5 GB RAM" | The blog: about 191 MB for the text-only weights and about 567 MB for the full multimodal model, quantized, on a Pixel 11 Pro | Rounded down; 0.5 GB is the full model (567 MB), text alone is about 0.19 GB |
| Text, code, images, video and audio | The same | Right |
| "make the text model the default embedding model" | Not on any Google page; the guide shows it as a choice in Unsloth's own desktop app's retrieval settings | Unsloth's app, not a Google claim |

*The two sources graded above, each linked where it is first named: Google's blog announcement of EmbeddingGemma 2, dated 2026-10-06, and Unsloth's EmbeddingGemma 2 page in its model documentation; both were read on 2026-10-06.*

**Quality, from the cards (full precision, 768 numbers).** Scores are averages on MTEB, the standard public leaderboard for embedding models; higher is better.

| Benchmark | EmbeddingGemma 2 | EmbeddingGemma 1 (308 M) |
|---|---:|---:|
| MTEB multilingual v2 | 61.36 | 61.15 |
| MTEB English v2 | 68.46 | 69.67 |
| MTEB code v1 | 78.68 | 68.76 |

The code score is up 9.92 points, which is the "about 14 %" relative gain the card describes (78.68 / 68.76 = 1.144). The card's main comparison table does not list English v2; its truncation table does, at 68.46 in the 768-number row, 1.21 below the first EmbeddingGemma's 69.67. At 256 numbers the card gives 60.41 multilingual, 67.78 English and 76.18 code. The card says "All results reported below use the full-precision checkpoint": no quantized build has a published figure. **No source compares the model with nomic-embed-text**, which is the comparison our stores need. (Sources: the EmbeddingGemma 2 card's comparison and truncation tables; the [EmbeddingGemma 1 card](https://ai.google.dev/gemma/docs/embeddinggemma/model_card)'s evaluation section.)

**Where the sources disagree**, read on 2026-10-06 between 16:30Z and 16:45Z:

| Item | One source | Another | What we did |
|---|---|---|---|
| Context length | The card: 8,192 tokens | `config.json`: `max_position_embeddings` 262144; the GGUF metadata: 262144; the Ollama library shows "256K" on every tag | Treated 8,192 as the limit and set it explicitly on every runtime |
| Sliding attention window (in some layers each token looks only at a fixed span of nearby tokens; this is that span) | The card: 1,024 tokens | `config.json`: 512; the GGUF metadata: 1,024 | Checked after the run by the audit (our own agent's pass, below), an extra, unregistered check: on the 67 of our 2,569 passages longer than 512 tokens the 16-bit (BF16) GGUF file agrees with the float32 reference path (defined below) at cosine 0.99999 or above (cosine: a similarity score, 1 for identical vectors) and the 8-bit (Q8_0) file at 0.99958 or above, so the disagreement had no effect on our passages |
| float16 | The card: do not use it | Unsloth ships an `-F16.gguf`; the ONNX repository ships `fp16` and `q4f16` builds (reported by a research pass, not re-checked) | Never ran any float16 file or dtype |

**The licence.** The repository is tagged Apache 2.0; the blog calls it "a commercially permissive Apache 2.0 license". That is a change: the first EmbeddingGemma is under the Gemma licence and gated behind manual approval. Apache 2.0 permits commercial use, hosting and redistribution, provided the licence text and attribution notices ship with the weights. One open question: the card also says "Deployments must adhere to the Gemma Prohibited Use Policy", a sentence Apache 2.0 does not contain; whether it binds anyone is a legal question this page does not answer. A research pass reported no LICENSE or NOTICE file in the repository; we did not re-check.

**Could it be run on release day?** At 16:41Z on 2026-10-06, only on code hours old:

- **llama.cpp**: [pull request #30054](https://github.com/ggml-org/llama.cpp/pull/30054), which adds the architecture (`gemma-embedding2`, mean pooling), merged at 16:20:31Z as commit `4fbc76d`. No release tag contained it (the newest at 16:03:55Z was b11445; at 18:08Z, b11446 and b11447 still lacked the merge, both published after it yet cut from commits that do not contain it): running it meant building from source.
- **transformers**: support is in version 5.19.0. At 16:39Z that version was not on PyPI (5.18.0 was the latest); by 18:08Z it was, and that is the wheel we installed. sentence-transformers 6.1.0 was already published.
- **vLLM**: [pull request #60254](https://github.com/vllm-project/vllm/pull/60254) was open, not in any release. Google's blog lists vLLM as supported anyway.
- **Ollama**: the [library](https://ollama.com/library/embeddinggemma-2/tags) lists `embeddinggemma-2:270m-bf16-text` (574 MB, a 16-bit text-only tag by its name) and two neighbours, the bare `:270m` (378 MB) and `latest`/`740m` (1.3 GB, the full multimodal model); on 2026-10-06 we did not read how any of the three is built, and no page we read stated the minimum Ollama version. The copy on our bench laptop, 0.32.14, was several releases behind (per Ollama's GitHub releases, 0.35.0 was out by 2026-09-28, and a 0.40.0 pre-release existed); it refused the pull (below). Read again on 2026-10-07, two things had moved or come into view. At 12:35Z (UTC) Ollama's releases showed 0.40.0 as a full release whose tag commit is dated 2026-10-06, where on 2026-10-06 we had seen only its pre-release. Between 14:41Z and 15:01Z (UTC) the registry's own configuration of each tag showed every `embeddinggemma-2` tag, the 16-bit text one included, to be a safetensors model for Ollama's MLX runner (MLX is Apple's machine-learning framework) that declares a minimum Ollama of 0.36.0, where on 2026-10-06 we had read only the library's tag list. This bench tried no release newer than 0.32.14; whether any loads the model is not measured on this page.

## The instrument: 68 questions over our own documents {#the-instrument}

The test set was frozen on 2026-08-19 for an earlier bench of ours (the first embedder bench, E-1) and was reused byte for byte. Nothing about it was chosen after seeing an EmbeddingGemma number.

**The documents.** 164 web pages the workshop had fetched for a different project (a tool that assembles source dossiers about a town), all about one historical gold-rush town, in English, heavy with proper nouns, dates and figures. The pages were cut into **2,569 passages ("windows") of at most 1,200 characters**, paragraphs packed greedily, never overlapping: 2,431,453 characters in all. Measured with EmbeddingGemma 2's tokenizer, the longest input is 831 tokens (67 windows exceed 512). A 1,200-character window cannot reach either model's limit (2,048 tokens for nomic-embed-text, 8,192 for EmbeddingGemma 2), so no window is truncated.

**The questions.** 68 factual claims about that town, drawn from a set of 113 whose supporting passages our own AI research agents had already checked against independent sources in July 2026 (an adjudicated set). They are statements, not questions in the grammatical sense: each was used verbatim as the search text, and this page calls them questions because that is their role in the search. They run 31 to 259 characters, median 104. Each claim came with the exact passages of the source pages that its reviewing agent had read to support it, and a window counts as a correct answer ("gold") for a claim when it contains one of those passages (a quoted passage with an ellipsis is split at the ellipsis, every part must appear, and parts under 12 characters are dropped). Of the 113, claims whose quoted text is nowhere in the 164 pages were excluded beforehand (32), as were 13 whose gold rule was too loose to measure anything; nothing was counted as a miss that the corpus could not answer. Gold windows per question: minimum 1, median 2, 90th percentile 7, maximum 18.

**The scoring.** For each embedding model, every window's vector is computed once and every question's vector is compared with all 2,569 by exact cosine similarity (no approximate index), ties broken by ascending window id. **recall@8** is the share of the 68 questions with at least one gold window in the top 8. Eight windows picked at random would score 0.0091, the floor every number below is read against. Reported beside it and gating nothing: recall at 1, 3 and 20, nDCG@10 (how high the correct passages sit within the top 10) and MRR@20 (one over the rank of the first correct passage, within the top 20). The significance test is McNemar's exact two-sided test on the paired per-question hits: with *b* the questions only the challenger finds and *c* those only the incumbent finds, the p-value is the exact binomial over those discordant questions.

**One stated weakness, in both directions.** The gold label is literally "this window contains these words". That is keyword search's home ground, which is why BM25 is on the board as a reference and not as a gate. A dense model (an embedding model, which both models under test are) that beats BM25 here has beaten it on hostile ground; a dense model that loses to it has not necessarily lost on friendly ground. Both readings were written down on 2026-08-19, before any dense model was scored.

**What the earlier bench had already measured**, on 2026-08-19 (its registration is a 602-line document committed at 2026-08-19 18:15:57Z; its report is in the data kit):

| Arm | What | recall@8 (of 68) |
|---|---|---:|
| L0-bm25 | BM25, k1 = 1.2, b = 0.75, over the same 2,569 windows | 0.853 (58) |
| A1-asym | `nomic-embed-text` with its documented prefixes, `search_document: ` / `search_query: ` | 0.647 (44) |
| A2-symdoc | `nomic-embed-text`, `search_document: ` on both sides | 0.588 (40) |
| A3-bare | `nomic-embed-text`, no prefixes | 0.559 (38) |

That bench registered four challengers and ran none, so until 2026-10-06 nomic-embed-text held its place in our stores by default, under that bench's rule that the incumbent stays unless a challenger clears the gate, not by any win. The first EmbeddingGemma (`embeddinggemma:300m`), one of the four, ran this time as an exploratory row.

**The pre-registration for this bench** (E-2) is a 704-line document that fixes the arms, the exact prefix strings, every gate, the power clause, the run protocol and seven predictions. It was committed at **2026-10-06 18:24:49Z** to the workshop's private repository and pushed six seconds later. The first EmbeddingGemma arm started at **18:41:48Z** (its cache header); the scorer, which computes every retrieval metric and significance test by calling the 2026-08-19 bench's scoring code unchanged, was pinned by a further commit at **18:47:24Z**; the first EmbeddingGemma recall number was computed at **18:50:41Z**. The body never changed after the freeze: the file at the end of the run is the frozen bytes plus dated addenda, each recording only a checksum, a receipt, an arm that could not run, one install-source change or one corrected timestamp (transformers came from the PyPI wheel, not the git tag the registration named, same version). An audit by a separate AI agent of ours, with its own code and not an outside reviewer, re-read these times from the commit and push logs and the result files' own headers. Both registrations name our machines and folders, so the data kit carries redacted copies, each listing what it replaced, and the digests of the unredacted files are withheld, because a digest of a file that was later redacted lets anyone confirm a guess at what was removed.

**The gate, in the registered words.** Two arms are gated: the **reference path** (`G2-768-ref`, the text model through Google's own sentence-transformers code in 32-bit floats, what the model is) and the **serving path** (`G2-768-q8`, the 8-bit quantized GGUF file on a llama.cpp server, what a server of ours would run; Q8_0 is 8-bit integer weights with one scale per block). Each is **ADMITTED TO PROPOSAL** (a pass would have produced a costed plan to switch, not a switch) only if it passes the integrity (G-E2), no-GPU (G-CPU) and, for the serving path, parity (G-0) gates in the table below, **and** `recall@8(arm) − recall@8(nomic) ≥ +0.100`, **and** McNemar exact two-sided p ≤ 0.05. "Anything less is NO MATERIAL DIFFERENCE". Against nomic's 44 of 68 that floor is **51 of 68**, gained almost one-sidedly (at 51 hits the challenger may lose at most one question nomic finds). Below-margin arms get their full numbers on the board and no "close second" narrative. A null is read under the registered power clause, the registration's own statement of what this test can see: with 68 questions it detects only large effects, and a smaller real gain could hide in the noise, so a null reads **"no effect of at least 10 points was detected at n = 68"**, never "the two models are equal".

## What we ran {#what-we-ran}

**The arms.** Every arm embeds the same 2,569 windows and 68 questions, except the derived arms, which reuse another arm's vectors.

| Arm | Class | What | Prefixes | Numbers per vector |
|---|---|---|---|---:|
| L0-bm25 | Reference | BM25 over the windows (2026-08-19's run, re-scored) | — | — |
| A1-asym | Incumbent | `nomic-embed-text`, the vectors cached on 2026-08-19 | `nomic`'s | 768 |
| A1-cpu | Check | The same nomic model, re-embedded fresh on the CPU | `nomic`'s | 768 |
| **G2-768-ref** | **gated** | EmbeddingGemma 2 text model, sentence-transformers, float32, CPU | Google's | 768 |
| **G2-768-q8** | **gated** | The Q8_0 GGUF on a llama.cpp server built from source, CPU | Google's | 768 |
| G2-768-bf16 | Exploratory | The BF16 GGUF (16-bit floats with float32's exponent range) on the same server | Google's | 768 |
| G2-768-ollama | Exploratory | The Ollama tag `embeddinggemma-2:270m-bf16-text` | Google's | 768 |
| G2-768-noprefix | Exploratory | The reference path with no prefix on either side | None | 768 |
| G1-300m | Exploratory | EmbeddingGemma 1 through Ollama (`embeddinggemma:300m`) | Google's | 768 |
| G2-512-ref, G2-256-ref, G2-128-ref | Exploratory, derived | The reference vectors cut to 512, 256 and 128 numbers and re-normalised | — | 512 / 256 / 128 |
| G2-256-q8 | Exploratory, derived | The Q8_0 vectors cut to 256 and re-normalised | — | 256 |
| RRF-nomic, RRF-G2 | Exploratory, derived | Reciprocal rank fusion of BM25 with nomic, and with the reference path (RRF: a window's score is the sum over both lists of 1 / (60 + its rank), BM25's list holding only the windows it scores above zero; k = 60, equal weights, fixed in advance) | — | — |
| RR-minilm-bm25, RR-minilm-rrfg2 | Exploratory | `cross-encoder/ms-marco-MiniLM-L6-v2` re-orders the top 50 of BM25, and of RRF-G2 | — | — |
| RR-bge-bm25, RR-bge-rrfg2 | Exploratory | `BAAI/bge-reranker-v2-m3` re-orders the same two top-50 pools | — | — |

Google's prefixes, as exact strings with their trailing space: `title: none | text: ` on every window (the pages carry no titles, and the card says `title: none` when none is available) and `task: search result | query: ` on every question. They are the strings in the model's own configuration under its `document` and `query` prompt names; the harness passed them as literal bytes, never by name.

**The machine.** One laptop's CPU, an Intel Core Ultra 9 290HX Plus (24 logical processors; AVX2 and AVX-VNNI, no AVX-512 and no AMX), 8 threads per path, 61 GiB of RAM. The laptop has a GPU and the bench never used it: every runtime was built or configured without one (the llama.cpp build links no GPU library, the PyTorch wheel is CPU-only, the private Ollama instance discovered only a CPU), every Ollama call carried `num_gpu: 0`, and a read of the GPU's process list at the start and the end of every run was empty, 12 runs of 12 (six embedding arms, two parity-gate runs and four rerank arms).

**The software.** llama.cpp at commit `4fbc76d` (the merge of #30054, 2026-10-06 16:20:31Z), built from source as a CPU-only release build; the server ran `--embeddings --pooling mean -c 8192 -b 2048 -ub 2048 -t 8 -tb 8 -np 1`. Python 3.13.14; torch 2.14.1 (CPU build); sentence-transformers 6.1.0; transformers 5.19.0 from PyPI; Ollama 0.32.14 for the nomic and EmbeddingGemma 1 arms, as a private instance on a loopback port with its own copy of the model store. The weights: `google/embeddinggemma-2` at revision `914f7f89…` (`model.safetensors`, 1,488,915,288 bytes, all three towers; only the text tower was loaded); `unsloth/embeddinggemma-2-GGUF` at revision `ba388827…` (`embeddinggemma-2-BF16.gguf`, 557,950,240 bytes; `embeddinggemma-2-Q8_0.gguf`, 309,855,520 bytes). Every file's sha256 is in the kit's manifest and was re-checked by the audit after the run.

**What ran, and when.** The run went from 18:35:40Z (the first nomic vector on the CPU) to 19:44:58Z, 2,637 calls per embedding arm. One registered arm could not run: Ollama 0.32.14 refused to pull `embeddinggemma-2:270m-bf16-text` at 18:29:19Z with "412: The model you are attempting to pull requires a newer version of Ollama that may be in pre-release", and the bench upgrades nothing, so `G2-768-ollama` is **NOT RUN** by name. An audit and a second, fresh verification followed, each by a separate AI agent of ours with its own code, not an outside reviewer; both are described below.

## The result {#the-result}

**The gated verdict, in the registered words: NO MATERIAL DIFFERENCE, on both gated arms.**

EmbeddingGemma 2 through Google's own reference code found a gold window in the top 8 for **45 of the 68** questions (recall@8 0.6618). nomic-embed-text, the incumbent, found one for **44 of 68** (0.6471). The difference is +0.0147, one question, against the +0.100 the gate requires. The two models disagree on 17 questions: 9 that only EmbeddingGemma 2 finds and 8 that only nomic finds; McNemar's exact two-sided p = 1.0. The serving path, the 8-bit file on the llama.cpp server, lands on the same 45 of 68, the same 9-against-8 split and the same p. Read under the registered power clause (explained above): no effect of at least 10 points was detected at n = 68. nomic-embed-text keeps its place in our stores on this corpus, now by measurement rather than by default. Nothing is adopted and no proposal is produced. (`e2-1-scores.json` → `ge3`; `arms/G2-768-ref.score.json`; `arms/G2-768-q8.score.json`.)

**Every gate, with its registered words and the number behind it.**

| Gate | What it asks | Result | Verdict | File |
|---|---|---|---|---|
| R-0, the instrument reproduces | Every number of the 2026-08-19 board comes back identical from the cached vectors and the same code | All 24 numbers identical to three decimals; the earlier prefix-gate JSON identical | PASS | `arms/R-0.json` |
| G-0, runtime parity (BF16 GGUF) | Cosine ≥ 0.999 to the reference path's vector on 200 registered windows and all 68 questions | Minimum 0.999986, median 0.999997; 0 of 268 below the floor; 0 NaN, Inf or zero vectors | PASS | `arms/G2-768-bf16.g0.json` |
| G-0, runtime parity (Q8_0 GGUF) | The same | Minimum 0.999507, median 0.999849; 0 of 268 below; 0 NaN, Inf or zero | PASS | `arms/G2-768-q8.g0.json` |
| G-0 (Ollama tag) | The same | The tag could not be pulled on Ollama 0.32.14 | NOT RUN | `arms/G2-768-ollama.score.json`; the registration's Addendum 1 |
| G-E2, integrity | 768 (or the cut's size) on 100.0 % of calls, no NaN, Inf or all-zero vector | 2,637 of 2,637 on every embedding arm; 268 of 268 on each parity set | PASS | `e2-1-scores.json` → `ge2` |
| G-CPU, no GPU | No process of the run's unit in the GPU's process list at start or end | 12 runs (6 embedding, 2 parity, 4 rerank), 24 readings, all empty | PASS | `arms/*.embed.json`, `arms/*.g0-run.json`, `arms/*.rerank.json` |
| R-1, the incumbent reproduces on CPU | `nomic` re-embedded fresh on the CPU equals its cached recall@8 at 4 decimals | 0.6471 = 0.6471; the same 44 questions; fresh-to-cached cosine minimum 0.9999944, median 0.9999983 over 2,637 texts | PASS | `arms/R-1.json` |
| **G-E3, admission (reference path)** | Δ ≥ +0.100 and p ≤ 0.05 | Δ = +0.0147 (45 − 44 of 68); b 9, c 8; p = 1.0 | **NO MATERIAL DIFFERENCE** | `e2-1-scores.json` → `ge3` |
| **G-E3, admission (serving path)** | The same, plus G-0 | Δ = +0.0147; b 9, c 8; p = 1.0 | **NO MATERIAL DIFFERENCE** | `e2-1-scores.json` → `ge3` |

**The rows that decide it**, in plain words (the full board, every arm and every metric, follows):

| Arm | recall@8 (of 68) | Δ vs nomic | `b` / `c` | `p` |
|---|---:|---:|---:|---:|
| Keyword search (BM25), the reference | 0.853 (58) | — | — | — |
| `nomic-embed-text`, the incumbent | 0.647 (44) | — | — | — |
| EmbeddingGemma 2 through Google's code, float32 (gated) | 0.662 (45) | +0.0147 | 9 / 8 | 1.0000 |
| EmbeddingGemma 2 as the Q8_0 file on llama.cpp (gated) | 0.662 (45) | +0.0147 | 9 / 8 | 1.0000 |

The gate asked for +0.100 (51 of 68) at p ≤ 0.05.

**The full board.** recall@k is the share of the 68 questions with at least one gold window in the top k; the number in parentheses is the count. Every row below the two gated rows carries the word EXPLORATORY and no verdict: the registration says an exploratory result that looks large is a reason for a new registration, never a finding of this one. "b / c" are the discordant counts against the comparison arm (b: questions only this row's arm finds; c: only the comparison arm), and p is McNemar's exact two-sided p over them.

| Arm | Class | recall@1 | recall@3 | **recall@8** (of 68) | recall@20 | nDCG@10 | MRR@20 | Δ vs nomic | `b` / `c` | `p` | Δ vs BM25 | `b` / `c` | `p` |
|---|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
| L0-bm25 | Reference | 0.471 | 0.603 | **0.853** (58) | 0.956 | 0.539 | 0.591 | — | — | — | — | — | — |
| A1-asym (nomic) | Incumbent | 0.324 | 0.426 | **0.647** (44) | 0.765 | 0.369 | 0.413 | — | — | — | — | — | — |
| A2-symdoc (nomic) | Reference | 0.294 | 0.426 | **0.588** (40) | 0.750 | 0.348 | 0.395 | −0.0588 | 1 / 5 | 0.2188 | −0.2647 | 2 / 20 | 0.0001 |
| A3-bare (nomic) | Reference | 0.294 | 0.441 | **0.559** (38) | 0.750 | 0.354 | 0.400 | −0.0882 | 2 / 8 | 0.1094 | −0.2941 | 1 / 21 | 0.0000 |
| A1-cpu (nomic, fresh) | Check | 0.324 | 0.426 | **0.647** (44) | 0.779 | 0.369 | 0.413 | +0.0000 | 0 / 0 | 1.0000 | −0.2059 | 2 / 16 | 0.0013 |
| **G2-768-ref** | **gated** | 0.338 | 0.500 | **0.662** (45) | 0.779 | 0.374 | 0.448 | **+0.0147** | **9 / 8** | **1.0000** | −0.1912 | 4 / 17 | 0.0072 |
| **G2-768-q8** | **gated** | 0.324 | 0.515 | **0.662** (45) | 0.779 | 0.370 | 0.439 | **+0.0147** | **9 / 8** | **1.0000** | −0.1912 | 4 / 17 | 0.0072 |
| G2-768-bf16 | EXPLORATORY | 0.324 | 0.500 | **0.662** (45) | 0.779 | 0.372 | 0.441 | +0.0147 | 9 / 8 | 1.0000 | −0.1912 | 4 / 17 | 0.0072 |
| G2-768-ollama | EXPLORATORY | — | — | NOT RUN | — | — | — | — | — | — | — | — | — |
| G2-768-noprefix | EXPLORATORY | 0.191 | 0.368 | **0.529** (36) | 0.632 | 0.266 | 0.306 | −0.1176 | 7 / 15 | 0.1338 | −0.3235 | 6 / 28 | 0.0002 |
| G1-300m (EmbeddingGemma 1) | EXPLORATORY | 0.382 | 0.544 | **0.676** (46) | 0.794 | 0.428 | 0.491 | +0.0294 | 6 / 4 | 0.7539 | −0.1765 | 3 / 15 | 0.0075 |
| G2-512-ref | EXPLORATORY | 0.279 | 0.500 | **0.662** (45) | 0.765 | 0.360 | 0.415 | +0.0147 | 10 / 9 | 1.0000 | −0.1912 | 5 / 18 | 0.0106 |
| G2-256-ref | EXPLORATORY | 0.265 | 0.471 | **0.647** (44) | 0.735 | 0.342 | 0.391 | +0.0000 | 8 / 8 | 1.0000 | −0.2059 | 4 / 18 | 0.0043 |
| G2-128-ref | EXPLORATORY | 0.147 | 0.368 | **0.515** (35) | 0.618 | 0.246 | 0.276 | −0.1324 | 3 / 12 | 0.0352 | −0.3382 | 2 / 25 | 0.0000 |
| G2-256-q8 | EXPLORATORY | 0.265 | 0.471 | **0.618** (42) | 0.735 | 0.338 | 0.391 | −0.0294 | 6 / 8 | 0.7905 | −0.2353 | 3 / 19 | 0.0009 |
| RRF-nomic | EXPLORATORY | 0.338 | 0.559 | **0.765** (52) | 0.941 | 0.447 | 0.482 | +0.1176 | 8 / 0 | 0.0078 | −0.0882 | 2 / 8 | 0.1094 |
| RRF-G2 | EXPLORATORY | 0.338 | 0.588 | **0.824** (56) | 0.897 | 0.455 | 0.488 | +0.1765 | 13 / 1 | 0.0018 | −0.0294 | 5 / 7 | 0.7744 |
| RR-minilm-bm25 | EXPLORATORY | 0.500 | 0.647 | **0.794** (54) | 0.912 | 0.528 | 0.599 | +0.1471 | 15 / 5 | 0.0414 | −0.0588 | 3 / 7 | 0.3438 |
| RR-minilm-rrfg2 | EXPLORATORY | 0.485 | 0.647 | **0.779** (53) | 0.912 | 0.521 | 0.591 | +0.1324 | 14 / 5 | 0.0636 | −0.0735 | 3 / 8 | 0.2266 |
| RR-bge-bm25 | EXPLORATORY | 0.412 | 0.603 | **0.897** (61) | 0.941 | 0.556 | 0.564 | +0.2500 | 19 / 2 | 0.0002 | +0.0441 | 7 / 4 | 0.5488 |
| RR-bge-rrfg2 | EXPLORATORY | 0.412 | 0.603 | **0.897** (61) | 0.941 | 0.540 | 0.562 | +0.2500 | 19 / 2 | 0.0002 | +0.0441 | 7 / 4 | 0.5488 |

*Every row is computed from the per-question ranks in `e2-1-perquery.jsonl` against the gold sets in the frozen sample. The BM25 row and the first three nomic rows (A1-asym, A2-symdoc, A3-bare) are the 2026-08-19 run re-scored; A1-cpu is nomic re-embedded fresh on the CPU for this bench. Each rerank arm's pool of 50 held a gold window for 65 of the 68 questions, the ceiling a reranker could reach.*

**The predictions, registered before any run, by point estimate.** The agent coordinating the bench wrote P-1 to P-3; the agent that wrote the registration added L-1 to L-4. A prediction is not a significance test.

| | Predicted | Result | Verdict |
|---|---|---|---|
| P-1 | EmbeddingGemma 2 beats nomic, but by less than 10 points (45 to 50 hits), reasoning from the card's own English score | Reference path 45 of 68, Q8_0 path 45: each one above nomic's 44 and short of the gate's 51, p = 1.0 | HELD |
| P-2 | BM25 still beats both | 58 against 45 | HELD |
| P-3 | Fusing BM25 with a dense model beats BM25 (59 or more) | RRF-nomic 52, RRF-G2 56 | **REFUTED** |
| L-1 | `nomic` re-run on the CPU finds the same 44 | The same 44 | HELD |
| L-2 | Both GGUF files pass parity, the 8-bit one less closely | Minimum 0.999507 against 0.999986 | HELD |
| L-3 | Cutting to 256 numbers costs at most 2 questions (registered for the reference path) | 44 against 45 on the reference vectors; on the Q8_0 file's vectors the same cut cost three (42 against 45), outside the prediction's scope | HELD |
| L-4 | Dropping the prefixes costs at least 1 question | 36 against 45 | HELD |

**Four exploratory rows worth a sentence each**, with no verdict attached. The two runtimes agree with the reference path at cosine 0.9995 or better and land on the same 45 questions, so a llama.cpp built at that day's merge commit (16:20:31Z) computes the same model as Google's own code; the quantized file's agreement is lower than the 16-bit file's and still 0.0005 above the floor at its worst of the 268 parity texts. EmbeddingGemma 1, the older model, found 46 of 68 (EXPLORATORY, with no parity reference; against nomic, b 6 / c 4, p 0.75): one question above its successor, inside the same noise, and the two cards' English scores also put it slightly ahead (69.67 against 68.46). Cutting the vector to 256 numbers costs one question on the reference vectors (44 against 45), consistent with the card's "close to lossless", and three on the Q8_0 file's vectors (42 against 45); cutting the reference vectors to 128 costs ten (35). Dropping the prefixes costs nine questions (36 against 45): a bare `encode()` call adds no prefix, because the model's configuration sets no default prompt.

## What the keyword search and the hybrids show {#keyword-search-and-the-hybrids}

**BM25 still beats every embedding model on this set, by point estimate.** Plain keyword search finds 58 of 68; EmbeddingGemma 2 trails it by 0.1912 (17 questions only BM25 finds against 4 only EmbeddingGemma finds, p = 0.0072, an exploratory comparison), and the gap over nomic (58 against 44) was already the earlier bench's. Prediction P-2 said so in advance, with the reason: the gold label is "this window contains these words", and the claims turn on dates, names and figures a keyword search matches exactly. An embedder upgrade was not expected to close a 14-question gap, and did not.

**A registered hybrid or rerank bench is the follow-up these rows suggest, not a finding of this run.** Against BM25, no fused or reranked arm beat it by more than chance: BM25 fused with EmbeddingGemma 2 (56) and BM25 fused with nomic (52) sit below BM25's 58, and the best reranked arm, the bge reranker over either first stage's top 50, is 61, three questions above BM25, a gap that chance explains easily (p = 0.55, EXPLORATORY). The large, low-p margins the fused and reranked arms show on the board are against the dense arms, not against BM25. Fusion lifts a dense model's 45 toward BM25's 58 without passing it, and a cross-encoder reading the question beside each candidate lifts the top of the ranking (recall@1 0.500 for the MiniLM reranker over BM25's pool, against BM25's 0.471, the gated arms' 0.338 and 0.324 and nomic's 0.324). The fusion constant, the pool of 50 and the reranker's 512-token limit were fixed in advance and never tuned on this sample (tuning them is a new registration on a held-out sample), and the MiniLM reranker was trained on MS MARCO's questions and passages, where these queries are statements.

## Speed: not cleanly measured, and not compared {#speed}

The registration made timings secondary. Their summaries (first call, documents per second, query p50 and p95, per arm) are in the kit's `e2-1-report.md`; the per-call milliseconds sit in the vector caches, which are not shipped. They are not compared on this page, for a reason the audit put a number on: the bench's own scoring and reranking jobs ran beside the embedding arms with uneven overlap, loading the 24-thread laptop by a mean of 0 extra cores during the nomic arm, 3.3 during the reference arm, 7.8 and 10.1 during the two llama.cpp arms, 7.5 during the no-prefix arm and 5.0 during the EmbeddingGemma 1 arm, while each arm asked for 8 threads. Rows taken under that load would read as a speed gap the data cannot support. Other workloads ran on the laptop too, and this CPU has no AVX-512 and no AMX, so BF16 values are converted in software (it converts float16 in hardware). A clean throughput row needs a quiet machine and its own run; if one is taken it will be added below, dated.

## The audit and the verify {#the-audit-and-the-verify}

Two passes read the run after it finished, each by a separate AI agent of ours writing its own code (not an outside reviewer), and neither changed a number.

**The audit** (19:52:58Z to 20:13Z) re-read every cache, re-ranked every arm with the same tie rule, and recomputed every metric, McNemar p-value and fusion with its own code, importing from the bench only the BM25 scoring function: 111 checks, 0 failed; 443 cells of the run's results notes compared at their printed precision, 0 mismatches. It confirmed the freeze preceded every EmbeddingGemma score from the commit and push logs and the result files' own timestamps. Then it attacked the instrument:

- **A swapped query prefix in a mislabelled arm.** The 68 questions were re-embedded with the document prefix in place of the query prefix, keeping the arm's label: recall@8 moved from 45 to 44. The questions spliced in bare, with no prefix, while the passages kept theirs: 45 to 40. A mislabelled prefix is invisible in the recall number and was caught three ways at the vector level: a fresh re-embed with the declared prefix reproduces the cache at cosine 1.000000 and any other prefix at 0.976 or less; the parity gate fails on all 68 questions (minimum 0.9287); the quantized file's full-run vectors disagree the same way.
- **One gold label removed** from a question with three gold windows: both sample checksums refused to run; bypassed, the reference arm's count moved from 45 to 44 and the discordant split from 9 / 8 to 8 / 8. The scorer reads the labels, and the pins stop a moved label before any scoring.
- **A parity set cast through float16.** Rounding vectors to 16-bit floats changes each component by at most 1.2 × 10⁻⁴ (measured on the reference's query vectors), and the parity gate still passes: the 8-bit file's set, cast that way, keeps its minimum of 0.999507. The gate checks agreement, not compute precision. The hazard the card names is inside the model (activations overflowing float16), and the gate refuses both planted faults that imitate it, a NaN and a vector set degraded to about 0.998.

**The verify** (20:13:44Z to 20:26Z), by a different agent, re-derived every gated number, gate verdict, prediction and table row with its own code: 108 checks, 0 failed; all 20 measured rows and 18 delta rows equal the run's results notes and the scores file. A planted fault (one reference-arm hit moved out of the top 8 in a scratch copy of the per-question file) failed 7 of the same checks, so they can fail and did not pass vacuously. It read the freeze commit (18:24:49Z) and the first EmbeddingGemma recall number (18:50:41Z) from git and the file headers, independently of the audit.

The audit filed no blocking finding and two important ones, both about prose rather than numbers: a sentence in the run's results notes that read a hybrid as a finding (it was reworded to what the rows show, above), and the speed rows' uneven co-load (the paragraph above). It also filed ten minor notes, among them a class label copied wrongly into four derived files' headers, rerank timings no file keeps, and the prefix check proposed below; the verify added four of its own. None changed a number. The verify applied both important fixes; this page's author then read the change once more, as our rule for any fix requires.

## How to run it yourself today, and the traps {#how-to-run-it}

As of 2026-10-06 18:08Z, when we last read the release state for these steps (Ollama's releases and registry were read once more on 2026-10-07, above). The versions will move; the traps mostly will not.

**The reference path (Google's own code, CPU is fine).** In a fresh Python environment: `pip install "sentence-transformers>=6.1.0" "transformers==5.19.0" torch pillow torchvision` (a CPU-only torch build works; `pillow` and `torchvision` are needed even for text-only use, because the model's processor imports them). Then:

```python
import torch
from sentence_transformers import SentenceTransformer
model = SentenceTransformer(
    "google/embeddinggemma-2",
    config_kwargs={"vision_config": None, "audio_config": None},  # text tower only
    model_kwargs={"dtype": torch.float32},                          # never float16
)
model.max_seq_length = 8192   # the default is effectively unbounded
q = model.encode("how do I reset my password", prompt="task: search result | query: ")
d = model.encode("To reset your password, open …", prompt="title: none | text: ")
```

**The llama.cpp path.** Build from source at commit `4fbc76d` or later: at 18:08Z on 2026-10-06 no release contained it (above), so a release that does not contain that merge will not load the file. A CPU-only build is `cmake -B build -DGGML_CUDA=OFF -DGGML_VULKAN=OFF && cmake --build build --config Release --target llama-server`. Serve the Q8_0 or BF16 file from `unsloth/embeddinggemma-2-GGUF` with `llama-server --embeddings --pooling mean -c 8192 -b 8192 -ub 8192 -t 8 -np 1 -m embeddinggemma-2-Q8_0.gguf`, and prepend the prefixes yourself: the server adds none. The model is non-causal, so a whole input must fit in one micro-batch: `-ub` must be at least your longest input, up to 8,192 tokens (our bench ran `-b 2048 -ub 2048`, enough for its longest input of 831 tokens; our preparation test of a 3,368-token input ran at 8,192). `/v1/embeddings` returns L2-normalised vectors. The server has no option to cut the vector; cut to 256 numbers and re-normalise in your own code, knowing that on our set, in exploratory rows, that cut cost three questions on the Q8_0 file's vectors (42 against 45) where it cost one on the float32 reference's (44 against 45).

**The Ollama path.** The 16-bit text-only tag, by its name, is `embeddinggemma-2:270m-bf16-text` (574 MB; the one this bench registered, and the one our bench laptop's Ollama could not pull). On 2026-10-06 no page we read stated the minimum version, and we did not read how the tags are built; read again on 2026-10-07 between 14:41Z and 15:01Z (UTC; above), every `embeddinggemma-2` tag was a safetensors model for Ollama's MLX runner and declared a minimum Ollama of 0.36.0. The bare `:270m` (the same digest as `270m-nvfp4-text`, a 4-bit build by its name) and `latest`/`740m` (the 1.3 GB multimodal model) are other builds; this page measured none of the three. Set the context length explicitly; the library shows "256K" on every tag.

**The traps, each measured or sourced above.**

1. **Never float16.** The card forbids it and says why (activations overflow; NaN or silently degraded vectors). Unsloth's `-F16.gguf` and the ONNX `fp16`/`q4f16` builds exist anyway. Use BF16, float32, or an integer quantization such as Q8_0; the Q8_0 file agreed with the float32 reference at cosine ≥ 0.9995 on all 268 texts we checked.
2. **8K, not 256K.** The card says 8,192 tokens; `config.json`, the GGUF metadata and the Ollama library all say 262,144. Treat 8,192 as the limit and set it on every runtime. (Our longest input was 831 tokens; on the 67 windows that exceed 512, the audit's extra check found both GGUF files agreeing with the reference path at cosine 0.99958 or above. In a smoke test before the run, a 3,368-token input through the BF16 GGUF agreed with the float32 reference at cosine 0.9999976, past both sliding-window sizes the sources disagree on.)
3. **The prefixes are not optional, and nothing adds them for you.** The model's `config_sentence_transformers.json` sets `default_prompt_name` to null, so `encode()` without a `prompt` adds nothing; a server adds nothing. On our set, in an exploratory row, the bare model finds 36 of 68 against 45 with the prefixes. Use `task: search result | query: ` for questions, `title: {title} | text: ` for passages (`title: none` when there is no title), and the same vector size on both sides.
4. **Both models are 768 numbers wide, so mixing them is silent.** A store that held some of each would pass every dimension check and simply return bad neighbours. If you switch, re-embed everything, and keep the model's name and version in a column beside every vector; a `vector(768)` type is not a model check.
5. **Pin the files.** Unsloth's GGUF repository changed at 16:23Z on release day; a re-upload would change the bytes under the same file names. Record the revision and the sha256 of what you ran. Ours: BF16 `f315cbbb…`, Q8_0 `6f1bd4ac…`, revision `ba388827…`.

## What this bench does not say {#what-this-bench-does-not-say}

- **It does not say the two models are equal.** At n = 68 the gate could register only an effect of 10 points or more, and even a 10-point gain had to come almost without losses to reach p ≤ 0.05; one question of difference is inside the noise.
- **It speaks to one corpus**: one town, one subject, one language, one century, with gold labels that favour exact keyword matches by construction and passages that never overlap, so a quote straddling a boundary is unfindable for every arm alike.
- **Our stores hold other content** (a conversation archive, a glossary, a rules index, game worlds), none of it one town's history and none of it measured here; a store heavier in code than this set could see a different answer, and code is where the card's gain is largest. This set is the only retrieval instrument we have, and the bar is set by the cost of re-embedding, whatever a store holds.
- **It says nothing about what the card says the model is best at**: code search, other languages, images, video or audio. A different store with different content could see a different answer.
- **The parity gate shows the runtimes agree with the reference; it cannot catch an error both share.** The code was hours old on both sides.
- **There is no published quality figure for the quantized build**, so our parity gate is the only check the Q8_0 file has had.
- **The timings are not a speed comparison** (above).
- **EmbeddingGemma 1's row has no parity reference**: its Hugging Face weights sit behind a manual-approval gate, and this bench took nothing from behind such a gate, so its number is the Ollama tag's with integrity checks only.
- **The earlier bench's nomic vectors came from an unrecorded device**; the R-1 check shows the recall survives a fresh CPU run exactly (the same 44), so they could serve as the baseline.
- **The freeze can be checked only by us.** The registration's commit sits in a private repository and the digests of the unredacted files are withheld (above), so the kit's redacted copy cannot prove when the original was committed; the witnesses to the order of events are our own agents' logs, quoted above with their times.

## What switching from nomic would take, and why we are not doing it {#what-switching-would-take}

We are not switching: the pass mark was 51 of 68, EmbeddingGemma 2 found 45, and the mark is set by what a switch costs. Our survey of every vector store we knew of (2026-10-06, 18:08Z to 18:25Z) found that cost:

- **Two copies of one nomic release, eight stores.** The Ollama tag this page measured is the embedder for six: the conversation archive (two of the six, one per column), an in-house glossary, this site's related-page vectors, the tool that assembled this bench's documents, and [RealKeep](https://realkeep.strata2signal.com)'s worlds. Nomic's own v1.5 checkpoint wrote RuleSage's retrieval index and one other store, a memory store our AI agents write notes to and read back from. All eight are 768 wide.
- **Every vector recomputed.** The archive and the glossary share one space and switch together; RealKeep's worlds (each a game's saved history, rebuilt by replaying its events) never re-embed on replay, so a switch there means a new world; every store's prefixes change.
- **A guard first.** No store records which installed copy of a model wrote a row (some stamp the model's name, none the bytes that answered), every dimension check is a check for 768, which EmbeddingGemma 2 passes, and only one reader filters by that name. A guard that records the model beside every vector and refuses a mismatched read comes first; it is planned either way.
- **A server that loads it.** Six of the eight stores embed through Ollama, and our bench laptop's Ollama, 0.32.14, refused the model's tag; a switch there needs a release that loads it, or a separate llama.cpp server, after the parity test below.

RuleSage's retrieval code is also wired to a second model, nomic-embed-vision-v1.5, for pictures; as of 2026-10-07 no product sends it a picture, and that picture path does not load (a defect recorded for its own repair, separate from this bench and not yet fixed). Whether EmbeddingGemma 2's single space for text and images could replace that pair is measured on a separate page, [EmbeddingGemma 2's shared space for text, images and audio, measured on what we publish](https://research.strata2signal.com/embeddinggemma-2-shared-space/).

## For this workshop: what changes and what does not {#for-this-workshop}

**Nothing changes.** nomic-embed-text stays the embedder behind every store we run: the eight text stores in the section above, this site's own related-page vectors among them (843 rows in their manifest, generated 2026-10-05). The gate for displacing it was set by the price of that re-embed, and the price did not change on 2026-10-06. The verdict rests on one instrument, 68 questions over one town's history pages, whose content differs from most of those stores; it is the only retrieval instrument we have, and the bar it sets comes from the re-embed's cost, whatever a store holds.

**Why we are glad to have run it anyway.** The incumbent now holds its place by a measurement with a stated power, not by default, against a model released that day, with the whole chain inside four hours of wall clock from the first read of the release (16:30Z) to the verify's close (20:26Z), and the build, registration, run, audit and verify inside two hours and twenty minutes of it (18:06Z to 20:26Z). The next challenger gets the same instrument and the same bar.

**What we will not do.** Adopt an embedder by editing a configuration line (the registration's non-adoption clause: a challenger that clears the gate produces a measured proposal naming every store, every row count and the priced re-embed, which is its own project with its own go). Mix two models' vectors in one store. Run this model in float16. Pull the bare Ollama tags.

**What this points at next**, none of it started when this bench closed on 2026-10-06 (UTC); each is the team's to decide, and a later bench that takes one up reports on its own page, dated, not here:

1. **An Ollama parity test, with an Ollama release that loads the tag** (0.32.14 refused the pull on 2026-10-06; this bench tried no newer release): put `embeddinggemma-2:270m-bf16-text` through the same 268-text parity gate against the reference arm's cached vectors, in a private instance. Six of our eight text stores embed through Ollama, mostly on machines whose releases this bench did not try, so this decides whether a switch there could stay on Ollama or would need a separate embedding server.
2. **A registered hybrid and rerank bench on a held-out sample**, with the fusion constant, the pool and the weights tuned on one half and scored on the other. The exploratory rows above are the reason for that registration, not a result.
3. **A larger and harder question set**: paraphrased questions rather than verbatim claims, and a second subject. Verbatim gold is keyword search's home ground; questions written the way people ask would test what dense models are for.
4. **A prefix-fingerprint check in every future run**: at the end of each arm, re-embed a few texts with the declared prefix and require cosine ≥ 0.9999 to the cache. The audit showed a mislabelled prefix moves recall by one question and is invisible from the score.
5. **A quiet-machine throughput row**, if anyone needs one.

## What to take with you {#what-to-take-with-you}

- **NO MATERIAL DIFFERENCE, in the registered words.** EmbeddingGemma 2 found 45 of 68 against nomic-embed-text's 44 (b 9 / c 8, McNemar exact p = 1.0), through Google's reference code and a quantized file on llama.cpp alike; the gate asked for 51. Read as "no effect of at least 10 points was detected at n = 68".
- **Keyword search still beats every embedding model on this set, by point estimate**: BM25 58 of 68, with prediction P-2 written before the run saying why (the gold is verbatim text). No fusion or reranking beat it by more than chance (BM25 fused with EmbeddingGemma 2 56; the best reranker 61, three above BM25 at p = 0.55; all exploratory).
- **The release-day runtimes compute the same model**: the BF16 and Q8_0 GGUF files on a llama.cpp built from that day's merge agree with the float32 reference at cosine ≥ 0.999986 and ≥ 0.999507 on 268 texts, and land on the same 45 questions.
- **The prefixes are worth nine questions** (36 without, 45 with; exploratory rows, on this set) and nothing adds them for you; the 256-number cut costs one on the reference vectors and three on the Q8_0 file's (44 and 42 against 45); the 128-number cut costs ten (35, reference vectors).
- **Both models are 768 wide, so a mixed store fails silently.** A switch means re-embedding every stored passage, with the model recorded beside every vector; and on 2026-10-06 the Ollama on our bench laptop (0.32.14) refused to pull the new model.

## How to check our work — and see it live {#how-to-check-our-work-and-see-it-live}

A data kit for this page ships separately, redacted for publication (machine names, folder paths, port numbers and internal tool names replaced, each replacement listed in the copy); a dated note here will say when it is up. It carries: both registrations and the earlier bench's report, as redacted copies; the scores file every table above is printed from (`e2-1-scores.json`); the per-question ranks (`e2-1-perquery.jsonl`: for each of the 20 measured arms and 68 questions, 1,360 rows, the hit at 8, the first gold rank, the gold window ids and the top-20 window ids); the parity gate's full output (`e2-1-g0.json`, the five worst texts included); the run manifest with every model file's sha256 and every library version (`e2-1-manifest.json`); the per-arm receipts (`arms/`, with the same redactions, and other jobs' names reduced to a count); the sha256 of every vector cache (`e2-1-vectors.sha256`; the caches themselves, 312,650,414 bytes, are not shipped); and the harness's own report with the 68-row hit map (`e2-1-report.md`). Recompute recall@8 for any arm from the per-question file with one rule: count the questions whose first gold rank is 8 or less (an empty rank counts as a miss), divide by 68. The scores file also prints the scorer's own gate-level label, `INCUMBENT HOLDS`, a code word inherited from the earlier bench meaning "no challenger cleared the margin"; the registered per-arm words are the ones on this page.

To reproduce the model side: the weights are on Hugging Face at [google/embeddinggemma-2](https://huggingface.co/google/embeddinggemma-2) (revision `914f7f89…`) and [unsloth/embeddinggemma-2-GGUF](https://huggingface.co/unsloth/embeddinggemma-2-GGUF) (revision `ba388827…`); the commands are in the section above; nomic-embed-text is `ollama pull nomic-embed-text` (`ollama list` should show its ID as `0a109f422b47`). The 164 pages, the 2,569 windows and the 68 claims are this workshop's own bench sample and are not published; the kit carries their checksum (`8258eeeb…` for the frozen sample) so that a later correction can point at the same bytes.

And the live door: [ask RuleSage a rules question](https://rulesage-live.strata2signal.com/), free and without an account, and open the wrench on the ruling (the panel behind every answer, which shows the search's own timing and the pages it cited). Its meaning-search runs nomic-embed-text v1.5 inside RuleSage's own process; this page measured Ollama's copy of the same release, not RuleSage's.

## The rest of the seminar {#the-rest-of-the-seminar}

- [Three librarians and a careful reader](https://research.strata2signal.com/three-librarians/) — how RuleSage finds the right page: keyword search, dense search and a reranker, with the shipped numbers.
- [Four homes for one reranker](https://research.strata2signal.com/four-homes-for-one-reranker/) — the MiniLM cross-encoder, one of the two that re-ordered the pools above, and what moving it between machines cost.
- [Seat trials](https://research.strata2signal.com/seat-trials/) — how a model earns a place in one of our apps: a frozen set, a registered bar, a measured proposal.
- [A short history of Ollama](https://research.strata2signal.com/a-short-history-of-ollama/) — the server most of our stores embed through, its release cadence and its pre-releases (the refusal above says the model may need one).
- [EmbeddingGemma 2's shared space for text, images and audio, measured on what we publish](https://research.strata2signal.com/embeddinggemma-2-shared-space/) — the companion page: the same model's images and code against the Nomic pair, and its audio with no incumbent, on what we publish; on 1,000 public COCO captions it puts the right photograph first for 729 against the pair's 497 (the pair on transformers 4.57.6; a set most likely inside both models' training data), on 316 of our own code queries 173 against nomic's 123 with the code-retrieval prefix, on our 14 photographs an exploratory lead at 12 scenes, and nothing in our stores changes.

[The whole shelf](https://research.strata2signal.com/) holds the rest: the benches behind the claims we publish, failures included. If there is a measurement you want next, [say so](https://strata2signal.com/contact/); the suggestion box is read.

## Who ran this, and thanks {#who-ran-this-and-thanks}

The models: **EmbeddingGemma 2** (`google/embeddinggemma-2`, Google, Apache 2.0, revision `914f7f89…`); its GGUF conversions from **Unsloth** (`unsloth/embeddinggemma-2-GGUF`, revision `ba388827…`, BF16 and Q8_0); **nomic-embed-text** from Nomic AI through Ollama (manifest digest `0a109f422b47`); **EmbeddingGemma 1** through Ollama (`embeddinggemma:300m`, digest `85462619ee72`, whose licence text in the Ollama store is the Gemma Terms of Use, not Apache 2.0); the rerankers `cross-encoder/ms-marco-MiniLM-L6-v2` (the Sentence-Transformers project) and `BAAI/bge-reranker-v2-m3` (the Beijing Academy of Artificial Intelligence), at pinned revisions. The runtimes: **llama.cpp** (ggml-org) at commit `4fbc76d`, **sentence-transformers** 6.1.0 over **transformers** 5.19.0 and **PyTorch** 2.14.1 (CPU build), and **Ollama** 0.32.14. The licences this shelf has read first-hand are on its [licences](https://research.strata2signal.com/licences/) page, nomic-embed-text's among them; EmbeddingGemma 2's Apache 2.0 tag and the card's prohibited-use sentence are described above as we found them, not as a legal reading.

A small human team asked for this bench on the day the model was released, set the bar before any EmbeddingGemma score existed, chose what the page would and would not claim, and signed the numbers. A fleet of AI agents wrote the registration, built the runtimes from source, ran the arms, audited and verified the run with independent code, and drafted this page under that team's rulings. No visitor's data was involved; the bench ran on the workshop's own documents, on one laptop's CPU, with its GPU idle.

Corrections and later measurements will be added below, each dated (UTC) with a window at both ends where one applies, each saying in plain words what it counts, and each comparing itself in one sentence to the reading it replaces.

## Changes {#changes}

- 2026-10-06 (UTC): drafts v1 to v4 (v2 a critic's fold; v3 a fresh check's three fixes; v4 a fresh critic's round folded, with the section on what switching from nomic would take); nothing released; the data kit ships separately, and a dated note above will say when it is up.
- 2026-10-08 (UTC): draft v5, after a fresh review of v4 by a separate AI agent of ours: four small wording fixes; the release-day Ollama line now records that 0.40.0 became a full release (its tag dated 2026-10-06, read on 2026-10-07), where v4 knew only a 0.40.0 pre-release; the sentence on RuleSage's picture model now says what its sources support (the retrieval code is wired to nomic-embed-vision-v1.5 for pictures, no product sends it a picture, the path does not load, a defect recorded for its own repair), where v4 said the store keeps image vectors; nothing released.
- 2026-10-08 (UTC): draft v6, after two fresh reads of v5 by AI agents of ours (one read as a stranger would read it, one a figures-and-laws check): the chain's wall clock now runs from the first read of the release (16:30Z) where v5 counted from the first build step (18:06Z); the 256-number cut's cost is given for the reference vectors (one question) and for the Q8_0 file's (three), where v5 gave only the reference's; "beat it by more than chance" replaces "separated"; the audit and verify are named as our own agents' passes; "by default" replaces "by contract"; "arm" is defined; the Ollama tags are described as read on each date (by name on 2026-10-06; as MLX-runner models requiring Ollama 0.36.0 on 2026-10-07), where v5 called the text tag plain and its neighbours MLX-built; "none started" is dated to the bench's close; the questions paragraph gives its starting count (113); the short version is shorter (196 to 173 words); a four-row plain-words table precedes the full board; nothing released.
- 2026-10-08 (UTC): the v6 shelf copy, after a fresh check of v6 by a separate AI agent of ours: the power clause is restated in the registration's own words (with 68 questions the test detects only large effects), where v6 said it could reliably detect a gain of 10 points or more; the llama.cpp recipe says a release that does not contain the merge will not load the file, where v6 said a release built before it; the how-to's 2026-10-07 registry reading gives its window (14:41Z to 15:01Z) and its past tense, where v6 gave the date only; the source caption says Google's blog, where v6 said its developer blog; two earlier lines of this log name the reviews as our own agents'; nothing released.
- 2026-10-09 (UTC): published; the byline dated.

<!-- derived 2026-10-09 (UTC) by tools/derive_md.py from the pour source.
     source html sha256: 4febfdfb795a3b1595c314e559fbeccd4dfc7ad5d1c71e2b9db28c20a2458fd4
     derivation sha256:  9e89a1ac8a0a6ac32b80dd906ea56a7bbc63e6c9cf70d934e32f84a4605b7c2a
     the {#id} on each heading is the anchor that heading carries on the page. -->
