{
 "schema_version": "1.0",
 "page_kind": "article",
 "pack_kind": "generate",
 "slug": "embeddinggemma-2-on-release-day",
 "title": "EmbeddingGemma 2, measured on release day against the embedder we already run",
 "dek": "Google's new open embedding model, EmbeddingGemma 2, against nomic-embed-text, the embedder our search stores already run, on 68 questions over our own documents behind a pass mark of 51 of 68 written down before any score: 45 against 44 of 68, the registered verdict NO MATERIAL DIFFERENCE (one question, inside the noise at n = 68), with plain keyword search at 58 of 68 still ahead of both and nothing we layered on it beating it by more than chance.",
 "published": "2026-10-09",
 "series": [
  "bench"
 ],
 "licence": "CC BY 4.0",
 "status": "pending-judge",
 "pack_note": "no judge is seated, so every chip is held",
 "notice": "items with published:false are held and are not this page's published words",
 "source": {
  "url": "https://research.strata2signal.com/embeddinggemma-2-on-release-day/",
  "md_url": "https://research.strata2signal.com/embeddinggemma-2-on-release-day/index.md",
  "md_sha": "6c01f67f6da16647979b613e5a752639e65f52d5fdcfa619926344464fb2f30e",
  "html_sha": "972cab3d6fb6db94bd5c0dd0d7caec6237be8770a22b43aa0dfabee4f420db40",
  "anchors_sha": "89538559d07153413f4d8656fe6224b6bdd3db0bb035f5a916cdc0e4b44367cb"
 },
 "short": {
  "paragraph": "On 2026-10-06 Google released EmbeddingGemma 2, an open embedding model under the Apache 2.0 licence: it turns text (and images, video and audio) into lists of numbers that search can compare. The same day we wrote down a test, froze it in a commit, and ran the model's text part against nomic-embed-text, the embedder our search stores already run, on 68 questions over 2,569 passages of our own documents. nomic-embed-text finds a correct passage in the top 8 for 44 of the 68; EmbeddingGemma 2 finds one for 45, through Google's own code and through an 8-bit copy of its weights on llama.cpp (an open-source model runner) alike. The pass mark, written down before any score existed, was 51 of 68, so the registered verdict is NO MATERIAL DIFFERENCE: one question of difference is inside the noise, which is not the same as saying the two are equal. Plain keyword search (BM25) finds 58 of 68, and nothing we layered on it beat it by more than chance. Nothing in our stores changes.",
  "from_dek": false,
  "counts": {
   "words": 9788,
   "minutes": 44,
   "tables": 9,
   "kit": false
  },
  "bullets": []
 },
 "sections": [
  {
   "id": "what-an-embedding-model-is",
   "heading": "What an embedding model is",
   "level": 2,
   "span": [
    3578,
    5176
   ],
   "chunks": [
    [
     3578,
     5176
    ]
   ],
   "chars": 1598,
   "digest": "semantic search uses arithmetic to compare lists of numbers generated by an embedding model. nomic-embed-text is the model used for the workshop's stored lists. swapping models requires recomputing every passage, which is why the cost of switching is a central concern.",
   "digest_skipped": null
  },
  {
   "id": "the-question",
   "heading": "The question",
   "level": 2,
   "span": [
    5176,
    5885
   ],
   "chunks": [
    [
     5176,
     5885
    ]
   ],
   "chars": 709,
   "digest": "the question asks if EmbeddingGemma 2's text model is better enough than nomic-embed-text to justify recomputing every stored vector. the gate requires an improvement of at least 10 points of recall@8 with a paired significance test at p ≤ 0.05.",
   "digest_skipped": null
  },
  {
   "id": "what-google-released",
   "heading": "What Google released on 2026-10-06, and what we checked",
   "level": 2,
   "span": [
    5885,
    13378
   ],
   "chunks": [
    [
     5885,
     13378
    ]
   ],
   "chars": 7493,
   "digest": "EmbeddingGemma 2 is a second open embedding model from Google. the text part is a 270 M-parameter model. the model card and various sources provide different information regarding context length, sliding attention windows, and float16 usage.",
   "digest_skipped": null
  },
  {
   "id": "the-instrument",
   "heading": "The instrument: 68 questions over our own documents",
   "level": 2,
   "span": [
    13378,
    20325
   ],
   "chunks": [
    [
     13378,
     20325
    ]
   ],
   "chars": 6947,
   "digest": "the test uses 68 factual claims about a historical gold-rush town. the corpus consists of 2,569 passages of at most 1,200 characters. scoring is based on recall@8, using McNemar's exact two-sided test for significance.",
   "digest_skipped": null
  },
  {
   "id": "what-we-ran",
   "heading": "What we ran",
   "level": 2,
   "span": [
    20325,
    24766
   ],
   "chunks": [
    [
     20325,
     24766
    ]
   ],
   "chars": 4441,
   "digest": "the bench ran several arms, including the incumbent nomic-embed-text and various EmbeddingGemma 2 configurations. the machine used was an Intel Core Ultra 9 290HX Plus. the software included llama.cpp, transformers, and Ollama.",
   "digest_skipped": null
  },
  {
   "id": "the-result",
   "heading": "The result",
   "level": 2,
   "span": [
    24766,
    34689
   ],
   "chunks": [
    [
     24766,
     34689
    ]
   ],
   "chars": 9923,
   "digest": null,
   "digest_skipped": null
  },
  {
   "id": "keyword-search-and-the-hybrids",
   "heading": "What the keyword search and the hybrids show",
   "level": 2,
   "span": [
    34689,
    36504
   ],
   "chunks": [
    [
     34689,
     36504
    ]
   ],
   "chars": 1815,
   "digest": "BM25 keyword search outperforms every embedding model on this set. while fusion and reranking lift the performance of dense models, no hybrid arm beat BM25 by more than chance.",
   "digest_skipped": null
  },
  {
   "id": "speed",
   "heading": "Speed: not cleanly measured, and not compared",
   "level": 2,
   "span": [
    36504,
    37564
   ],
   "chunks": [
    [
     36504,
     37564
    ]
   ],
   "chars": 1060,
   "digest": "timings were not cleanly measured or compared because the scoring and reranking jobs ran beside the embedding arms, creating uneven load on the CPU.",
   "digest_skipped": null
  },
  {
   "id": "the-audit-and-the-verify",
   "heading": "The audit and the verify",
   "level": 2,
   "span": [
    37564,
    40932
   ],
   "chunks": [
    [
     37564,
     40932
    ]
   ],
   "chars": 3368,
   "digest": "two separate AI agents performed an audit and a verify. the audit checked for mislabelled prefixes, removed gold labels, and float16 casting. the verify re-derived every gated number and gate verdict with its own code.",
   "digest_skipped": null
  },
  {
   "id": "how-to-run-it",
   "heading": "How to run it yourself today, and the traps",
   "level": 2,
   "span": [
    40932,
    45978
   ],
   "chunks": [
    [
     40932,
     45978
    ]
   ],
   "chars": 5046,
   "digest": null,
   "digest_skipped": null
  },
  {
   "id": "what-this-bench-does-not-say",
   "heading": "What this bench does not say",
   "level": 2,
   "span": [
    45978,
    48200
   ],
   "chunks": [
    [
     45978,
     48200
    ]
   ],
   "chars": 2222,
   "digest": "this bench does not claim the models are equal, nor does it cover code search or other languages. it is specific to one corpus and one subject.",
   "digest_skipped": null
  },
  {
   "id": "what-switching-would-take",
   "heading": "What switching from nomic would take, and why we are not doing it",
   "level": 2,
   "span": [
    48200,
    50459
   ],
   "chunks": [
    [
     48200,
     50459
    ]
   ],
   "chars": 2259,
   "digest": "switching from nomic would require re-embedding every vector in eight stores. a guard to record model names beside vectors is planned.",
   "digest_skipped": null
  },
  {
   "id": "for-this-workshop",
   "heading": "For this workshop: what changes and what does not",
   "level": 2,
   "span": [
    50459,
    53481
   ],
   "chunks": [
    [
     50459,
     53481
    ]
   ],
   "chars": 3022,
   "digest": null,
   "digest_skipped": null
  },
  {
   "id": "what-to-take-with-you",
   "heading": "What to take with you",
   "level": 2,
   "span": [
    53481,
    54976
   ],
   "chunks": [
    [
     53481,
     54976
    ]
   ],
   "chars": 1495,
   "digest": "the verdict is NO MATERIAL DIFFERENCE. keyword search still beats every embedding model on this set. prefixes are important, and mixing models in one store fails silently.",
   "digest_skipped": null
  },
  {
   "id": "how-to-check-our-work-and-see-it-live",
   "heading": "How to check our work — and see it live",
   "level": 2,
   "span": [
    54976,
    57557
   ],
   "chunks": [
    [
     54976,
     57557
    ]
   ],
   "chars": 2581,
   "digest": null,
   "digest_skipped": "credits"
  },
  {
   "id": "the-rest-of-the-seminar",
   "heading": "The rest of the seminar",
   "level": 2,
   "span": [
    57557,
    59329
   ],
   "chunks": [
    [
     57557,
     59329
    ]
   ],
   "chars": 1772,
   "digest": null,
   "digest_skipped": "credits"
  },
  {
   "id": "who-ran-this-and-thanks",
   "heading": "Who ran this, and thanks",
   "level": 2,
   "span": [
    59329,
    61248
   ],
   "chunks": [
    [
     59329,
     61248
    ]
   ],
   "chars": 1919,
   "digest": null,
   "digest_skipped": "credits"
  },
  {
   "id": "changes",
   "heading": "Changes",
   "level": 2,
   "span": [
    61248,
    63890
   ],
   "chunks": [
    [
     61248,
     63890
    ]
   ],
   "chars": 2642,
   "digest": "the log records various drafts and updates to the page between 2026-10-06 and 2026-10-09.",
   "digest_skipped": null
  }
 ],
 "chips": [
  {
   "id": "c-fca23a3a",
   "q": "how does semantic search work with embedding models?",
   "a": "Semantic search uses arithmetic to turn questions and passages into lists of numbers. The passages whose lists point closest to the question's list are selected as candidates. These stored lists are computed once per passage and must be recomputed if the model is swapped.",
   "cites": [
    "what-an-embedding-model-is"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-f97564dc",
   "q": "what metric determines if a new model is worth adopting?",
   "a": "The decision rests on whether a challenger achieves an improvement of at least 10 points of recall@8 over the incumbent model, paired with a significance test where p ≤ 0.05. Anything less is recorded as no material difference.",
   "cites": [
    "the-question"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-5caec16c",
   "q": "what are the technical specifications of embeddinggemma 2?",
   "a": "The text model has 270 M parameters and produces 768-number vectors. It supports a context of 8,192 tokens. Google recommends running inference in bfloat16 or float32 to avoid NaN or degraded vectors caused by activations exceeding float16's range.",
   "cites": [
    "what-google-released"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-46691826",
   "q": "how was the test set for this benchmark constructed?",
   "a": "The test set consists of 164 web pages about a historical gold-rush town, divided into 2,569 passages of at most 1,200 characters. The 68 questions are factual claims used as search text to find gold windows containing specific supporting passages.",
   "cites": [
    "the-instrument"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  },
  {
   "id": "c-c9bd59a6",
   "q": "which specific model arms were tested in this run?",
   "a": "The test included the nomic-embed-text incumbent, the EmbeddingGemma 2 reference path using float32, and the EmbeddingGemma 2 serving path using a Q8_0 GGUF file. Other exploratory arms included various vector cuts, rerankers, and hybrid fusion methods.",
   "cites": [
    "what-we-ran"
   ],
   "published": false,
   "publish_state": "held",
   "quote": "",
   "quote_cite": ""
  }
 ],
 "related": [
  {
   "slug": "embeddinggemma-2-shared-space",
   "why": "shares ground with § Speed: not cleanly measured, and not compared · § Speed and memory: measured, contended, not compared"
  },
  {
   "slug": "ten-minutes-with-living-artists",
   "why": "shares ground with § What this bench does not say · § What to take with you"
  },
  {
   "slug": "the-open-call",
   "why": "shares ground with § The instrument: 68 questions over our own documents · § Three scenarios, one small cove, one draw of its world"
  }
 ],
 "thanks": "## Who ran this, and thanks {#who-ran-this-and-thanks}\n\nThe models: **EmbeddingGemma 2** (`google/embeddinggemma-2`, Google, Apache 2.0, revision `914f7f89…`); its GGUF conversions from **Unsloth** (`unsloth/embeddinggemma-2-GGUF`, revision `ba388827…`, BF16 and Q8_0); **nomic-embed-text** from Nomic AI through Ollama (manifest digest `0a109f422b47`); **EmbeddingGemma 1** through Ollama (`embeddinggemma:300m`, digest `85462619ee72`, whose licence text in the Ollama store is the Gemma Terms of Use, not Apache 2.0); the rerankers `cross-encoder/ms-marco-MiniLM-L6-v2` (the Sentence-Transformers project) and `BAAI/bge-reranker-v2-m3` (the Beijing Academy of Artificial Intelligence), at pinned revisions. The runtimes: **llama.cpp** (ggml-org) at commit `4fbc76d`, **sentence-transformers** 6.1.0 over **transformers** 5.19.0 and **PyTorch** 2.14.1 (CPU build), and **Ollama** 0.32.14. The licences this shelf has read first-hand are on its [licences](https://research.strata2signal.com/licences/) page, nomic-embed-text's among them; EmbeddingGemma 2's Apache 2.0 tag and the card's prohibited-use sentence are described above as we found them, not as a legal reading.\n\nA small human team asked for this bench on the day the model was released, set the bar before any EmbeddingGemma score existed, chose what the page would and would not claim, and signed the numbers. A fleet of AI agents wrote the registration, built the runtimes from source, ran the arms, audited and verified the run with independent code, and drafted this page under that team's rulings. No visitor's data was involved; the bench ran on the workshop's own documents, on one laptop's CPU, with its GPU idle.\n\nCorrections and later measurements will be added below, each dated (UTC) with a window at both ends where one applies, each saying in plain words what it counts, and each comparing itself in one sentence to the reading it replaces.",
 "kit": null,
 "seat_class": {
  "writer": "gemma-class",
  "judge": null,
  "writer_runtime": "vllm"
 },
 "bench": {
  "bullets_written": 3,
  "bullets_kept": 0,
  "digests_written": 15,
  "digests_kept": 12,
  "chips_written": 8,
  "chips_kept": 5,
  "chips_grounded": 0,
  "chips_published": 0,
  "dropped_by": {
   "new_noun": 0,
   "figure": 0,
   "length": 0,
   "cite": 0,
   "judge": 0,
   "quote": 0,
   "directive": 0,
   "redaction": 0,
   "profanity": 0,
   "identifier": 0,
   "bare_figure": 1,
   "house_voice": 0,
   "editorial": 0,
   "hand-edit": 8
  },
  "new_noun_tokens_checked": [
   "0.0147",
   "0.05",
   "0.100",
   "1",
   "10",
   "137",
   "164",
   "192",
   "2",
   "2's",
   "200",
   "2026-10-06",
   "2026-10-09",
   "270",
   "290HX",
   "44",
   "45",
   "569",
   "58",
   "68",
   "768-number",
   "8",
   "831",
   "9",
   "AI",
   "BM25",
   "CPU",
   "Core",
   "DIFFERENCE",
   "EmbeddingGemma",
   "GGUF",
   "Google",
   "Intel",
   "M",
   "M-parameter",
   "MATERIAL",
   "McNemar's",
   "NaN",
   "Ollama",
   "Plus",
   "Python",
   "Q80",
   "Ultra",
   "bfloat16",
   "bm25",
   "float16",
   "float16's",
   "float32"
  ],
  "new_noun_tokens_withheld": 0,
  "figure_definition": "v3",
  "figure_boundary": "guarded",
  "drops": [
   {
    "kind": "bullet",
    "reason": "bare_figure",
    "item": "the longest input measured with EmbeddingGemma 2's tokenizer is 831",
    "detail": "the cited section writes this figure as '831 tokens' and the bullet prints '831' alone"
   },
   {
    "kind": "bullet",
    "reason": "hand-edit",
    "item": null,
    "item_chars": 47,
    "text_sha": "bb560299b311",
    "figure": "270",
    "withheld": "one bullet removed from this pack by hand",
    "removed_utc": "2026-10-09",
    "why": "ungrammatical: the line cuts the cited sentence off before its noun, so a summary line ends on a 270 M-parameter with nothing after it and never names the model it describes, on a page about two models; a summary line a stranger cannot parse fails the stranger read",
    "detail": "removed from this pack by hand on 2026-10-09 (UTC), after the writer ran and after the gates had counted: ungrammatical: the line cuts the cited sentence off before its noun, so a summary line ends on a 270 M-parameter with nothing after it and never names the model it describes, on a page about two models; a summary line a stranger cannot parse fails the stranger read. Not a gate drop, and recorded here so the ledger adds up."
   },
   {
    "kind": "bullet",
    "reason": "hand-edit",
    "item": null,
    "item_chars": 53,
    "text_sha": "ec1d48bf93c2",
    "figure": "137",
    "withheld": "one bullet removed from this pack by hand",
    "removed_utc": "2026-10-09",
    "why": "names no model: the cited sentence gives 137 M parameters as the size, read from the Ollama store, of the nomic-embed-text copy the bench ran as the incumbent; on a page about EmbeddingGemma 2 the line reads as the challenger, and it ends on 137 M-parameter with no noun after it",
    "detail": "removed from this pack by hand on 2026-10-09 (UTC), after the writer ran and after the gates had counted: names no model: the cited sentence gives 137 M parameters as the size, read from the Ollama store, of the nomic-embed-text copy the bench ran as the incumbent; on a page about EmbeddingGemma 2 the line reads as the challenger, and it ends on 137 M-parameter with no noun after it. Not a gate drop, and recorded here so the ledger adds up."
   },
   {
    "kind": "digest",
    "reason": "hand-edit",
    "item": null,
    "item_chars": 202,
    "text_sha": "560b71642a38",
    "section": "the-result",
    "withheld": "one section digest removed from this pack by hand",
    "removed_utc": "2026-10-09",
    "why": "the digest sets 45 of 68 against 44 and gives the bar only as the +0.100 margin; the page prints the pass mark beside that result as 51 of 68, the count that shows the challenger six questions short, and a line that sets 45 against 44 carries it",
    "detail": "removed from this pack by hand on 2026-10-09 (UTC), after the writer ran and after the gates had counted: the digest sets 45 of 68 against 44 and gives the bar only as the +0.100 margin; the page prints the pass mark beside that result as 51 of 68, the count that shows the challenger six questions short, and a line that sets 45 against 44 carries it. Not a gate drop, and recorded here so the ledger adds up."
   },
   {
    "kind": "digest",
    "reason": "hand-edit",
    "item": null,
    "item_chars": 178,
    "text_sha": "c9934f179b78",
    "section": "how-to-run-it",
    "withheld": "one section digest removed from this pack by hand",
    "removed_utc": "2026-10-09",
    "why": "the digest says the model can be run through Ollama; the page says the Ollama on the bench laptop, 0.32.14, could not pull the tag, and the section says every tag read again on 2026-10-07 declared a minimum Ollama of 0.36.0 and that this page measured none of the tags, so no Ollama release is shown here to run it",
    "detail": "removed from this pack by hand on 2026-10-09 (UTC), after the writer ran and after the gates had counted: the digest says the model can be run through Ollama; the page says the Ollama on the bench laptop, 0.32.14, could not pull the tag, and the section says every tag read again on 2026-10-07 declared a minimum Ollama of 0.36.0 and that this page measured none of the tags, so no Ollama release is shown here to run it. Not a gate drop, and recorded here so the ledger adds up."
   },
   {
    "kind": "digest",
    "reason": "hand-edit",
    "item": null,
    "item_chars": 147,
    "text_sha": "3a2335c81305",
    "section": "for-this-workshop",
    "withheld": "one section digest removed from this pack by hand",
    "removed_utc": "2026-10-09",
    "why": "the digest makes an Ollama parity test and a larger question set the next steps of the workshop; the section lists them as what the result points at next, none of it started when the bench closed and each for the team to decide",
    "detail": "removed from this pack by hand on 2026-10-09 (UTC), after the writer ran and after the gates had counted: the digest makes an Ollama parity test and a larger question set the next steps of the workshop; the section lists them as what the result points at next, none of it started when the bench closed and each for the team to decide. Not a gate drop, and recorded here so the ledger adds up."
   },
   {
    "kind": "chip",
    "reason": "hand-edit",
    "item": null,
    "item_chars": 262,
    "text_sha": "a4e70324ea44",
    "chip_id": "c-935a376b",
    "withheld": "one chip removed from this pack by hand",
    "removed_utc": "2026-10-09",
    "why": "the answer sets 45 of 68 against 44 and gives the bar only as the +0.100 margin; the page prints the pass mark beside that result as 51 of 68, six questions above the challenger, and a line that sets 45 against 44 carries it",
    "detail": "removed from this pack by hand on 2026-10-09 (UTC), after the writer ran and after the gates had counted: the answer sets 45 of 68 against 44 and gives the bar only as the +0.100 margin; the page prints the pass mark beside that result as 51 of 68, six questions above the challenger, and a line that sets 45 against 44 carries it. Not a gate drop, and recorded here so the ledger adds up."
   },
   {
    "kind": "chip",
    "reason": "hand-edit",
    "item": null,
    "item_chars": 274,
    "text_sha": "259597989d07",
    "chip_id": "c-253af7a0",
    "withheld": "one chip removed from this pack by hand",
    "removed_utc": "2026-10-09",
    "why": "the answer gives the 58 of 68 of BM25 without the condition the page sets beside it, that no fused or reranked arm beat it by more than chance",
    "detail": "removed from this pack by hand on 2026-10-09 (UTC), after the writer ran and after the gates had counted: the answer gives the 58 of 68 of BM25 without the condition the page sets beside it, that no fused or reranked arm beat it by more than chance. Not a gate drop, and recorded here so the ledger adds up."
   },
   {
    "kind": "chip",
    "reason": "hand-edit",
    "item": null,
    "item_chars": 296,
    "text_sha": "120c8c8ff5db",
    "chip_id": "c-65dcf66e",
    "withheld": "one chip removed from this pack by hand",
    "removed_utc": "2026-10-09",
    "why": "the answer lists the packages the reference path needs with no version; the section pins transformers 5.19.0 and sentence-transformers 6.1.0 or later, as read on 2026-10-06, and the page says support for the model is in transformers 5.19.0, so the list without its pins need not load the model",
    "detail": "removed from this pack by hand on 2026-10-09 (UTC), after the writer ran and after the gates had counted: the answer lists the packages the reference path needs with no version; the section pins transformers 5.19.0 and sentence-transformers 6.1.0 or later, as read on 2026-10-06, and the page says support for the model is in transformers 5.19.0, so the list without its pins need not load the model. Not a gate drop, and recorded here so the ledger adds up."
   }
  ]
 },
 "generated_utc": "2026-10-09T12:13:57Z"
}
