# How an embedding model turns anything into a list of numbers

*An embedding model reads a sentence, a photograph or a few seconds of sound and writes a fixed-length list of numbers, so that a program can find related things by comparing lists. This page explains what happens inside one, how one model can hold text, images, audio and video in one space and hand you vectors of several sizes, where it fails and plain keyword search still wins, and how to choose one. Our own figures come from two bench pages we published on 2026-10-09 (UTC) about Google's EmbeddingGemma 2; everything else is cited to its paper or model card.*

*Published 2026-10-10 (UTC) · A small (human) team and a fleet of AI agents.*

**the short version:** An embedding model turns a piece of content into a vector, a fixed-length list of numbers, placed so that things that belong together land close together; search becomes "which stored vectors point the way the question's does?". The model learns where to put things from millions of pairs that belong together, a question and its answer, a photograph and its caption, which is also how one model can put a photograph beside the sentence that describes it. The rest of this page is the mechanism, with examples small enough to check with a calculator, and what we measured on EmbeddingGemma 2, the open embedding model Google released on 2026-10-06.

7,914 words · about 36 minutes (at 220 words/min) · 2 tables · data kit: no

https://research.strata2signal.com/how-an-embedding-model-works/

---

- **An embedder is a transformer that reads the input, a layer that sets how many numbers come out, and an averaging step that turns its many per-word vectors into one, trained on millions of pairs that belong together.** EmbeddingGemma 2's vectors are 768 numbers long.
- **One model can put photographs and sentences in one space because it was trained on photograph-caption pairs, but the two kinds still sit apart.** We put 1,000 public COCO photographs, one caption each, into one search index with 2,569 text passages about something else: searching by caption, EmbeddingGemma 2 (at its default 280 image tokens) still found the right photograph first for 721 of 1,000. With the other 999 photographs' captions added to that index as text, only 139 did. Nomic AI's image and text models (the image model run on an older release, 4.57.6, of the transformers library it loads through), searched the same way, found none: every passage outranked every photograph.
- **Some vectors can be cut short and still work, if you rescale them afterwards.** EmbeddingGemma 2 is trained so that its first 512, 256 or 128 numbers make a usable vector on their own (Matryoshka training, after the nesting dolls). On our 68-question text test (a right passage in the top 8; 32-bit reference vectors; the cuts exploratory) it found one for 45 questions at the full 768 numbers and at 512, 44 at 256, 35 at 128.
- **Plain keyword search still wins where the answer repeats the question's words.** On that 68-question test, where every right answer contains a quoted source passage word for word, BM25 (the classic keyword-ranking formula) found a right passage in the top 8 for 58 to EmbeddingGemma 2's 45. On 316 searches for a function by its description, it came last: the right function first for 87 of 316, against EmbeddingGemma 2's 173 with its code-retrieval prefix.
- **Two models with the same vector length do not share a space.** Mix two 768-number models' vectors in one index and search returns wrong neighbours without an error; even one model, run through two different programs, can drift.

*If you're new here: [strata→signal](https://strata2signal.com) is a small workshop that builds on its own machines and writes up what it measures. Every bench figure in the body of this page is from the two pages named below, quoted with its condition and linked to the section it was read from; everything else is from the paper or model card named beside it. Machines are named by their role, never their hostname.*

*What we measured on.* Our figures come from two bench pages published on 2026-10-09 (UTC), [EmbeddingGemma 2, measured on release day](https://research.strata2signal.com/embeddinggemma-2-on-release-day/) (text; "the release-day page" below) and [EmbeddingGemma 2's shared space for text, images and audio](https://research.strata2signal.com/embeddinggemma-2-shared-space/) (images, audio, code; "the shared-space page"), on these sets:

- **The 68-question text test**: 68 factual claims about one historical town, each used as a search over 2,569 passages of at most 1,200 characters cut from 164 web pages about the town. A right answer is a passage that holds, word for word, a source passage the claim was checked against, which favours keyword search. The figure is how many of the 68 find a right passage in the top 8 results (**recall@8**; recall@k in general counts hits in the top k) ([release-day page](https://research.strata2signal.com/embeddinggemma-2-on-release-day/#the-instrument)).
- **1,000 COCO photographs**: COCO is a widely used public set of everyday photographs with human-written captions. We took the first 1,000-image fold of its standard 5,000-image test split, searched by each photograph's first caption, and counted the captions whose own photograph came first (**recall@1**) ([shared-space page](https://research.strata2signal.com/embeddinggemma-2-shared-space/#what-the-images-show)).
- **Our own images**: 14 photographs, in 12 scenes, with the 40 descriptions published beside them, and 23 other released images (screenshots and paintings among them) with their 39 published texts, all on this site ([shared-space page](https://research.strata2signal.com/embeddinggemma-2-shared-space/#what-we-tested)).
- **316 code searches**: each a docstring (the description written at the top of a function) from code in our published data kits, searching 746 functions for the one it belongs to; the right function ranked first counts ([shared-space page](https://research.strata2signal.com/embeddinggemma-2-shared-space/#what-the-code-shows)).

A figure marked **exploratory** was measured but was not part of the test written down before the run, so it is a hint, not a verdict.

## The core idea: a position on a map {#the-core-idea}

An **embedding model** (an embedder) reads an input and writes a **vector**, a fixed-length list of numbers; EmbeddingGemma 2 writes 768 of them for a sentence, a photograph or a few seconds of sound alike ([its model card](https://huggingface.co/google/embeddinggemma-2)). Each number is a coordinate, so a vector is a point in a space of 768 directions, or **dimensions**, and the model is trained so that inputs that belong together land close together. Search becomes geometry: turn every document into a vector once and store it, turn each question into one when it arrives, return the stored vectors nearest to it. [RuleSage](https://rulesage.strata2signal.com/), our board-game rules helper, searches this way beside a keyword search.

"Nearest" usually means highest **cosine similarity**, the cosine of the angle between two vectors: their **dot product** (multiply the lists number by number and add up) divided by both vectors' **lengths** (a vector's length is the square root of the sum of its numbers squared), from 1 (same direction) through 0 (at right angles) to −1 (opposite). A **runtime** is the program that loads a model's file and runs it, and some return vectors already scaled to length 1, **normalised**: the server in llama.cpp (an open-source model runner) does ([release-day page](https://research.strata2signal.com/embeddinggemma-2-on-release-day/#how-to-run-it)), and so does Google's own code for the model, built on the sentence-transformers library and called "Google's code" below, whose pipeline ends in a normalise step (the `modules.json` published with [the card](https://huggingface.co/google/embeddinggemma-2)). For such vectors cosine and dot product are the same number ([Karpukhin et al., 2020, "Dense Passage Retrieval for Open-Domain Question Answering", arXiv:2004.04906](https://arxiv.org/abs/2004.04906), §3.1).

**A toy to check by hand.** Real models do not label their dimensions, but for teaching, take three labelled dimensions, *animals, home, cars*, and texts already placed at length 1:

- the question "my dog keeps chewing the sofa": (0.8, 0.6, 0.0)
- A, "stopping a puppy chewing furniture": (0.6, 0.8, 0.0)
- B, "choosing a sofa fabric": (0.0, 1.0, 0.0)
- C, "replacing a car's brake pads": (0.0, 0.0, 1.0)
- D, "a dog's first car journey": (0.6, 0.0, 0.8)

The cosine with A is 0.8 × 0.6 + 0.6 × 0.8 + 0 × 0 = **0.96**; with B, 0.6; with D, 0.48; with C, 0. A first and C last is what a person would pick too.

One caution: in a real model no dimension means "animals". Deep models "tend to diffuse 'information' across the entire representation vector" ([Kusupati et al., 2022, "Matryoshka Representation Learning", arXiv:2205.13147](https://arxiv.org/abs/2205.13147), §1); what the model learned is spread across directions, not kept in named columns. And a cosine is not a probability, nor on one scale across models. In a trained model unrelated inputs rarely score near 0: Liang and colleagues found the average cosine between pairs of inputs inside three trained models between 0.47 and 0.56, with even the minimum positive ([Liang et al., 2022, "Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning", arXiv:2203.02053](https://arxiv.org/abs/2203.02053), §2.1), and EmbeddingGemma 2 puts a caption at an average cosine of 0.546 to a passage about something else ([shared-space page](https://research.strata2signal.com/embeddinggemma-2-shared-space/#what-the-images-show)); the modality-gap section below shows how far apart two models' scales sit.

## From text to a vector {#from-text-to-a-vector}

A text embedder is a **transformer**, the network type behind chat models, with different last steps. For EmbeddingGemma 2's text model, from its card:

1. **Tokens.** The text is cut into **tokens**, word pieces from a fixed vocabulary of 262,144.
2. **A vector per token**, looked up in a table with one row of 512 numbers per vocabulary entry. The card counts the table at 140 M **parameters** (the learned numbers a model is made of; the card calls the table the "embedder", a narrower use of the word than this page's) beside a 130 M-parameter transformer: 262,144 × 512 is about 134 million, so about half the text model is that table.
3. **Context.** 24 transformer layers, each 512 numbers wide, let every token's vector take in the others, in both directions (**bidirectional attention**), where a text-generating model looks only backwards ([Schechter Vera et al., 2025, "EmbeddingGemma: Powerful and Lightweight Text Representations", arXiv:2509.20354](https://arxiv.org/abs/2509.20354), §1, §2.1). That report describes the first EmbeddingGemma; EmbeddingGemma 2 has no report yet, so where this page describes training it leans on that report and on the new model's card. (Runtimes call this "non-causal": [release-day page](https://research.strata2signal.com/embeddinggemma-2-on-release-day/#how-to-run-it).)
4. **Projection.** A projection maps each token's 512 numbers to the 768 the model publishes (the card's "512→768"), so the output width is chosen here, independently of the transformer's own width. In the code Google ships this step comes before the average, token by token (the `1_Pooling` configuration published with the card pools 768-wide vectors).
5. **Pooling** collapses the per-token vectors into one; EmbeddingGemma 2 takes their average (**mean pooling**). This is how an input of any length becomes one fixed-length vector. One average has to stand for everything in its input, which is one reason long documents are cut into passages before embedding; our text bench cut 164 web pages into 2,569 passages of at most 1,200 characters ([release-day page](https://research.strata2signal.com/embeddinggemma-2-on-release-day/#the-instrument)).
6. **Normalise** to length 1.

Why not average an off-the-shelf language model's outputs and skip the training? Sentence-BERT tried that with BERT (Google's 2018 language model) and found it "yields rather bad sentence embeddings, often worse than averaging GloVe embeddings" (an older word-vector method); fine-tuning on sentence pairs fixed it. The same paper shows why one vector per text pays: the most similar pair among 10,000 sentences took about 65 hours to find with BERT reading every pair together, about 5 seconds with one vector per sentence ([Reimers and Gurevych, 2019, "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks", arXiv:1908.10084](https://arxiv.org/abs/1908.10084)). What makes the space useful is the training.

## Contrastive training: learning from pairs {#contrastive-training}

An embedder learns from **pairs** that belong together. The card for nomic-embed-text-v1.5, an open text embedder from Nomic AI and the one our own stores run (the search indexes our products keep; this page compares EmbeddingGemma 2 against it throughout, and with its companion image model, nomic-embed-vision-v1.5, it makes up what we call the Nomic pair), lists forum question-answer pairs and title-body pairs for its first stage, then "higher quality labeled datasets such as search queries and answers from web searches" ([its model card](https://huggingface.co/nomic-ai/nomic-embed-text-v1.5)).

Training takes a **batch** of pairs, embeds all of them, and scores every question against every passage. True pairs sit on the diagonal of that grid; every other cell is a wrong pairing, a **negative**, free with the batch (**in-batch negatives**), and the **loss**, the single number training nudges the model to make smaller, rewards each true pair for outscoring the wrong ones in its row. This **contrastive loss** (InfoNCE, in the papers) is the field's standard recipe. OpenAI's CLIP, the 2021 image-and-text model this page returns to below, traces it to earlier work ([Radford et al., 2021, "Learning Transferable Visual Models From Natural Language Supervision", arXiv:2103.00020](https://arxiv.org/abs/2103.00020), §2.3; [van den Oord, Li and Vinyals, 2018, "Representation Learning with Contrastive Predictive Coding", arXiv:1807.03748](https://arxiv.org/abs/1807.03748)); the first EmbeddingGemma's report ([Schechter Vera et al., 2025](https://arxiv.org/abs/2509.20354), §2.2, beside two other losses) and the Gemini Embedding 2 report ([Shanbhogue et al., 2026, "Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini", arXiv:2605.27295](https://arxiv.org/abs/2605.27295), §3.2) both train with it. Real recipes mask the cells that are accidentally right (two identical questions in one batch; the EmbeddingGemma report's equation 3), and CLIP scores the columns as well as the rows (§2.3).

For one row: divide each cosine by a small number τ, the **temperature**; raise *e* to each; take the true pair's share of the total; the loss is minus its logarithm (a softmax followed by cross-entropy, if those names are familiar). **By hand**, with cosines 0.8 for the true passage and 0.3 and 0.1 for two others:

- τ = 1: *e*^0.8, *e*^0.3, *e*^0.1 are 2.23, 1.35, 1.11; the true pair's share is 0.475; loss 0.74.
- τ = 0.1: *e*^8, *e*^3, *e*^1 are 2,981, 20.1, 2.72; share 0.992; loss 0.008.
- Swap the scores, so a wrong passage gets 0.8 and the true one 0.3: at τ = 0.1 the share falls to 0.0067 and the loss jumps to 5.0.

A small temperature makes the loss care sharply about which passage wins. CLIP learned its temperature from a start equivalent to 0.07, and used 32,768 pairs per batch, so each true pair was scored against 32,767 wrong ones at every step (§2.5).

Two refinements matter in practice. **Hard negatives** are wrong answers that look right: Dense Passage Retrieval's best model added one per question, a passage keyword search ranked high that did not hold the answer ([Karpukhin et al., 2020](https://arxiv.org/abs/2004.04906), §3.2). And **stages**: the first EmbeddingGemma trained on a large, noisy set of pairs, then a smaller, cleaner one with hard negatives, while also learning to copy a larger model's vectors (**distillation**, from Gemini Embedding; report §2.2, §2.3). None of these text models starts from nothing: nomic's begins "from a long-context BERT model" (its card), EmbeddingGemma from a Gemma 3 language model (report §1), and Gemini Embedding 2 from Gemini, which its report calls the embedding model's "pre-training" stage (§3.1); CLIP, below, is the exception that trained from scratch. The pairs reshape a network that already reads, so an embedder learns not meaning in general but whatever its training pairs said belongs together.

## Images, audio and video, and the shared space {#images-audio-video}

A transformer reads a sequence of vectors, so every other kind of input, every other **modality**, needs its own **encoder**: a whole separate model (a **tower**, as in CLIP's two-tower design) or a front end that turns raw data into a sequence a shared transformer reads.

**Images.** The Vision Transformer cuts an image into a grid of small square **patches** and feeds them in like words ([Dosovitskiy et al., 2020, "An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale", arXiv:2010.11929](https://arxiv.org/abs/2010.11929)). The Gemma 4 report (Gemma 4 is Google's open model family) describes its vision encoders as Vision Transformers with 16-pixel patches that keep the image's shape, pooled down to at most a set number of **soft tokens**, vectors that take the place of word tokens in the transformer's input but come from the image rather than a vocabulary ([Gemma Team, 2026, "Gemma 4 Technical Report", arXiv:2607.02770](https://arxiv.org/abs/2607.02770), §2.1). EmbeddingGemma 2, whose card says it "builds upon the architectural and capability advancements of Gemma 4", gives an image up to 280 soft tokens by default (one of five budgets from 70 to 1,120, by setting), and describes its vision encoder no further; our 1,600-pixel photographs took 256 to 273 tokens at the default, because an image is resized keeping its shape ([shared-space page](https://research.strata2signal.com/embeddinggemma-2-shared-space/#how-to-run-it)).

**Audio and video.** Google says EmbeddingGemma 2 shares Gemma 4's audio encoder ([announcement, 2026-10-06](https://blog.google/innovation-and-ai/technology/developers-tools/embeddinggemma-2/)), which the Gemma 4 report describes as taking audio "in 40ms chunks with Mel filterbank inputs" (a **spectrogram**, energy per frequency band over time, on a hearing-based scale) through Conformer layers (a transformer variant with added convolutions, built for speech) (§2.2); the card gives 25 tokens per second of mono audio at 16 kHz. Video is sampled as still frames, by default one per second at 140 tokens each. All of it shares one 8,192-token budget, so one input can mix modalities; about 29 images, 58 video frames or 327 seconds of audio fill it (card). Where those tokens go, the card says this much: a photograph, a clip or a sound sits inside the text as a placeholder token from the model's own vocabulary, every modality draws on "a single shared 8,192-token context", and a mixed input, a photograph with its caption, say, "returns a single embedding". The rest follows the design the next paragraph calls the second, which its card points to: one transformer reads the whole sequence, and the same projection, pooling and normalising that end the text pipeline turn it into one 768-number vector.

**What makes the space shared is training, not plumbing.** Two encoders trained separately have no reason to agree on coordinates. CLIP trained an image encoder and a text encoder together, from scratch, with the contrastive loss above on 400 million (image, text) pairs from the web ([Radford et al., 2021](https://arxiv.org/abs/2103.00020), §2.2 to §2.5). Two designs have followed. **A pair of models trained to share a space**: Nomic's image model was trained against its finished text model, held fixed (its parameters not updated) ([Nussbaum, Duderstadt and Mulyar, 2024, "Nomic Embed Vision: Expanding the Latent Space", arXiv:2406.18587](https://arxiv.org/abs/2406.18587), §1). **One model whose front ends feed one transformer**: Gemini Embedding 2, per its report, a bidirectional transformer initialised from Gemini with mean pooling and a projection, trained contrastively on "single-modality tasks, multimodal tasks, as well as cross-modal tasks" ([Shanbhogue et al., 2026](https://arxiv.org/abs/2605.27295), §3.1, §3.2). For EmbeddingGemma 2 we found no training report as of 2026-10-09 (UTC); its card lists "paired text, image, video, and audio data to align representations across different modalities" among its training data, and the announcement says it is "built from the same technology as Gemini Embedding models". Its card points to the second design: the placeholder tokens, the single shared context and the single embedding for a mixed input, above, are one transformer's.

On 1,000 COCO photographs (the public captioned set described above; one fold of its standard test split, one caption each), EmbeddingGemma 2 put the right photograph first for 729 captions at its default 280 image tokens, the Nomic pair for 497 (with the query recipe Nomic's model cards give, the better of the two we ran, and on an older release of the transformers library, 4.57.6, because the release our own product pins refuses to load Nomic's image model) ([shared-space page](https://research.strata2signal.com/embeddinggemma-2-shared-space/#what-the-images-show)); the evaluation section below says why a lead on a public set is weaker evidence than it looks.

## The modality gap {#the-modality-gap}

A shared space does not mix images and text evenly. Liang and colleagues (cited above) showed that in CLIP-style models the two sit "at arm's length", in separate regions of the space ([Liang et al., 2022](https://arxiv.org/abs/2203.02053)): each freshly initialised encoder's outputs fill a narrow cone of its own, so the modalities start apart (§2), and at CLIP's low learned temperature the contrastive loss was lowest with the gap in place (§4).

It bites when images and text share an index, and it shows how far apart two models' cosine scales can sit. On COCO ([shared-space page](https://research.strata2signal.com/embeddinggemma-2-shared-space/#what-the-images-show)): with EmbeddingGemma 2 a caption sits at an average cosine of 0.743 to its own photograph, 0.535 to a random other photograph and 0.546 to a random passage of unrelated text; with the Nomic pair (transformers 4.57.6), 0.066 (our product's recipe) or 0.086 (the model cards' recipe) to its own photograph and about 0.49 to a random passage: to a Nomic caption, any unrelated passage looks closer than its own photograph. We then put the 1,000 photographs in one list with 2,569 passages about something else:

- EmbeddingGemma 2 still put the right photograph first for 721 of 1,000 captions, against 729 with photographs alone.
- The Nomic pair, an image model fitted to a finished text model, the first design above, for none: every passage outranked every photograph.
- With the other 999 captions added as documents (text about the same things), EmbeddingGemma 2 kept the photograph first for 139; a caption's nearest neighbour was usually another caption.

An obvious repair, standardising each modality's scores separately before ranking (a per-modality z-score), helped on COCO: EmbeddingGemma 2 beside the related text went from 139 to 405, and the Nomic pair beside the unrelated passages from none to 323 or 443 (its two recipes). On our own 37 images beside the same 2,569 passages, though, it took EmbeddingGemma 2 from 0.792 (recall@1 with each of 12 scenes weighted equally) to 0.000. All exploratory ([shared-space page](https://research.strata2signal.com/embeddinggemma-2-shared-space/#what-the-images-show)), and no general fix. The rule: **one model is not one ranked list**. Filter by modality, or rank each separately and merge.

## Matryoshka sizes: one vector, several lengths {#matryoshka-sizes}

A model's width is set by its projection (768 for EmbeddingGemma 2; 3,072 for Gemini Embedding 2, [its report](https://arxiv.org/abs/2605.27295), §3.2), and every extra number costs storage and search time for every item, forever. **Matryoshka Representation Learning** (MRL, after the nesting dolls) trains one vector so that its first numbers already make a good shorter vector: the loss is computed on several prefixes (the first 128 numbers, the first 256, and so on) and summed ([Kusupati et al., 2022](https://arxiv.org/abs/2205.13147), §3), so the model learns to put the most useful information first; an ordinary model's vector, cut to its first numbers or compressed after the fact by SVD (singular value decomposition, the standard way to keep a matrix's strongest directions), lost far more accuracy at small sizes in the same paper (§4.2; the plain cut is its "random feature selection" baseline, Appendix D). EmbeddingGemma 2 is trained this way for 768, 512, 256 and 128 (card).

**Keep the first numbers, then normalise again.** By hand in four dimensions: a question (0.5, 0.5, 0.5, 0.5), d1 = (0.2, 0.4, 0.4, 0.8) and d2 = (0.1, 0.7, 0.1, 0.7), all length 1; full cosines 0.9 and 0.8, so d1 is the better match. Keep the first two numbers: (0.5, 0.5), (0.2, 0.4) and (0.1, 0.7). Their lengths are now 0.707, 0.447 and 0.707, so the raw dot products, 0.3 and 0.4, put d2 first: the cut kept more of d2's length, not more of its meaning. Divide by the lengths, which is what normalising again does, and they are cosines again, 0.949 and 0.8: d1 is back on top. (The question's own length is the same for every document, so only the documents' lengths reorder anything.) This bites wherever a store scores by plain dot product, trusting every vector to be length 1; a cosine score divides by the lengths itself. The card warns that skipping the second step "degrades ranking quality silently", and that queries and documents must share a size.

| Size | Card: MTEB English v2 | Card: MTEB multilingual v2 | Card: MTEB code v1 | Card: MMEB v2 overall | Ours: questions with a right passage in the top 8, of 68 |
|---:|---:|---:|---:|---:|---:|
| 768 | 68.46 | 61.36 | 78.68 | 59.01 | 45 |
| 512 | 68.41 | 61.17 | 77.24 | 58.38 | 45 |
| 256 | 67.78 | 60.41 | 76.18 | 56.24 | 44 |
| 128 | 65.68 | 57.89 | 71.41 | 45.65 | 35 |

*Card columns: average scores on MTEB, the standard public benchmark for text embedders (its English, multilingual and code collections), and on MMEB v2, a multimodal embedding benchmark (both under Evaluation below), higher is better, full-precision checkpoint ([EmbeddingGemma 2 card](https://huggingface.co/google/embeddinggemma-2), truncation table). Ours: recall@8 (questions with a right passage among the top 8 results) on 68 questions over 2,569 passages of our own documents, the text model in 32-bit floats through Google's code, each cut re-normalised; the cuts are exploratory rows ([release-day page](https://research.strata2signal.com/embeddinggemma-2-on-release-day/#the-result)).*

Both say little is lost down to 256 and a real drop comes at 128; the card calls 256 "close to lossless". The card's multimodal column falls further, and the card says 128 "degrades multimodal quality substantially and should be validated against your own workload before adoption". On our 14 photographs a 256-number cut was free when searching for a photograph by its description and cost one scene of 12 when searching for a description by its photograph (exploratory; [shared-space page](https://research.strata2signal.com/embeddinggemma-2-shared-space/#what-the-images-show)). One caution from our text bench: cutting to 256 the vectors from an 8-bit copy of the model (its weights rounded to 8 bits, served by llama.cpp; next section) cost three questions (42 against 45), where cutting Google's 32-bit vectors cost one (both exploratory; [release-day page](https://research.strata2signal.com/embeddinggemma-2-on-release-day/#the-result)). Savings stack, and so can losses.

## Storage precision: how many bytes per number {#storage-precision}

"Precision" means three different things around an embedder, and only one of them sets your database's size.

**The precision the model computes in.** EmbeddingGemma 2's card: "Run inference in `bfloat16` or `float32`. Do not use `float16`." Its internal values exceed float16's range, and it then "returns NaN or silently degraded embeddings rather than raising an error". (float32 is 4 bytes per number; float16 and bfloat16 are 2, and bfloat16 keeps float32's range, per the card.)

**The precision of the weights.** Runtimes such as llama.cpp load **quantised** files, weights rounded to 8 bits or fewer. Against Google's 32-bit reference, EmbeddingGemma 2's 8-bit file (309,855,520 bytes) agreed at cosine 0.999507 or better on 268 texts and found the same 45 of 68; a 6-bit file, 0.997544 and 44; a 5-bit file, 0.991992 and 42, both under the 0.999 the agreement check asks for but run on the questions anyway; a 4-bit file fell to 0.966964, low enough that the check counts it broken, and was never run on them (the three smaller files exploratory; [release-day page](https://research.strata2signal.com/embeddinggemma-2-on-release-day/#the-result); [shared-space page](https://research.strata2signal.com/embeddinggemma-2-shared-space/#what-the-code-shows)). Read plainly: the 8-bit copy found exactly what the reference found, and each smaller file cost more; only a measurement on your own test says where yours stops.

**The precision of the stored vectors**, independent of the other two:

| Stored as | Bytes per number | One vector of 768 | One vector of 256 | A million vectors of 768 |
|---|---:|---:|---:|---:|
| float32 | 4 | 3,072 bytes | 1,024 bytes | 3.07 GB |
| float16 or bfloat16 | 2 | 1,536 bytes | 512 bytes | 1.54 GB |
| int8 | 1 | 768 bytes | 256 bytes | 0.77 GB |
| binary, one bit per number | 1/8 | 96 bytes | 32 bytes | 0.10 GB |

*Arithmetic: vectors only, without the index built on them; GB is 10⁹ bytes.*

Our one measurement of the second row: rounding EmbeddingGemma 2's float32 vectors to float16 moved each number by at most 0.00012 (on our 68 questions' vectors), and the 8-bit copy's vectors, rounded the same way, still agreed with the reference at cosine 0.999507 or better, unchanged to six places. The check was made to test our agreement gate, not search; recall with float16 vectors was not measured ([release-day page](https://research.strata2signal.com/embeddinggemma-2-on-release-day/#the-audit-and-the-verify)). Storing in float16 is not the hazard the card forbids; computing in it is.

The smaller rows need a method. **int8** maps each dimension's range, read from a sample of real vectors, onto 256 levels; **binary** keeps each number's sign and compares vectors by counting differing bits (the **Hamming distance**). Shakir, Aarsen and Lee report that binary vectors with **rescoring** (fetch a few times more candidates by Hamming distance, then re-score them against the full-precision question) kept "up to ~96%" of retrieval performance at 32 times less storage on their best model, mxbai-embed-large-v1 on MTEB's 15 retrieval sets, about 92.5 % without, and warn that "quantization doesn't universally work with all embedding models" ([Shakir, Aarsen and Lee, 2024-03-22, "Binary and Scalar Embedding Quantization for Significantly Faster & Cheaper Retrieval", Hugging Face blog](https://huggingface.co/blog/embedding-quantization)); rescoring comes from Binary Passage Retriever, a model trained from the start to produce binary codes (not binarised after training), which shrank a passage index from 65 GB to 2 GB without losing accuracy on two question-answering benchmarks ([Yamada, Asai and Hajishirzi, 2021, "Efficient Passage Retrieval with Hashing for Open-domain Question Answering", arXiv:2106.00882](https://arxiv.org/abs/2106.00882)). Some embedders are trained with this in mind: the first EmbeddingGemma added a "spread-out" loss partly so that its vectors survive storage quantisation (report §2.2). To measure it on your own data: store the same vectors in each shape, and check against float32 how many of each document's nearest neighbours survive and how far recall moves.

## Task prefixes: telling the model what a text is for {#task-prefixes}

A question and its answer are not alike: one is short and asks, the other is long and states. So many embedders are trained with a short fixed string in front of each text saying which side it is on, and its tokens are averaged into the vector with the text's (the `1_Pooling` configuration published with EmbeddingGemma 2's card includes the prompt): E5, a 2022 family of open embedders, "break[s] the symmetry by adding two prefix identifiers 'query:' and 'passage:'" ([Wang et al., 2022, "Text Embeddings by Weakly-Supervised Contrastive Pre-training", arXiv:2212.03533](https://arxiv.org/abs/2212.03533), §4.1), and nomic-embed-text's card says a prefix such as `search_query: ` "*must*" be included. EmbeddingGemma 2's card lists seven, among them `task: search result | query: ` for a search question and `title: {title} | text: ` for a document (`title: none` without one). Images, audio and video take none.

On our text bench, EmbeddingGemma 2 found 36 of 68 without prefixes and 45 with them (exploratory; [release-day page](https://research.strata2signal.com/embeddinggemma-2-on-release-day/#the-result)). On 316 code-search queries over our published code, the *right* prefix mattered too: the code-retrieval prefix put the right function first for 173 of 316 (0.547), the search prefix for 0.497, none for 0.402, and Nomic's `search_query: ` prefix, the other model's convention, for 0.335 (the last three are extra runs by our independent check after the test closed, so they inform but decide nothing; [shared-space page](https://research.strata2signal.com/embeddinggemma-2-shared-space/#what-the-code-shows)).

Two traps. **Nothing adds the prefix for you**: EmbeddingGemma 2's configuration sets no default prompt, so a bare `encode()` call in Google's code sends the text unprefixed, and a llama.cpp server adds none. And **a wrong prefix hides**: re-embedding the 68 questions with the document prefix moved recall by one question (45 to 44), invisible in a score, while at the vector level the declared prefix reproduced the stored vectors at cosine 1.000000 and any other at 0.976 or less ([release-day page](https://research.strata2signal.com/embeddinggemma-2-on-release-day/#the-audit-and-the-verify)). That gives a cheap check for the end of every embedding job: re-embed a few stored texts with the prefix you think you used, and require a cosine of 0.9999 or more to the stored vectors ([release-day page](https://research.strata2signal.com/embeddinggemma-2-on-release-day/#for-this-workshop)).

## Where embedders fail, and why keyword search still wins on exact matches {#where-embedders-fail}

**BM25** is the classic keyword-ranking formula behind most full-text search ([Robertson and Zaragoza, 2009, "The Probabilistic Relevance Framework: BM25 and Beyond", Foundations and Trends in Information Retrieval 3(4), doi:10.1561/1500000019](https://doi.org/10.1561/1500000019), §3.4): for each query word a document contains, a score that grows with the word's count (with diminishing returns), shrinks for documents longer than average and favours rare words. No model, no training.

Its strengths mirror an embedder's weaknesses. Keyword methods "are sensitive to highly selective keywords and phrases", while a **dense model** (an embedding model, as opposed to keyword search) "might lack sufficient capacity to represent salient phrases which appear rarely" ([Karpukhin et al., 2020](https://arxiv.org/abs/2004.04906), §5.3, Appendix C). BEIR, which tested retrievers on 18 datasets outside their training, found in 2021 that "BM25 is a robust baseline" while dense and sparse retrievers "often underperform other approaches" ([Thakur et al., 2021, "BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models", arXiv:2104.08663](https://arxiv.org/abs/2104.08663)). The average has since moved: E5 was "the first model that outperforms the strong BM25 baseline on the BEIR retrieval benchmark without using any labeled data" ([Wang et al., 2022](https://arxiv.org/abs/2212.03533), abstract). Where the exact words are the signal, BM25 still wins, as below.

Our benches show both sides ([release-day page](https://research.strata2signal.com/embeddinggemma-2-on-release-day/#keyword-search-and-the-hybrids); [shared-space page](https://research.strata2signal.com/embeddinggemma-2-shared-space/#what-the-code-shows)):

- On the 68 factual claims about one historical town, where a right answer is a passage containing a quoted source passage word for word, BM25 found one in the top 8 for 58, EmbeddingGemma 2 for 45, nomic-embed-text for 44. The test's written plan, committed before any score, had predicted it: answers copied word for word are keyword search's home ground.
- On the 316 code-search queries (a docstring as the query, its function as the document), BM25 came last: 87 right functions first, against nomic's 123 and EmbeddingGemma 2's 173 of 316 with its code-retrieval prefix, even though docstrings written beside their code share its identifiers, which favours keyword search.

The geometry has a ceiling too. Weller and colleagues showed that the number of different top-k result lists (the k best results for a query) that queries can draw from one collection is limited by the embedding's dimension; on their deliberately simple LIMIT dataset ("who likes apples?" over documents like "Jon likes apples"), strong embedders "struggle to reach even 20% recall@100" while BM25 "comes close to perfect", and on a 46-document version with synonyms swapped in, BM25 drops by nearly 90 % and falls below most embedders ([Weller et al., 2025, "On the Theoretical Limitations of Embedding-Based Retrieval", arXiv:2508.21038](https://arxiv.org/abs/2508.21038), §3, §5). Neither kind of search contains the other.

That is why RuleSage runs both kinds and merges their ranked lists by **reciprocal rank fusion** ([Cormack, Clarke and Büttcher, 2009, "Reciprocal Rank Fusion Outperforms Condorcet and Individual Rank Learning Methods", SIGIR 2009, doi:10.1145/1571941.1572114](https://doi.org/10.1145/1571941.1572114)): the scores are thrown away, and each list gives a result 1 / (60 + its rank) there, summed. A result first in one list only scores 1/61, about 0.0164; one third in both scores 2/63, about 0.0317, so agreement beats one list's enthusiasm ([Three librarians and a careful reader](https://research.strata2signal.com/three-librarians/) walks through it). On our 68 questions, fusion lifted EmbeddingGemma 2 to a right passage in the top 8 for 56, from 45, without passing BM25's 58; the best re-ranked list (bge-reranker-v2-m3, a re-ranker: a slower model that re-reads a shortlist, explained below, re-ordering BM25's top 50) reached 61, a gap over BM25 that chance explains easily (p = 0.55; exploratory; [release-day page](https://research.strata2signal.com/embeddinggemma-2-on-release-day/#keyword-search-and-the-hybrids)). Fusion is not free either: on the code set, where BM25 was the weak list, fusing it with EmbeddingGemma 2 put the right function first for 0.408 of queries, below EmbeddingGemma 2 alone at 0.547 (exploratory; [shared-space page](https://research.strata2signal.com/embeddinggemma-2-shared-space/#what-the-code-shows)).

## Evaluation, and the contamination problem {#evaluation-and-contamination}

Model-card scores mostly come from **MTEB**, the Massive Text Embedding Benchmark, whose launch paper found that "no particular text embedding method dominates across all tasks" ([Muennighoff et al., 2022, "MTEB: Massive Text Embedding Benchmark", arXiv:2210.07316](https://arxiv.org/abs/2210.07316)); EmbeddingGemma 2's card adds image, document, video and audio benchmarks beside it.

A public benchmark can be trained on, or on data like it. MTEB's maintainers note that "the highest ranking models achieve their scores by training on benchmark tasks, even though models with lower scores might generalize better" ([Chung et al., 2025, "Maintaining MTEB: Towards Long Term Usability and Reproducibility of Embedding Benchmarks", arXiv:2506.21182](https://arxiv.org/abs/2506.21182), §5.1), and describe the leaderboard's **zero-shot score**, one minus the share of the benchmark's datasets a model trained on; the first EmbeddingGemma's report excluded from its comparisons any model trained on more than 25 % of the MTEB data ([Schechter Vera et al., 2025](https://arxiv.org/abs/2509.20354), §4.2). Contamination can be measured, and is not always large: CLIP's authors checked 35 evaluation sets against their training data and found a median overlap of 2.2 % and at most 0.6 % of accuracy gained from it ([Radford et al., 2021](https://arxiv.org/abs/2103.00020), §5).

Our COCO lead carries the caveat: COCO has been public since 2014, and EmbeddingGemma 2's card reports scores on MMEB v2, a multimodal embedding benchmark that extends MMEB ([Meng et al., 2025, "VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents", arXiv:2507.04590](https://arxiv.org/abs/2507.04590)); the original MMEB lists COCO caption-to-image and image-to-caption retrieval, 100K and 113K training pairs, among its 20 training datasets ([Jiang et al., 2024, "VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks", arXiv:2410.05160](https://arxiv.org/abs/2410.05160), Table 1), and whether EmbeddingGemma 2 trained on it is not stated ([shared-space page](https://research.strata2signal.com/embeddinggemma-2-shared-space/#what-this-bench-does-not-say)). Our own 14 photographs were first published on this site from 2026-09-25 to 2026-10-05 ([shared-space page](https://research.strata2signal.com/embeddinggemma-2-shared-space/#what-this-bench-does-not-say)), after the pre-training data cutoff of January 2025 that the card states. On those, searching by the 40 descriptions published with them, EmbeddingGemma 2 put the right photograph first for 33 and the Nomic pair for 13 (exploratory: 12 scenes are too few to decide anything; [shared-space page](https://research.strata2signal.com/embeddinggemma-2-shared-space/#what-the-images-show)).

Our alternative is a **pre-registration**: questions from our own documents, frozen with the pass mark before any score. The test asked whether EmbeddingGemma 2 beat the model our search already runs, nomic-embed-text (44 of 68), by enough to pay for re-embedding everything: 10 percentage points of recall@8, so 51 of 68. It found 45. With 68 questions the test sees only large effects, so the verdict reads "no effect of at least 10 points was detected at n = 68", never "the two models are equal" ([release-day page](https://research.strata2signal.com/embeddinggemma-2-on-release-day/#the-instrument)).

## Single vector or late interaction {#single-vector-or-late-interaction}

**Late interaction** keeps a vector per token instead of one per input. ColBERT stores a small vector (128 numbers in the paper) for every document token; each query token finds its best-matching document token, and the score is the sum of those best matches, built from maximum-similarity (**MaxSim**) operators ([Khattab and Zaharia, 2020, "ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT", arXiv:2004.12832](https://arxiv.org/abs/2004.12832), §3.3, §4.1.2). By hand: query tokens (1, 0) and (0, 1) against document tokens (0.6, 0.8), (0.8, 0.6) and (1, 0). The first query token's best match scores 1.0, the second's 0.8; the score is 1.8. A second document whose three tokens all point (0.707, 0.707) scores only 1.41 by MaxSim; averaged into one vector each, though, it points exactly the query's way (cosine 1.0) and beats the first document (0.97). Late interaction prefers the document with strong evidence for each word, which one averaged vector cannot offer.

The price is space: a 200-token passage is 200 × 128 = 25,600 numbers against 768 for one vector; ColBERTv2's authors say late interaction "inflates the space footprint of these models by an order of magnitude" ([Santhanam et al., 2021, "ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction", arXiv:2112.01488](https://arxiv.org/abs/2112.01488)). ColBERT was competitive with models that read query and document together "while executing two orders-of-magnitude faster" (Khattab and Zaharia, abstract).

That reading-together model is the **cross-encoder** or **re-ranker** ([Nogueira and Cho, 2019, "Passage Re-ranking with BERT", arXiv:1901.04085](https://arxiv.org/abs/1901.04085)): too slow for a whole collection, so it re-orders a shortlist. On our 68 questions, a small cross-encoder (ms-marco-MiniLM-L6-v2) re-ordering BM25's top 50 put a right passage first for 0.500 of questions, against BM25's 0.471 and EmbeddingGemma 2's 0.338 (the reference path; exploratory; [release-day page](https://research.strata2signal.com/embeddinggemma-2-on-release-day/#keyword-search-and-the-hybrids)).

## Why vectors from different models never mix {#why-vectors-never-mix}

Each model invents its own coordinates: even two random initialisations of one network land in different cones ([Liang et al., 2022](https://arxiv.org/abs/2203.02053), §2.3), and cosine is unchanged when a space is rotated, so two models can learn one arrangement in unrelated coordinates. By hand: a second model learns the toy's exact arrangement with its axes in another order, *cars, animals, home*, so A becomes (0.0, 0.6, 0.8). Score the first model's question, (0.8, 0.6, 0.0), against the second model's documents: A 0.36, B 0, C ("brake pads") 0.8, and D ("a dog's first car journey") a perfect 1.0. Nonsense, and no error.

No error is the danger. EmbeddingGemma 2 and nomic-embed-text both write 768 numbers, so a store holding some of each "would pass every dimension check and simply return bad neighbours"; "a `vector(768)` type is not a model check" ([release-day page](https://research.strata2signal.com/embeddinggemma-2-on-release-day/#how-to-run-it)). One model through two runtimes can differ too: llama.cpp's server matched Google's code on text at cosine 0.9997 or better (40 texts), but its image vectors fell as low as 0.978 on our 14 photographs, 13 of them below the 0.999 agreement we require of a second runtime ([shared-space page](https://research.strata2signal.com/embeddinggemma-2-shared-space/#what-the-images-show)).

Translating between spaces is a learning problem of its own; one 2025 paper does it without paired data and warns that an attacker holding only vectors can then "extract sensitive information about the underlying documents" ([Jha et al., 2025, "Harnessing the Universal Geometry of Embeddings", arXiv:2505.12540](https://arxiv.org/abs/2505.12540)), and a 2023 paper recovered 92 % of 32-token texts exactly from their vectors ([Morris, Kuleshov, Shmatikov and Rush, 2023, "Text Embeddings Reveal (Almost) As Much As Text", arXiv:2310.06816](https://arxiv.org/abs/2310.06816)), two reasons to guard stored vectors like the text itself. The rule: one model, one prefix scheme, one size and one runtime per index; record them beside every vector; re-embed everything when you switch.

## Cost and hardware {#cost-and-hardware}

An embedder costs compute once per item at indexing and once per query at search; the first is the big bill, due again at every model switch. Our figures come from one laptop CPU (an Intel Core Ultra 9 290HX Plus, 8 threads; [shared-space page](https://research.strata2signal.com/embeddinggemma-2-shared-space/#what-we-tested)) with other jobs running, and no GPU was used, so they are rough costs, not a fair race between the two models ([shared-space page](https://research.strata2signal.com/embeddinggemma-2-shared-space/#speed)):

- **Text:** EmbeddingGemma 2 through Google's code, 9.12 CPU-seconds per 1,000 tokens (CPU time summed over threads), or 2.53 CPU-hours per million; nomic-embed-text loaded by RuleSage's own code, 5.71 CPU-seconds per 1,000 tokens. Both on short captions, where each call's fixed cost dominates.
- **Images:** at the default 280 tokens, a median 14.80 CPU-seconds per image (1.97 s of wall clock), or 41.1 CPU-hours per 10,000; at 70 tokens, 3.145 CPU-seconds; at 1,120 tokens, 97.11 CPU-seconds.
- **Audio:** 3.560 CPU-seconds per 10-second window, with the text, image and audio parts all loaded.

More image tokens are not automatically better. The card expects them to help ("increasing input sequence length will improve fine-grained visual understanding"); on our material, 1,120 cost 6.6 times the CPU of 280 and did worse on seven interface screenshots from our pages, finding the right one from its alt text for 2 against 5; 70 cost a fifth and put the right COCO photograph first for 30 fewer captions in 1,000 (exploratory; [shared-space page](https://research.strata2signal.com/embeddinggemma-2-shared-space/#what-the-images-show)). Load only the parts you need: the card's text-only setup is 270 M of 740 M parameters.

Searching is cheap by comparison: exact search costs 768 multiply-adds per stored vector, which is what our benches did over a few thousand items; past that, databases build an **approximate nearest-neighbour** index such as HNSW, a layered graph whose search cost its authors report scales logarithmically ([Malkov and Yashunin, 2016, "Efficient and robust approximate nearest neighbor search using Hierarchical Navigable Small World graphs", arXiv:1603.09320](https://arxiv.org/abs/1603.09320)), which adds memory on top of the storage table above. The bill that decides most switches is the re-embed: the one large store of ours we counted, a conversation archive (our searchable record of past working sessions), held 701,937 message vectors on 2026-10-06 (UTC); its token count was not read, so no total is printed, but at the rate above every 100 million text tokens is about 250 CPU-hours ([shared-space page](https://research.strata2signal.com/embeddinggemma-2-shared-space/#what-moving-would-take)).

## How to choose an embedder: a checklist {#how-to-choose}

1. **Write your test before reading the leaderboard**: real queries from your content, their right answers marked, frozen before any score, with the gain worth a re-embed decided in advance.
2. **Check what the leaderboard numbers were trained on**: the zero-shot score, the training-data disclosures.
3. **Run BM25 on the same test**, and if your queries carry names, codes or numbers, a fusion too.
4. **Load only the modalities you need**; images and text in one index means filtering or ranking per modality.
5. **Set any similarity cutoff on your own test**: no fixed cosine means "relevant" across models or modalities (on average, 0.546 was an unrelated passage for EmbeddingGemma 2 above, and 0.066 a caption's own photograph for the Nomic pair).
6. **Use the documented prefixes exactly, and check them** at the end of each embedding job.
7. **Choose a size and a storage precision by measuring them** on your test; re-normalise after cutting; one size for queries and documents.
8. **Respect the compute precision**: for EmbeddingGemma 2, bfloat16 or float32, never float16.
9. **Read the context limit (the longest input, in tokens, the model is meant to read) from the card, and set it on every runtime yourself**: EmbeddingGemma 2's card says 8,192 tokens, while its configuration file, the llama.cpp file's metadata and Ollama's listing all say 262,144 (yes, the same number as its vocabulary) ([release-day page](https://research.strata2signal.com/embeddinggemma-2-on-release-day/#how-to-run-it)).
10. **One model per index, with its provenance beside every vector**: model, version, prefix, size, runtime and file checksum.
11. **Price the switch**: CPU-seconds per item, times your corpus, times every store holding the old model's vectors.
12. **Read the licence first-hand**: EmbeddingGemma 2 is tagged Apache 2.0 and its card adds a use policy; the first EmbeddingGemma is under the Gemma licence ([release-day page](https://research.strata2signal.com/embeddinggemma-2-on-release-day/#what-google-released)).

## What to take with you {#what-to-take-with-you}

- **An embedder is a transformer, a projection and a pooling step, trained on pairs**; the pairs decide what "belongs together" means, and the projection sets the width.
- **A shared space keeps a gap**: on 1,000 COCO captions, EmbeddingGemma 2 put the right photograph first for 721 beside unrelated text and 139 beside related text; the Nomic pair (transformers 4.57.6), for none. One model is not one ranked list.
- **Matryoshka vectors cut cleanly if re-normalised**: a right passage in the top 8 for 45 of our 68 questions at 768 and 512, 44 at 256, 35 at 128 (32-bit reference vectors; the cuts exploratory).
- **Keyword search wins where answers repeat the query's words**: BM25 put a right passage in the top 8 for 58 of 68 against EmbeddingGemma 2's 45 on our verbatim set, yet came last on our code set (the right function first for 87 of 316, against 173 with EmbeddingGemma 2's code-retrieval prefix).
- **Same width is not same space**: two 768-number models mixed in one index fail silently, and one model's image vectors through two runtimes agreed only to cosine 0.978 at worst on our 14 photographs, 13 of them below our 0.999 mark.

## How to check our work — and see it live {#how-to-check-our-work-and-see-it-live}

Every bench figure here is on the two pages linked throughout, with its conditions, its gates and the audits that re-derived it; the toy examples need only a calculator. And the live door: [ask RuleSage a rules question](https://rulesage-live.strata2signal.com/), free and without an account, and open the wrench on the ruling (the panel behind every answer, which shows the search's own timing and the pages it cited).

## The rest of the seminar {#the-rest-of-the-seminar}

This page's companions:

- [EmbeddingGemma 2, measured on release day against the embedder we already run](https://research.strata2signal.com/embeddinggemma-2-on-release-day/) — the text bench, the runtimes and the traps.
- [EmbeddingGemma 2's shared space for text, images and audio, measured on what we publish](https://research.strata2signal.com/embeddinggemma-2-shared-space/) — images, audio, code, the gap and the costs.
- [Three librarians and a careful reader](https://research.strata2signal.com/three-librarians/) — keyword search, embeddings, fusion and a re-ranker working together in RuleSage.
- [How a vision model sees](https://research.strata2signal.com/how-a-vision-model-sees/) — what an image becomes on its way into a language model.

[The whole shelf](https://research.strata2signal.com/) holds the rest: the benches behind the claims we publish, failures included. If there is a measurement you want next, [say so](https://strata2signal.com/contact/); the suggestion box is read.

## Who ran this, and thanks {#who-ran-this-and-thanks}

This page stands on the authors of every paper and model card linked where it is cited, and on Google DeepMind and Nomic AI for open models whose cards say what they are; the cards and Google's announcement were read on 2026-10-09 (UTC). The licences this shelf has read first-hand are on its [licences](https://research.strata2signal.com/licences/) page. A small human team asked for this explainer and chose what it would cover; a fleet of AI agents read the sources, computed the examples and drafted it under that team's rulings.

<!-- derived 2026-10-10 (UTC) by tools/derive_md.py from the pour source.
     source html sha256: fefb702ac09b3d0d4766b53298633ff1bf5a045c32cb917d856e6b76c40ee88f
     derivation sha256:  d5357b819fc5bbc51d632dbde8974d3b1b9083f3b3354d827abe3d2ba3894abc
     the {#id} on each heading is the anchor that heading carries on the page. -->
