EmbeddingGemma 2 Guide: Local Multimodal Embeddings for Code

TL;DR
EmbeddingGemma 2 is a 740M open model that embeds code, images, video and audio in one space. Specs, Matryoshka storage math, and where it breaks.
EmbeddingGemma 2 is Google's new open embedding model: 740M parameters, built on Gemma 4, licensed Apache 2.0, and it maps text, code, images, video and audio into one 768-dimensional vector space that you can truncate to 512, 256 or 128 dimensions. Google announced it on October 6, 2026 with weights on Hugging Face and Kaggle. For builders, the headline is not the audio or video. It is a code retrieval score that jumped almost 10 points over the first EmbeddingGemma, in a 270M text-only configuration that runs on a laptop.
Last updated: October 9, 2026
This guide covers what shipped, how it compares to EmbeddingGemma 1 and the 7 MB Ternlight browser encoder, the storage math at each Matryoshka cut point measured against a real repository, and the places it breaks. The frame throughout is local RAG, codebase retrieval and agent memory on your own machine.
What shipped#
Every number in this section comes from Google's launch post, the EmbeddingGemma 2 model card or the developer guide. These are vendor-reported results on the full-precision checkpoint. We have not reproduced them.
- Modular size. 740M parameters in total: a 270M text backbone (130M transformer plus 140M embedder), a 170M vision encoder and a 300M audio encoder. You load only what you need: 270M for text and code, 440M with vision, 570M with audio, 740M for everything.
- One shared space. All four configurations load from the same checkpoint and project into the same 768-dimensional space, so a query embedded with the text-only setup can be scored against documents embedded with the full model, per the developer guide.
- Matryoshka truncation. Output vectors can be cut to 512, 256 or 128 dimensions, up to a 6x storage reduction.
- 8,192-token context, four times EmbeddingGemma 1, shared across modalities. An image costs 280 tokens by default (about 29 per input), a video frame 140 tokens (about 58 frames at 1 frame per second), and audio 25 tokens per second (about 327 seconds of 16 kHz mono).
- On-device footprint. With quantization on a Pixel 11 Pro, Google reports about 191MB of active RAM for the text-only weights and about 567MB for the full multimodal model.
- Task prefixes. Text inputs take short instruction prefixes. For code, queries use
task: code retrieval | query: {query}and documents usetitle: {filename} | text: {code}. Images, video and audio take no prefix. - Training data cutoff: January 2025, per the model card.
The headline benchmarks from the model card, at 768 dimensions:
| Benchmark | EmbeddingGemma 2 | EmbeddingGemma 1 |
|---|---|---|
| MTEB multilingual v2 (mean task) | 61.36 | 61.15 |
| MTEB code v1 (NDCG@10) | 78.68 | 68.76 |
| MIEB lite (image) | 64.64 | - |
| MMEB v2 VisDoc (NDCG@5) | 67.84 | - |
| MMEB v2 video (Hit@1) | 50.67 | - |
| MSEB retrieval (audio, MRR@10) | 69.54 | - |
Read the first two rows together. Multilingual text is flat (61.15 to 61.36). Code is up 9.92 points, about 14% relative. If your current index is prose, upgrading buys you little on text alone. If your index is a codebase, this is the release that matters.
Google's own demo in the developer guide is exactly that use: it embedded the Hugging Face transformers codebase with the 270M text-only setup and had a Gemma 4 26B A4B agent query it for code. That pairing is the one to copy if you already run Gemma 4 locally, since Google says the two share a text tokenizer and audio encoder, which lowers the combined memory footprint.
EmbeddingGemma 2 vs EmbeddingGemma 1 vs Ternlight#
| EmbeddingGemma 2 | EmbeddingGemma 1 | Ternlight base | |
|---|---|---|---|
| Released | Oct 6, 2026 | Sept 4, 2025 | Shown on HN, July 2026 |
| Size | 740M total, 270M text-only | 308M | 7 MB package |
| Base | Gemma 4 | Gemma 3 | Distilled from MiniLM, ternary weights |
| Inputs | Text, code, image, video, audio | Text | Text |
| Output dims | 768, MRL to 512/256/128 | 768, MRL to 512/256/128 | 384 |
| Context | 8,192 tokens | 2K tokens | Not stated |
| MTEB code v1 | 78.68 | 68.76 | Not reported |
| Footprint claim | ~191MB text-only, ~567MB full (quantized, Pixel 11 Pro) | Under 200MB RAM (quantized) | ~5ms per embedding in the browser |
| Runs via | sentence-transformers, transformers.js, Ollama, llama.cpp, MLX, vLLM, LiteRT | sentence-transformers, Ollama, llama.cpp, MLX, LiteRT | npm package, WASM |
EmbeddingGemma 1 figures are from Google's original launch post. Ternlight figures are from our Ternlight write-up and its repository.
The decision is mostly about the job, not the leaderboard:
- Pick EmbeddingGemma 2 (270M text-only) when you index code for an agent or a semantic code finder. The code gain is the reason to switch, and the text-only load keeps the footprint close to EmbeddingGemma 1.
- Pick EmbeddingGemma 2 (440M, text plus vision) when screenshots, diagrams, slides or PDF pages belong in the same index as code and notes. This is the configuration for agent memory that has to remember what a UI looked like.
- Stay on EmbeddingGemma 1 if your index is multilingual prose, it works, and re-embedding is expensive. The text numbers barely moved, and nothing in Google's material says EmbeddingGemma 1 vectors are compatible with the new space, so switching means a full re-embed.
- Pick Ternlight when the hard constraint is a browser tab and a single-digit-megabyte download, and short text is all you embed.
For where this sits next to the generative models you might pair it with, our local models roundup keeps the current picks.
The Monday run: size the index before you download anything#
The useful first step is not downloading 1.3GB of weights. It is finding out how many vectors your repo turns into, how much storage each Matryoshka cut costs, and which files will not fit in one 8,192-token input. This stdlib-only script does that. Token counts are a rough estimate (bytes divided by four), not the real tokenizer.
# eg2_plan.py - size an EmbeddingGemma 2 index for a repo before downloading anything.
# Stdlib only. Token counts are estimates (bytes / 4), not the real tokenizer.
import math, os, random, subprocess, sys
repo = sys.argv[1] if len(sys.argv) > 1 else "."
CODE = (".ts", ".tsx", ".js", ".mjs", ".py", ".go", ".rs")
IMG = (".png", ".jpg", ".jpeg", ".webp")
CHUNK_TOKENS = 1024 # one vector per ~1K-token chunk
CTX = 8192 # model card: shared 8,192-token window
IMG_TOKENS = 280 # model card: default tokens per image
files = subprocess.run(["git", "-C", repo, "ls-files"], capture_output=True, text=True).stdout.split()
chunks = over_ctx = images = 0
for f in files:
p = os.path.join(repo, f)
if f.endswith(CODE):
try:
toks = os.path.getsize(p) / 4
except OSError:
continue
chunks += max(1, math.ceil(toks / CHUNK_TOKENS))
over_ctx += toks > CTX
elif f.endswith(IMG):
images += 1
vectors = chunks + images
print(f"code chunks: {chunks} files over 8,192 est. tokens: {over_ctx} images: {images}")
print(f"images per single input at {IMG_TOKENS} tokens each: {CTX // IMG_TOKENS}")
print(f"total vectors: {vectors}")
print("dim float32_MB bf16_MB")
for d in (768, 512, 256, 128):
print(f"{d:<5} {vectors*d*4/1e6:>10.1f} {vectors*d*2/1e6:>8.1f}")
random.seed(0)
v = [random.gauss(0, 1) for _ in range(768)]
n = math.sqrt(sum(x * x for x in v))
v = [x / n for x in v]
for d in (512, 256, 128):
t = v[:d]
print(f"norm after slicing to {d}: {math.sqrt(sum(x * x for x in t)):.3f} (re-normalize to 1.000)")
We ran it with Python 3.11 against this site's own repository on October 9, 2026. This is the verbatim output:
code chunks: 11492 files over 8,192 est. tokens: 122 images: 2251
images per single input at 280 tokens each: 29
total vectors: 13743
dim float32_MB bf16_MB
768 42.2 21.1
512 28.1 14.1
256 14.1 7.0
128 7.0 3.5
norm after slicing to 512: 0.848 (re-normalize to 1.000)
norm after slicing to 256: 0.611 (re-normalize to 1.000)
norm after slicing to 128: 0.430 (re-normalize to 1.000)
Three things fall out of that before any model runs:
- Storage is not the constraint at repo scale. A whole production Next.js codebase plus every image in it is 42MB of float32 vectors at full width. The 6x Matryoshka saving matters at millions of vectors (Google's guide puts a million 768-dimension bf16 vectors at roughly 1.5GB versus 250MB at 128), not at fourteen thousand. At repo scale, pick the dimension for quality, not bytes.
- 122 files are too long for one input. Anything past 8,192 tokens has to be chunked, and long files are usually the ones an agent most needs to find inside. Chunk by function or by around 1K tokens and keep the file path in the
title:field. - Truncation without re-normalization is silently wrong. The model card warns that a sliced vector is no longer unit length and cosine scores degrade without an error. The norms above come from a random vector, where energy is spread evenly; a Matryoshka-trained vector front-loads more of its signal, so real norms will sit higher, but still below 1. Pass
normalize_embeddings=Trueand the library handles it.
The indexing script: code plus screenshots in one space#
Here is the shape of the real run, built from the calls documented in the model card and the developer guide (sentence-transformers v6.1.0 or later). Be clear on its status: our test machine could not reach Hugging Face or the Ollama registry when we wrote this, so we have not executed this script or measured its retrieval quality. Every call in it appears in Google's documentation, and the sizing run above is the part we verified.
import pathlib
import numpy as np
import torch
from sentence_transformers import SentenceTransformer
# float16 is unsafe for this model (NaN or degraded vectors); use bf16 or float32.
dtype = torch.bfloat16 if torch.cuda.is_available() and torch.cuda.is_bf16_supported() else torch.float32
model = SentenceTransformer(
"google/embeddinggemma-2",
model_kwargs={"torch_dtype": dtype},
config_kwargs={"audio_config": None}, # text + vision = 440M, audio encoder never loaded
truncate_dim=256, # one dimension for every call
)
labels, docs = [], []
for p in pathlib.Path("lib").rglob("*.ts"):
text = p.read_text(errors="ignore")
for i in range(0, len(text), 4000): # ~1K-token chunks
labels.append(f"{p}:{i}")
docs.append(f"title: {p} | text: {text[i:i + 4000]}")
vecs = list(model.encode(docs, batch_size=16, normalize_embeddings=True))
for p in pathlib.Path("public/images").rglob("*.png"):
labels.append(str(p))
vecs.append(model.encode({"image": str(p)}, normalize_embeddings=True)) # media: no prefix
query = model.encode("where is the pricing card rendered?", prompt_name="CodeRetrieval", normalize_embeddings=True)
scores = model.similarity(query, np.stack(vecs))[0]
for i in scores.argsort(descending=True)[:5]:
print(f"{scores[i]:.3f} {labels[i]}")
Why 256: the model card's truncation table shows code retrieval losing 2.5 points at 256 (78.68 to 76.18) but 7.27 points at 128 (71.41). The multimodal side falls harder. MMEB v2 overall goes from 59.01 at full width to 56.24 at 256 and 45.65 at 128. Google's own guidance is that 128 suits text-only indexes and first-stage shortlisting, and that multimodal use should be validated before going that low. For a mixed code-and-screenshot index, 256 is the floor.
Where the vectors live is a separate decision. For a repo-sized index, SQLite with a vector extension or an embedded store is plenty. Our vector database comparison covers when you outgrow that. If your agent already navigates a local code graph, embeddings are the fuzzy layer on top: the graph answers "who calls this", the vectors answer "where is the thing that does roughly this".
One more design note from the developer guide: if you start text-only and add images later, reload the model with the vision encoder enabled. Embeddings you already computed do not need to be recomputed, because every configuration shares the same space.
Where it breaks#
- Context is shared. 8,192 tokens covers one long file, or about 29 images, or five and a half minutes of audio, not all of them at once. Interleaved inputs draw down the same budget.
- float16 fails quietly. The model card says activations exceed float16 range, producing NaN or degraded embeddings without raising an error. Use bfloat16 on hardware that supports it and float32 on most CPUs.
- Runtime support is uneven across modalities. The ggml-org GGUF for llama.cpp is the text model only (Q8_0 at 310MB, BF16 at 558MB). The Ollama library page lists
270m(378MB, text),440m(714MB, text and image),570m(990MB, listed as text) and740m(1.3GB, text and image), so audio is not exposed as an input there yet. The same page shows a 256K context field; trust the model card's 8,192. - Truncation saves index bytes, not RAM. As one commenter on the HN thread pointed out, this is Matryoshka on the output vectors, not a nested architecture, so a 128-dimension index still runs the full encoder.
- Text-only users get little. Multilingual MTEB moved by 0.21 points. The gain is in code and in the new modalities.
- Newer APIs are unknown territory. With a January 2025 data cutoff, the model has never seen libraries and framework APIs released since then. Embeddings still work on unfamiliar code, but expect weaker matches on brand-new names, and keep a keyword fallback.
What people are actually saying#
The Hacker News thread passed 400 points within a day. The main points:
- The license is the story for some. Simon Willison argued that embeddings are the worst place for a closed, hosted-only model: you store millions of vectors, and if the vendor retires the model you pay to re-embed all of them. Apache 2.0 weights mean you can always run the exact model yourself, even if you would rather pay someone to host it.
- Practitioner numbers on a laptop. minimaxir reported throughput from his own local embedding tool on an M3 Pro: about 78 embeddings per second on short texts, 4 per second for images, 6 per second for 30-second audio chunks, and 0.2 per second per minute of video. Another commenter said it replaced both EmbeddingGemma 1 and an image-text CLIP model on a site they run, with lower RAM use.
- The sharpest pushback is on text. One early tester said it was neither better nor faster for text-only retrieval and clustering, which matches the flat multilingual score. Another reported the MediaPipe decision example misclassifying its own sample sentence, scoring a refund request as non-financial.
- Open questions. Several asked why Google did not compare it against SigLIP 2 for images or against commercial text embedders, and one asked whether binary quantization would beat Matryoshka truncation for storage. Neither got a definitive answer in the thread.
The fair read: the community is excited about local multimodal retrieval and the license, and skeptical that text-only users need to move. That is the same split the benchmarks show.
FAQ#
What is EmbeddingGemma 2?#
An open embedding model from Google DeepMind, released October 6, 2026 under Apache 2.0. It has 740M parameters, is built on Gemma 4, and maps text, code, images, video and audio into a single 768-dimensional space, with Matryoshka truncation to 512, 256 or 128 dimensions.
Is EmbeddingGemma 2 better than EmbeddingGemma 1?#
For code, yes: MTEB code rose from 68.76 to 78.68 in Google's model card. For multilingual text, it is essentially the same (61.15 to 61.36). It also adds images, video and audio, and quadruples context to 8,192 tokens. If you only embed prose and your current index works, there is little reason to re-embed.
How much memory does EmbeddingGemma 2 need?#
Google reports about 191MB of active RAM for the quantized text-only weights and about 567MB for the full multimodal model on a Pixel 11 Pro. On Ollama the tags range from 378MB (270m) to 1.3GB (740m). Load only the encoders you need to stay at the low end.
Which embedding dimension should I use?#
768 or 512 when recall matters most or the index is multimodal. 256 is the practical default for mixed code and image indexes, losing 2.5 points on code retrieval for a 3x storage cut. Keep 128 for large text-only indexes or a first-pass shortlist before re-ranking. Queries and documents must use the same dimension.
Can I run EmbeddingGemma 2 with Ollama or llama.cpp?#
Yes for text. ollama pull embeddinggemma-2 fetches the 740m tag by default, with 270m, 440m and 570m tags also listed. The ggml-org GGUF for llama.cpp covers the text model. For images, video and audio, use sentence-transformers or transformers, where the model card documents the multimodal input format.
Is EmbeddingGemma 2 good for agent memory?#
It fits the job well when memory includes more than text. One model can put a code chunk, a screenshot of the UI it renders and a note about it in the same space, so a single query can pull all three. Treat it as one layer of agent memory, alongside structured state and a keyword index.
Sources#
- EmbeddingGemma 2: an open, lightweight multimodal embedding model - Google launch post, October 6, 2026 (also linked from Google DeepMind's blog)
- EmbeddingGemma 2 model card - architecture, benchmarks, truncation table, best practices
- google/embeddinggemma-2 on Hugging Face - weights
- EmbeddingGemma 2 on Kaggle - weights
- EmbeddingGemma 2: The Developer Guide - encoder configurations, dimension guidance, codebase demo
- Introducing EmbeddingGemma - EmbeddingGemma 1 specs, September 4, 2025
- embeddinggemma-2 on Ollama - tags and sizes
- ggml-org/embeddinggemma-2-GGUF - llama.cpp build
- Ternlight repository - browser embedding model used in the comparison
- Hacker News discussion - community reaction, read October 9, 2026
Continue Reading#
- Gemma 4: The Open Model Guide for Developers - the generative model EmbeddingGemma 2 is built on and pairs with
- Ternlight: A 7 MB Embedding Model That Runs Entirely in the Browser - the tiny text-only alternative for browser tabs
- Vector Database Comparison for RAG and AI Agents - where to store the vectors once you have them
- Local Code Graphs Are the Agent Context Layer - the structural index that embeddings complement
- Ollama vs LM Studio vs vLLM vs llama.cpp - choosing the local runtime that serves both models
Get the next deep dive like this in your inbox
One email a week on Google and the rest of the AI dev stack. Free.
Read next on local and open-weight models
Gemma 4: The Open Model Guide for Developers
Gemma 4 ships byte-for-byte open weights from Google DeepMind. How developers deploy it locally, fine-tune it, and ship agents on top of it.
11 min readTernlight: A 7 MB Embedding Model That Runs Entirely in the Browser
Ternlight ships a ternary-quantized sentence encoder at 7 MB that runs semantic search at 5ms per embedding - entirely client-side via WASM, no API calls required. Here is how it works, what HN thinks, and where browser-side embeddings make sense.
6 min readVector Database Comparison for RAG and AI Agents
pgvector, Pinecone, Qdrant, Weaviate, Chroma, Milvus, and Turbopuffer compared on hosting model, filtering, scale, and cost for RAG.
7 min readNew here? Start with
Technical content at the intersection of AI and development. Building with AI agents, Claude Code, and modern dev tools - then showing you exactly how it works.







