EmbeddingGemma 2: One Small Open Model for Searching Text, Code, Images, Video and Audio

Google's new open model puts text, code, images, video and audio into one shared vector space, and it runs in about 567 MB of RAM. Here's how it works and how to try it.

Anshu Pathak
•
8 min read
EmbeddingGemma 2: One Small Open Model for Searching Text, Code, Images, Video and Audio

Most search setups have a hidden mess behind them.

Text goes through one model. Photos go through another. Audio gets turned into a transcript first, then goes through the text model. Then somebody writes glue code to merge all the results.

Google's new EmbeddingGemma 2 is built to remove that mess. One open model, one shared vector space, five kinds of content: text, code, images, video and audio.

It launched on October 6, 2026, with open weights under the Apache 2.0 license. In this post I'll cover what it is, what changed from version 1, what the numbers say, and how to run it yourself.

💡 Quick note: The benchmark and memory numbers in this post come from Google's launch post, developer guide and docs (linked at the end). They're vendor numbers, so test on your own data before you rely on them.


EmbeddingGemma 2 at a glance

Released

October 6, 2026

Made by

Google DeepMind

Built on

Gemma 4

Size

740M parameters for everything, 270M for text and code only

Inputs

Text, code, images, video, audio

Vector size

768 dimensions (can shrink to 512, 256 or 128)

Context window

8,192 tokens

License

Apache 2.0

RAM on a Pixel 11 Pro (quantized)

About 191 MB text-only, about 567 MB full model


A 60-second refresher on embeddings

An embedding model reads some content and turns it into a list of numbers. That list is called a vector.

Content with a similar meaning ends up with similar vectors. So "ocean waves at sunset" lands close to a beach photo taken at dusk, and far away from a spreadsheet.

Search, RAG, recommendations and clustering all lean on this. You embed your data once, store the vectors, and compare each new query's vector against them.

Here's the catch with older setups. A text model can only compare text with text. To search photos using words, you needed a second model, and the two vector spaces didn't match.

EmbeddingGemma 2 puts everything in the same space. So a text query can be compared directly with a photo, a video clip or a sound recording.


What's new in version 2

The first EmbeddingGemma was text-only. Google says it passed 20 million downloads, so a sequel made sense.

Version 2 changes four big things:

  • More input types. Text, code, images, video and audio. The vision encoder also handles visual documents like PDFs, slides and charts.

  • Better at code. The MTEB Code score goes from 68.76 to 78.68. That's a gain of 9.92 points, or about 14%.

  • Longer context. 8K tokens, which is 4x what version 1 handled.

  • Modular design. You only load the parts you need.

It's also built on Gemma 4, so it shares Gemma 4's text tokenizer and audio encoder design. That detail matters later, when we get to on-device RAG.


How it's built

Each type of input gets its own encoder. They all end up in the same 768-dimension space.

text / code ───► text model (270M) ─────────┐
                                              │
 images / video ► vision encoder (+170M) ────┼──► shared backbone ──► 768-dim vector
                                              │
 audio ─────────► audio encoder (+300M) ─────┘

Because the parts are separate, you can leave out what you don't need:

Setup

Parameters

What it handles

Text + code

270M

Text and code only

Text + vision

440M

Adds images, video, visual documents

Text + audio

570M

Adds audio

Everything

740M

All five types

📌 Good to know: All four setups load from the same checkpoint and share one vector space. A query embedded with the 270M text-only setup can be compared directly with documents embedded by the full model.

Google's guide also says that if you start with a text-only index and add images later, you don't need to re-embed what you already have.

That opens up a neat setup. A small text-only model can handle search queries on a phone, while a bigger machine does the heavy media indexing.


What the numbers say

Here's what Google reports:

  • Code search: MTEB Code goes from 68.76 to 78.68.

  • Quality for its size: Google says it leads sub-1B multimodal embedders on benchmarks like MTEB Code and MAEB (the audio one), and beats some specialist models more than twice its size.

  • Text quality: Multilingual text performance matches version 1.

  • Memory: On a Pixel 11 Pro, with quantization, it needs about 191 MB of active RAM for text-only weights and about 567 MB for the full multimodal model.

  • Input budget: The 8K window fits up to 5.5 minutes of audio, 29 images or 58 video frames, or a mix of them.

⚠️ Heads up: These are Google's own results. The full tables are in the model card. Run a small test on your own data before you commit.


Shrinking vectors with Matryoshka

Vectors cost storage. A million of them adds up fast.

EmbeddingGemma 2 is trained with Matryoshka Representation Learning (MRL). In plain words, the most important information sits at the start of the vector. So you can chop it down to 512, 256 or 128 numbers and keep much of the quality.

Here's what that looks like in bfloat16, per million vectors:

Dimensions

Storage

Quality notes (from Google's guide)

768

about 1.5 GB

Full quality

512

about 1 GB

Recommended along with 768 for multimodal and visual document search

256

about 0.5 GB

Keeps most of the quality on text and code, and about 95% on image, video and speech

128

about 0.25 GB

Around 90% on text and code, but only around 75% on image, video and speech

The 768 and 128 storage numbers (about 1.5 GB and 250 MB) come from Google's guide. The 512 and 256 rows are the same math: vectors × dimensions × 2 bytes.

💡 Tip: Use 128 for big text-only indexes and for first-pass shortlisting before a re-ranker. Google says to check your own data before using 128 for multimodal queries.

Two rules to remember:

  1. Queries and documents must use the same dimension.

  2. Pass normalize_embeddings=True when you truncate.


Try it yourself

Google's guide uses the sentence-transformers library. You need version 6.1.0 or newer.

pip install -U "sentence-transformers[image,audio,video]" transformers

Step 1: Load the model

from sentence_transformers import SentenceTransformer

MODEL_ID = "google/embeddinggemma-2"

# Full model: all modalities (740M parameters)
model = SentenceTransformer(MODEL_ID)

Want a lighter load? Turn off the encoders you don't need. They never get loaded into memory, so you save on both weights and peak usage.

# Text only (270M parameters)
text_only_model = SentenceTransformer(
    MODEL_ID,
    config_kwargs={"vision_config": None, "audio_config": None},
)

# Text, images and video (440M parameters)
text_image_model = SentenceTransformer(
    MODEL_ID,
    config_kwargs={"audio_config": None},
)

# Text and audio (570M parameters)
text_audio_model = SentenceTransformer(
    MODEL_ID,
    config_kwargs={"vision_config": None},
)

Step 2: Embed text and code

The model was trained with short task instructions. You switch them on with prompt_name. For search, queries and documents use different ones:

query = "What causes the northern lights?"
document = "The northern lights are caused by charged particles from the sun.."

query_emb = model.encode(query, prompt_name="SearchQuery")
doc_emb = model.encode(document, prompt_name="Document")

print(model.similarity(query_emb, doc_emb))

Step 3: Search across media

For images, video and audio, pass a dictionary keyed by type. No prompt needed.

# One text query against a photo and a sound recording
image_emb = model.encode({"image": "sunset_beach.jpg"})
audio_emb = model.encode({"audio": "ocean_waves.wav"})
query_emb = model.encode("ocean waves at sunset", prompt_name="SearchQuery")

print(model.similarity(query_emb, image_emb))
print(model.similarity(query_emb, audio_emb))

You can even mix types into one embedding. Mark where each item goes with <|image|>, <|video|> or <|audio|>:

# One embedding for a product listing with text, a photo and a video
listing_emb = model.encode({
    "text": "Waterproof trail shoe. <|image|> Grip test on wet rock: <|video|>",
    "image": "trail_shoe.jpg",
    "video": "grip_test.mp4",
})
query_emb = model.encode("waterproof trail shoes", prompt_name="SearchQuery")

print(model.similarity(query_emb, listing_emb))

Step 4: Shrink the vectors

query_emb = model.encode(
    query,
    prompt_name="SearchQuery",
    truncate_dim=256,
    normalize_embeddings=True,
)

To use one size for every call, set it at load time: SentenceTransformer(MODEL_ID, truncate_dim=256).


Here's a small script that puts those pieces together. It indexes the photos and sound clips in a folder, then searches them with plain text.

from pathlib import Path
from sentence_transformers import SentenceTransformer

model = SentenceTransformer("google/embeddinggemma-2")

# 1. Index: embed every photo and sound clip once
index = []
for path in Path("media").iterdir():
    suffix = path.suffix.lower()
    if suffix in {".jpg", ".jpeg", ".png"}:
        emb = model.encode({"image": str(path)})
    elif suffix == ".wav":
        emb = model.encode({"audio": str(path)})
    else:
        continue
    index.append((path.name, emb))

# 2. Search with text
def search(text, top_k=3):
    q = model.encode(text, prompt_name="SearchQuery")
    scored = [(float(model.similarity(q, emb)), name) for name, emb in index]
    return sorted(scored, reverse=True)[:top_k]

print(search("ocean waves at sunset"))

For a real app, save the vectors in a vector database instead of a Python list. Google points to Qdrant for that.

📌 File notes: Video is sampled at 1 frame per second by default. Audio should be 16 kHz mono.


Where I'd use it

1. Local code search for coding agents. The MTEB Code jump is the headline here. Google's demo embedded the whole Hugging Face transformers codebase using the 270M text-only setup. An agent (Gemma 4 26B A4B with the Pi agent harness) then used those embeddings to search it.

2. Searching your own media, offline. Photos, voice memos and video clips never leave the device. Google shows this in the AI Edge Gallery app with Instant Media Search and Video Moments Finder.

3. Private RAG on a phone or laptop. Pair it with a Gemma 4 model for answers. Since both share a tokenizer and audio encoder design, running them together uses less total memory than two unrelated models.

4. Product and catalog search. Interleaved inputs let one listing, with its text, photo and video, become a single vector. A text query can then match the whole listing.

5. Fine-tuning for your own domain. Google points to Unsloth for guides on fine-tuning EmbeddingGemma 2 on your own data.


Where you can run it

The weights are on Hugging Face and Kaggle. Google says Model Garden availability is coming soon.

Where

Options

Python and servers

transformers, sentence-transformers, vLLM, SGLang

Desktop apps

Ollama, LM Studio, llama.cpp (GGUF), MLX

Browser

transformers.js, WebGPU

Mobile and edge

LiteRT, MediaPipe

Vector storage

Qdrant


EmbeddingGemma 2 or Gemini Embedding 2?

Google also has a hosted option, Gemini Embedding 2. It takes text, images, video, audio and PDFs, and supports output sizes from 128 to 3072 dimensions.

EmbeddingGemma 2 is the one to pick when you want open weights, offline use, or tight memory limits. If you're fine calling an API, the hosted model may be simpler. Google says EmbeddingGemma 2 is built from the same technology as the Gemini embedding models.


Things to watch

  • Benchmarks are Google's. Treat them as a starting point, not proof.

  • 128 dimensions hurts multimodal search. Text and code hold up. Image, video and speech drop to around 75% of full quality in Google's numbers.

  • Match your dimensions. Queries and documents must be the same size.

  • Use the task prompts for text. SearchQuery for queries, Document for documents. Media goes in without a prompt.

  • The 8K window is shared. Images, video frames and audio all use it up. Anything longer than the limits above (like a 20-minute recording) needs to be split into chunks.

  • Check your library version. You need sentence-transformers 6.1.0 or newer.


My take

If you've been stitching together a text model, an image model and a speech step just to search your own stuff, EmbeddingGemma 2 is worth a weekend test.

What I like most is the modular design. You can start with the 270M text-only setup, and add vision or audio later without redoing your old vectors. That's a gentle way to grow a project.

The open license and the small memory footprint matter too. Private, offline, multimodal search is a real option now, not a demo.

My advice is simple. Pick 50 to 100 real items from your own data, write a few queries, and see whether the right results come back. That tells you more than any leaderboard.


Comments (1)

Join the discussion by logging into your account.

Igor Ganapolsky

Igor Ganapolsky

First PostWeekend Warrior
17 hours ago

The part that bites on a mixed index is that "don't re-embed" and Matryoshka truncation are two different contracts. A 270M text query can sit next to a full-model image vector only while both sides use the same length and you normalize after the cut. If the phone query is chopped to 128 to hit about 0.25 GB per million vectors, the guide's own numbers say image, video, and speech keep around 75% of full quality while text and code stay near 90%. Adding photos later is safe only if you stored 768 or 256 up front. A prefix you already threw away cannot grow back, and the visual signal lives further down the vector. I'd keep the index at 256 for anything that might gain images, and use 128 only as a first-pass shortlist on text and code.

Anshu Pathak

Passionate developer sharing knowledge about modern web technologies and best practices.

Subscribe to Anshu Pathak's Newsletter

Direct email dispatches when new stories are published. Zero algorithms.