{"schemaVersion":"1.0","type":"TechArticle","types":["Article","TechArticle"],"slug":"embeddinggemma-2-one-small-open-model-for-searching-text-code-images-video-and-audio-lhknc","url":"https://zyvop.com/embeddinggemma-2-one-small-open-model-for-searching-text-code-images-video-and-audio-lhknc","title":"EmbeddingGemma 2: One Small Open Model for Searching Text, Code, Images, Video and Audio","subtitle":"Google's new open model puts text, code, images, video and audio into one shared vector space, and it runs in about 567 MB of RAM. Here's how it works and how to try it.","tldr":"EmbeddingGemma 2 is a 740M-parameter open model from Google that searches five kinds of content with a single set of vectors. Here's what's new, what the numbers mean, and a short code walkthrough.","keywords":["RAG","Google Gemma","EmbeddingGemma 2","Embeddings","on-device AI"],"entities":["Anshu Pathak","RAG","Google Gemma","EmbeddingGemma 2","Embeddings","on-device AI","ZyVOP"],"keyTakeaways":["Most search setups have a hidden mess behind them.","Text goes through one model.","Photos go through another."],"headings":["EmbeddingGemma 2 at a glance","A 60-second refresher on embeddings","What's new in version 2","How it's built","What the numbers say","Shrinking vectors with Matryoshka","Try it yourself","Step 1: Load the model","Step 2: Embed text and code","Step 3: Search across media","Step 4: Shrink the vectors","A tiny local media search","Where I'd use it","Where you can run it","EmbeddingGemma 2 or Gemini Embedding 2?","Things to watch","My take","Sources and links"],"outboundLinks":["https://ai.google.dev/gemma/docs/embeddinggemma/model_card_2","https://blog.google/innovation-and-ai/technology/developers-tools/embeddinggemma-2/","https://developers.googleblog.com/en/embeddinggemma-2-the-developer-guide/","https://ai.google.dev/gemma/docs/embeddinggemma","https://huggingface.co/google/embeddinggemma-2","https://www.kaggle.com/models/google/embeddinggemma-2","https://developers.googleblog.com/en/google-ai-edge-with-embeddinggemma-2/","https://developers.google.com/edge/gallery","https://huggingface.co/spaces/google/embedding-draw-challenge","https://unsloth.ai/docs/models/embeddinggemma-2","https://qdrant.tech/blog/embeddinggemma-2/","https://ai.google.dev/gemini-api/docs/models/gemini-embedding-2-preview"],"contentText":"Most search setups have a hidden mess behind them. Text goes through one model. Photos go through another. Audio gets turned into a transcript first, then goes through the text model. Then somebody writes glue code to merge all the results. Google's new EmbeddingGemma 2 is built to remove that mess. One open model, one shared vector space, five kinds of content: text, code, images, video and audio. It launched on October 6, 2026, with open weights under the Apache 2.0 license. In this post I'll cover what it is, what changed from version 1, what the numbers say, and how to run it yourself. 💡 Quick note: The benchmark and memory numbers in this post come from Google's launch post, developer guide and docs (linked at the end). They're vendor numbers, so test on your own data before you rely on them. EmbeddingGemma 2 at a glance Released October 6, 2026 Made by Google DeepMind Built on Gemma 4 Size 740M parameters for everything, 270M for text and code only Inputs Text, code, images, video, audio Vector size 768 dimensions (can shrink to 512, 256 or 128) Context window 8,192 tokens License Apache 2.0 RAM on a Pixel 11 Pro (quantized) About 191 MB text-only, about 567 MB full model A 60-second refresher on embeddings An embedding model reads some content and turns it into a list of numbers. That list is called a vector. Content with a similar meaning ends up with similar vectors. So \"ocean waves at sunset\" lands close to a beach photo taken at dusk, and far away from a spreadsheet. Search, RAG, recommendations and clustering all lean on this. You embed your data once, store the vectors, and compare each new query's vector against them. Here's the catch with older setups. A text model can only compare text with text. To search photos using words, you needed a second model, and the two vector spaces didn't match. EmbeddingGemma 2 puts everything in the same space. So a text query can be compared directly with a photo, a video clip or a sound recording. What's new in version 2 The first EmbeddingGemma was text-only. Google says it passed 20 million downloads, so a sequel made sense. Version 2 changes four big things: More input types. Text, code, images, video and audio. The vision encoder also handles visual documents like PDFs, slides and charts. Better at code. The MTEB Code score goes from 68.76 to 78.68. That's a gain of 9.92 points, or about 14%. Longer context. 8K tokens, which is 4x what version 1 handled. Modular design. You only load the parts you need. It's also built on Gemma 4, so it shares Gemma 4's text tokenizer and audio encoder design. That detail matters later, when we get to on-device RAG. How it's built Each type of input gets its own encoder. They all end up in the same 768-dimension space. text / code ───► text model (270M) ─────────┐ │ images / video ► vision encoder (+170M) ────┼──► shared backbone ──► 768-dim vector │ audio ─────────► audio encoder (+300M) ─────┘Because the parts are separate, you can leave out what you don't need: Setup Parameters What it handles Text + code 270M Text and code only Text + vision 440M Adds images, video, visual documents Text + audio 570M Adds audio Everything 740M All five types 📌 Good to know: All four setups load from the same checkpoint and share one vector space. A query embedded with the 270M text-only setup can be compared directly with documents embedded by the full model. Google's guide also says that if you start with a text-only index and add images later, you don't need to re-embed what you already have. That opens up a neat setup. A small text-only model can handle search queries on a phone, while a bigger machine does the heavy media indexing. What the numbers say Here's what Google reports: Code search: MTEB Code goes from 68.76 to 78.68. Quality for its size: Google says it leads sub-1B multimodal embedders on benchmarks like MTEB Code and MAEB (the audio one), and beats some specialist models more than twice its size. Text quality: Multilingual text performance matches version 1. Memory: On a Pixel 11 Pro, with quantization, it needs about 191 MB of active RAM for text-only weights and about 567 MB for the full multimodal model. Input budget: The 8K window fits up to 5.5 minutes of audio, 29 images or 58 video frames, or a mix of them. ⚠️ Heads up: These are Google's own results. The full tables are in the model card. Run a small test on your own data before you commit. Shrinking vectors with Matryoshka Vectors cost storage. A million of them adds up fast. EmbeddingGemma 2 is trained with Matryoshka Representation Learning (MRL). In plain words, the most important information sits at the start of the vector. So you can chop it down to 512, 256 or 128 numbers and keep much of the quality. Here's what that looks like in bfloat16, per million vectors: Dimensions Storage Quality notes (from Google's guide) 768 about 1.5 GB Full quality 512 about 1 GB Recommended along with 768 for multimodal and visual document search 256 about 0.5 GB Keeps most of the quality on text and code, and about 95% on image, video and speech 128 about 0.25 GB Around 90% on text and code, but only around 75% on image, video and speech The 768 and 128 storage numbers (about 1.5 GB and 250 MB) come from Google's guide. The 512 and 256 rows are the same math: vectors × dimensions × 2 bytes. 💡 Tip: Use 128 for big text-only indexes and for first-pass shortlisting before a re-ranker. Google says to check your own data before using 128 for multimodal queries. Two rules to remember: Queries and documents must use the same dimension. Pass normalize_embeddings=True when you truncate. Try it yourself Google's guide uses the sentence-transformers library. You need version 6.1.0 or newer. pip install -U \"sentence-transformers[image,audio,video]\" transformersStep 1: Load the model from sentence_transformers import SentenceTransformer MODEL_ID = \"google/embeddinggemma-2\" # Full model: all modalities (740M parameters) model = SentenceTransformer(MODEL_ID)Want a lighter load? Turn off the encoders you don't need. They never get loaded into memory, so you save on both weights and peak usage. # Text only (270M parameters) text_only_model = SentenceTransformer( MODEL_ID, config_kwargs={\"vision_config\": None, \"audio_config\": None}, ) # Text, images and video (440M parameters) text_image_model = SentenceTransformer( MODEL_ID, config_kwargs={\"audio_config\": None}, ) # Text and audio (570M parameters) text_audio_model = SentenceTransformer( MODEL_ID, config_kwargs={\"vision_config\": None}, )Step 2: Embed text and code The model was trained with short task instructions. You switch them on with prompt_name. For search, queries and documents use different ones: query = \"What causes the northern lights?\" document = \"The northern lights are caused by charged particles from the sun..\" query_emb = model.encode(query, prompt_name=\"SearchQuery\") doc_emb = model.encode(document, prompt_name=\"Document\") print(model.similarity(query_emb, doc_emb))Step 3: Search across media For images, video and audio, pass a dictionary keyed by type. No prompt needed. # One text query against a photo and a sound recording image_emb = model.encode({\"image\": \"sunset_beach.jpg\"}) audio_emb = model.encode({\"audio\": \"ocean_waves.wav\"}) query_emb = model.encode(\"ocean waves at sunset\", prompt_name=\"SearchQuery\") print(model.similarity(query_emb, image_emb)) print(model.similarity(query_emb, audio_emb))You can even mix types into one embedding. Mark where each item goes with &lt;|image|&gt;, &lt;|video|&gt; or &lt;|audio|&gt;: # One embedding for a product listing with text, a photo and a video listing_emb = model.encode({ \"text\": \"Waterproof trail shoe. &lt;|image|&gt; Grip test on wet rock: &lt;|video|&gt;\", \"image\": \"trail_shoe.jpg\", \"video\": \"grip_test.mp4\", }) query_emb = model.encode(\"waterproof trail shoes\", prompt_name=\"SearchQuery\") print(model.similarity(query_emb, listing_emb))Step 4: Shrink the vectors query_emb = model.encode( query, prompt_name=\"SearchQuery\", truncate_dim=256, normalize_embeddings=True, )To use one size for every call, set it at load time: SentenceTransformer(MODEL_ID, truncate_dim=256). A tiny local media search Here's a small script that puts those pieces together. It indexes the photos and sound clips in a folder, then searches them with plain text. from pathlib import Path from sentence_transformers import SentenceTransformer model = SentenceTransformer(\"google/embeddinggemma-2\") # 1. Index: embed every photo and sound clip once index = [] for path in Path(\"media\").iterdir(): suffix = path.suffix.lower() if suffix in {\".jpg\", \".jpeg\", \".png\"}: emb = model.encode({\"image\": str(path)}) elif suffix == \".wav\": emb = model.encode({\"audio\": str(path)}) else: continue index.append((path.name, emb)) # 2. Search with text def search(text, top_k=3): q = model.encode(text, prompt_name=\"SearchQuery\") scored = [(float(model.similarity(q, emb)), name) for name, emb in index] return sorted(scored, reverse=True)[:top_k] print(search(\"ocean waves at sunset\"))For a real app, save the vectors in a vector database instead of a Python list. Google points to Qdrant for that. 📌 File notes: Video is sampled at 1 frame per second by default. Audio should be 16 kHz mono. Where I'd use it 1. Local code search for coding agents. The MTEB Code jump is the headline here. Google's demo embedded the whole Hugging Face transformers codebase using the 270M text-only setup. An agent (Gemma 4 26B A4B with the Pi agent harness) then used those embeddings to search it. 2. Searching your own media, offline. Photos, voice memos and video clips never leave the device. Google shows this in the AI Edge Gallery app with Instant Media Search and Video Moments Finder. 3. Private RAG on a phone or laptop. Pair it with a Gemma 4 model for answers. Since both share a tokenizer and audio encoder design, running them together uses less total memory than two unrelated models. 4. Product and catalog search. Interleaved inputs let one listing, with its text, photo and video, become a single vector. A text query can then match the whole listing. 5. Fine-tuning for your own domain. Google points to Unsloth for guides on fine-tuning EmbeddingGemma 2 on your own data. Where you can run it The weights are on Hugging Face and Kaggle. Google says Model Garden availability is coming soon. Where Options Python and servers transformers, sentence-transformers, vLLM, SGLang Desktop apps Ollama, LM Studio, llama.cpp (GGUF), MLX Browser transformers.js, WebGPU Mobile and edge LiteRT, MediaPipe Vector storage Qdrant EmbeddingGemma 2 or Gemini Embedding 2? Google also has a hosted option, Gemini Embedding 2. It takes text, images, video, audio and PDFs, and supports output sizes from 128 to 3072 dimensions. EmbeddingGemma 2 is the one to pick when you want open weights, offline use, or tight memory limits. If you're fine calling an API, the hosted model may be simpler. Google says EmbeddingGemma 2 is built from the same technology as the Gemini embedding models. Things to watch Benchmarks are Google's. Treat them as a starting point, not proof. 128 dimensions hurts multimodal search. Text and code hold up. Image, video and speech drop to around 75% of full quality in Google's numbers. Match your dimensions. Queries and documents must be the same size. Use the task prompts for text. SearchQuery for queries, Document for documents. Media goes in without a prompt. The 8K window is shared. Images, video frames and audio all use it up. Anything longer than the limits above (like a 20-minute recording) needs to be split into chunks. Check your library version. You need sentence-transformers 6.1.0 or newer. My take If you've been stitching together a text model, an image model and a speech step just to search your own stuff, EmbeddingGemma 2 is worth a weekend test. What I like most is the modular design. You can start with the 270M text-only setup, and add vision or audio later without redoing your old vectors. That's a gentle way to grow a project. The open license and the small memory footprint matter too. Private, offline, multimodal search is a real option now, not a demo. My advice is simple. Pick 50 to 100 real items from your own data, write a few queries, and see whether the right results come back. That tells you more than any leaderboard. Sources and links EmbeddingGemma 2 launch post (Google) EmbeddingGemma 2: The Developer Guide EmbeddingGemma docs Model card Weights on Hugging Face and Kaggle Google AI Edge blog post on EmbeddingGemma 2 Google AI Edge Gallery Embedding Draw Challenge demo Unsloth fine-tuning guide Qdrant and EmbeddingGemma 2 Gemini Embedding 2 model page","contentHash":"sha256:4c595757672eebbc1a4e65f6bbd9dabaf2f5c344e3e9e2de19e5cfece680dd28","authorName":"Anshu Pathak","authorUrl":"https://zyvop.com/author/anshu","authorSameAs":[],"category":null,"tags":["RAG","Google Gemma","EmbeddingGemma 2","Embeddings","on-device AI"],"audience":"Software engineers and developers building applications with RAG","tone":"Instructional, practical, code-first","readingTimeMinutes":9,"wordCount":2014,"faqs":null,"primaryTopic":"RAG","publishedAt":"2026-10-07T05:10:09.840Z","updatedAt":"2026-10-07T05:10:09.840Z","canonicalUrl":"https://zyvop.com/embeddinggemma-2-one-small-open-model-for-searching-text-code-images-video-and-audio-lhknc"}