Google DeepMind has released EmbeddingGemma 2, an open multimodal model that maps text, code, images, video and audio into a single 768-dimensional vector space. It contains 740 million parameters, supports an 8K token context window, and carries an Apache 2.0 license. The release targets on-device search, classification and privacy-first retrieval-augmented generation.
In this article
Weights are available immediately on Hugging Face and Kaggle. Builders can run it via Ollama, llama.cpp GGUF and LiteRT.
What an Embedding Model Does
These models convert content into numerical vectors that capture meaning. Similar items cluster together, making them easy to search and compare. In a RAG pipeline, these vectors allow an LLM to retrieve fresh information it was not trained on. Generating embeddings locally keeps data on the device, cuts latency and works offline.
One Vector Space for Every Modality
EmbeddingGemma 2 is built on the Gemma 4 architecture. A text query can retrieve a photo. A voice memo can retrieve a video clip. Interleaved inputs, such as a product listing with text, images and a demo video, produce a single embedding.
The design is modular. It has three parts:
- Text and code backbone: 270M parameters (130M transformer plus 140M embedder)
- Vision encoder: 170M parameters, optional
- Audio encoder: 300M parameters, optional
Developers load only what they need: 270M for text, 440M for text and vision, 570M for text and audio, or 740M for everything. All setups share one vector space. A query embedded with the text-only setup can match documents embedded by the full model.
The context window is 8,192 tokens, 4x larger than version 1. That fits about 29 images, 58 video frames or 5.5 minutes of audio.
Benchmarks
Google research team reports leading scores among sub-1B multimodal embedders on MTEB Code and MAEB. Full-precision results at 768 dimensions:
| Benchmark | EmbeddingGemma 2 | EmbeddingGemma 1 |
|---|---|---|
| MTEB multilingual v2 | 61.36 | 61.15 |
| MTEB Code v1 | 78.68 | 68.76 |
| MIEB lite (image) | 64.64 | n/a |
| MMEB v2 overall | 59.01 | n/a |
| MSEB retrieval (sound) | 69.54 | n/a |
| MAEB (audio) | 49.39 | n/a |
Source: EmbeddingGemma 2 model card
Code retrieval gains 9.92 points, roughly 14%. Multilingual text quality holds steady. Bigger models still lead some boards. Qwen3-VL-Embedding-2B reports 73.2 on its own MMEB-V2 run, with about 2.7x the parameters and no audio support.
Built for Phones and Laptops
With quantization on a Pixel 11 Pro, active RAM is about 191MB for text-only weights. The full multimodal model needs about 567MB. Quantization-aware training compresses weights to INT4 and INT8. The Google AI Edge team measured 37.3 ms per image on a MacBook M5 Pro GPU, using a 70-token vision budget.
Matryoshka Representation Learning (MRL) lets developers truncate vectors to 512, 256 or 128 dimensions. Moving from 768 to 128 dimensions cuts storage up to 6x. At 256 dimensions, MTEB multilingual only slips from 61.36 to 60.41. At 128 dimensions, MMEB drops to 45.65, so Google recommends 128d mainly for text-only workloads.
Interactive Explainer
EmbeddingGemma 2 vs. Closest Competitors
| Feature | EmbeddingGemma 2 | EmbeddingGemma 1 | Qwen3-VL-Embedding-2B | LCO-Embedding-Omni-3B | Gemini Embedding 2 |
|---|---|---|---|---|---|
| Developer | Google DeepMind | Google DeepMind | Alibaba Qwen | LCO-Embedding (research) | |
| Parameters | 740M (270M text-only) | 308M | 2B | 3B backbone (5B listed on HF) | Not disclosed |
| Text / code | Yes | Yes | Yes | Yes | Yes |
| Images | Yes | No | Yes | Yes | Yes |
| Video | Yes | No | Yes | Yes | Yes |
| Audio | Yes | No | No | Yes | Yes |
| Output dims (MRL) | 768 (512, 256, 128) | 768 (down to 128) | Up to 2048 (64 to 2048) | Not stated | 3072 (128 to 3072) |
| Context | 8,192 tokens | 2K tokens | 32K tokens | Not stated | 8,192 tokens |
| Languages | 100+ | 100+ | 30+ | Not stated | 100+ |
| License / access | Apache 2.0, open weights | Open weights (Gemma terms) | Apache 2.0, open weights | Apache 2.0, open weights | Paid API only |
| Published on-device RAM | ~191MB text, ~567MB full | Under 200MB | Not published | Not published | Cloud only |
| Source | Model card | Docs | HF card | HF card | API docs |
How to Run It
It runs on sentence-transformers v6.1.0+, Transformers, vLLM, SGLang, MLX, llama.cpp, Ollama, LM Studio, LiteRT and MediaPipe. Qdrant covers vector storage and Unsloth covers fine-tuning. ML Kit support for Android, with NPU acceleration, is coming within weeks.
pip install -U "sentence-transformers[image,audio,video]" transformers
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("google/embeddinggemma-2")
q = model.encode("What causes the northern lights?", prompt_name="SearchQuery")
d = model.encode("Charged particles from the sun.", prompt_name="Document")
print(model.similarity(q, d))
On Ollama, run ollama pull embeddinggemma-2. Tags range from 270m (378MB) to 740m (1.3GB). Demos live in Google AI Edge Gallery. See the developer guide for more.



