saved
EmbeddingGemma 2: an open, lightweight multimodal embedding model
Hraness wrote this summary from a saved copy of the source. Quotations are taken word for word from the source.
gist
Google engineer Sahil Dua launches EmbeddingGemma 2, a 740M-parameter open multimodal embedder under Apache 2.0 built on Gemma 4. It puts code, images, video, and audio in one embedding space for on-device search and privacy-first RAG. Modular encoders start at 270M for text-only; Matryoshka truncates 768-d vectors to 128-d for up to 6x storage savings. An 8K context window covers minutes of audio or dozens of images locally.
ideas
- Multimodal shared space on Gemma 4. EmbeddingGemma 2 expands the text-only first release to unify code, images, video, and audio for cross-modal retrieval on consumer hardware.
- Modular and storage-efficient. Text-only needs 270M parameters; optional vision (170M) and audio (300M) encoders complete the stack; MRL truncates vectors from 768 to 512, 256, or 128 dimensions.
- Stronger code and on-device footprint. MTEB Code rises 9.92 points to 78.68; quantized text-only weights use ~191MB active RAM on a Pixel 11 Pro, ~567MB for the full multimodal model.
- 8K context for local media. Four times EmbeddingGemma 1, enough for about 5.5 minutes of audio, 29 images, or 58 video frames interleaved on-device.
- Ship where developers already build. Weights on Hugging Face and Kaggle; MediaPipe, LiteRT, transformers, Ollama, llama.cpp, and partners cover on-device and server paths.
quotes
“The developer community’s response blew past our expectations.”
“expanding beyond text to unify code, images, video, and audio in a shared embedding space”
“This provides up to 6x storage reduction for local vector databases and memory usage.”
“Requires as little as 270M parameters for text-only workloads”