Google released EmbeddingGemma 2 on Tuesday, a 740-million-parameter open model that handles text, code, image, video, and audio retrieval while using about 567MB of active RAM with quantization during the company's tests on a Pixel 11 Pro. Constructed on Gemma 4, the model projects all five input types into an identical 768-dimensional vector space. Photos no longer require captions and audio files don't need transcripts before either can be queried alongside text. Google published the weights under Apache 2.0, and on-device deployment through LiteRT and MediaPipe Tasks is ready now, with Android ML Kit integration and NPU acceleration planned for the coming weeks.
The complete model contains 740 million parameters, but engineers only activate the encoders their data needs. Text and code operate on a 270-million-parameter foundation, which Google measured at roughly 191MB of active RAM on the same device. Adding the vision encoder for photos and video pushes the model to 440 million parameters, while incorporating the audio encoder instead lifts it to 570 million, and loading both reaches the full 740 million. Each configuration maps into the same embedding space, so developers who begin with a text-only index can introduce image or audio search later without re-embedding stored content. The company expanded the context window fourfold from 2,048 to 8,192 tokens, which handles up to 5.5 minutes of audio, 29 images, or 58 video frames in one input. Video gets sampled at one frame per second by default, so those 58 frames represent just under a minute of footage.
On a phone, the index can rival the model itself for storage, since a million 768-dimensional bfloat16 vectors consume approximately 1.5GB. Google trained EmbeddingGemma 2 with Matryoshka Representation Learning, which allows engineers to truncate those embeddings to 512, 256, or 128 dimensions without retraining. At 256 dimensions, that million-vector index shrinks to about 500MB. The company reports the shorter embeddings preserve most of their full-size quality on text and code and roughly 95% on image, video, and speech retrieval. The trade gets steeper at 128 dimensions, where Google places text and code at around 90% but multimodal retrieval at about 75%, and the company advises testing that setting on actual data before depending on it for multimodal queries. Code operates on the same 270-million-parameter base as text, and Google reports an MTEB Code score of 78.68 for EmbeddingGemma 2, up from 68.76 for the original version.
The flexibility to load only necessary encoders addresses a core challenge in on-device deployment: developers can start lean and scale up modalities without rebuilding their entire infrastructure. Because all encoders map to the same vector space, a text search system can evolve into a multimodal one incrementally, preserving existing embeddings and avoiding the storage and compute costs of re-indexing. The Matryoshka approach extends that efficiency to the index itself, letting teams balance retrieval quality against memory constraints on devices with limited resources. Google's Video Moments Finder demo indexed video frames and audio chunks locally, then searched them with plain text to locate a specific moment, generating no captions or transcripts in the process. Instant Media Search applied the same method to photos and videos on a phone, storing embeddings in SQLite and refreshing results as the user typed. The company also demonstrated classification through MediaPipe Decision, which compares an incoming embedding against candidate descriptions instead of generating a response, evaluating 500 chess positions per turn in under 100 milliseconds in an on-device demo.
Google has shown multimodal retrieval working on its flagship phone so far, but performance with larger indexes, different hardware, and applications beyond Google demos remains to be seen. EmbeddingGemma 2 borrows Gemma 4's text tokenizer and audio encoder architecture, so the two models require less memory when running together on a device. The modular design and shared embedding space give developers a path to multimodal search without the overhead of separate systems for each data type, while Matryoshka truncation offers a practical trade-off between index size and retrieval accuracy. For teams building local RAG pipelines or on-device agents, the model's ability to handle code, images, and audio in a single framework could eliminate the tokenization overhead and latency of sending queries to the cloud. The 78.68 MTEB Code score and sub-100-millisecond classification times suggest the efficiency gains are real, even if the full picture won't emerge until the model faces production workloads and hardware beyond Google's own devices.

