Developers

Google's EmbeddingGemma 2 Unifies Multimodal Search on Mobile Devices

Google has released EmbeddingGemma 2, a 740-million-parameter model that handles text, code, images, video and audio search on a single device without requiring transcription or captioning. The open-source model uses modular encoders and Matryoshka compression to fit within mobile constraints.

4 min read
Your phone’s vector index might be bigger than the AI model running it

Developers can now deploy a unified multimodal search system on phones using Google's latest embedding model, released Tuesday. EmbeddingGemma 2 consolidates text, code, image, video and audio retrieval into a 740-million-parameter architecture that consumed approximately 567MB of active RAM with quantization when tested on a Pixel 11 Pro.

The model is built atop Gemma 4 and projects all five input modalities into a shared 768-dimensional vector space. This approach eliminates the need to caption images or transcribe audio before indexing them alongside text. Google distributed the model weights under Apache 2.0, with on-device deployment available immediately through LiteRT and MediaPipe Tasks. An Android ML Kit integration featuring NPU acceleration for compatible devices will arrive within the coming weeks.

Modular encoders, one vector space

Rather than forcing developers to load the entire model, EmbeddingGemma 2 uses modular encoders that load only the components needed for specific data types. The text and code foundation requires 270 million parameters and approximately 191MB of active RAM on the same test device. Adding the vision encoder for images and video brings the parameter count to 440 million; the audio encoder alone reaches 570 million; both together comprise the full 740 million. All configurations map into the same embedding space, allowing teams to expand from text-only indexing to include image or audio search without re-embedding previously stored data.

Google also quadrupled the context window from 2,048 to 8,192 tokens, which the company says covers as much as 5.5 minutes of audio, 29 images or 58 video frames in a single input.

Video frames are sampled at one frame per second by default, so the 58-frame capacity translates to roughly one minute of video. Google's Video Moments Finder demonstration indexes video frames and audio segments locally, then searches them using plain text queries to locate specific moments without generating captions or transcripts. The Instant Media Search feature applies the same technique to photos and videos stored on a phone, maintaining embeddings in SQLite and refreshing results as the user types.

Matryoshka shrinks the index

On mobile devices, the vector index can consume as much storage as the model itself. A million 768-dimensional bfloat16 vectors occupy roughly 1.5GB. Google trained EmbeddingGemma 2 using Matryoshka Representation Learning, a technique that permits developers to reduce embeddings to 512, 256 or 128 dimensions without retraining. At 256 dimensions, a million-vector index shrinks to approximately 500MB while retaining most of the full-size quality for text and code retrieval and roughly 95% quality for image, video and speech retrieval.

The compression trade-off becomes more pronounced at 128 dimensions, where Google reports text and code performance around 90% but multimodal retrieval at approximately 75%. The company recommends testing this configuration on actual data before depending on it for multimodal queries. This efficiency-versus-quality balance has become a recurring theme in embedding model releases this season, as demonstrated when Cohere's faster query model showed minimal retrieval degradation in its own benchmarks.

On-device code search

Code retrieval operates on the same 270-million-parameter base as text processing. Google achieved an MTEB Code score of 78.68 for EmbeddingGemma 2, a substantial improvement from 68.76 for the original EmbeddingGemma. To demonstrate performance in an agent workflow, Google indexed the Hugging Face Transformers repository using the text-only configuration and paired the index with Gemma 4 26B A4B running in Pi. This setup mirrored the agent framework behind a workaround for an MCP server that previously consumed 18,000 tokens before executing any action. In Google's implementation, EmbeddingGemma 2 handled repository retrieval while the larger model controlled agent behavior.

The embedding model also supports classification through MediaPipe Decision, which compares incoming embeddings against candidate descriptions rather than generating responses. A chess demonstration showed this approach evaluating 500 options per turn in under 100 milliseconds. For applications where agents waste tokens on decisions that never required generated output, this provides a substantially more efficient alternative.

Leaner local RAG pipelines

EmbeddingGemma 2 incorporates Gemma 4's text tokenizer and audio encoder design, reducing memory requirements when both models run together on a device. Google has demonstrated multimodal retrieval functioning on its flagship phone. Real-world performance with larger indexes, diverse hardware platforms and non-Google applications remains to be seen.

Source: The New Stack · Reporting supplemented by The Silicon Ledger staff.