A voice memo can become a query for a video clip, or a text query can search audio recordings. These are among the applications Google describes for EmbeddingGemma 2, which puts text, code, images, video, and audio into a shared embedding space without requiring content to be sent to a cloud embedding service.
The 740-million-parameter model, announced on October 6, 2026, expands Google’s earlier text-focused EmbeddingGemma into multimodal retrieval.
Developers have three practical choices to work through: which encoders to load, how much memory the application needs beyond the weights, and how aggressively to shrink its vector index. Downloadable weights and an Apache-2.0 license make local deployment possible. Google’s memory and performance figures still need validation against a developer’s own hardware and content.
Load Only the Encoders Your Search Needs
EmbeddingGemma 2’s full parameter count includes three modular components: a 270M text configuration, a 170M vision encoder, and a 300M audio encoder.
The following nominal parameter budgets are calculated from Google’s component counts:
| Required inputs | Components | Parameters |
|---|---|---|
| Text and code | Text configuration | 270M |
| Text plus images or video | Text configuration and vision encoder | 440M |
| Text plus audio | Text configuration and audio encoder | 570M |
| Full multimodal inputs | All three components | 740M |
Google’s description does not require a separate code encoder or list a standalone video encoder alongside vision and audio. A codebase-search tool therefore needn’t load the media encoders. A photo-search application can omit audio, while a voice-recording archive can omit vision. The complete model provides the broadest input coverage for a mixed-media library.
With a common output space, an application can compare a text query with indexed media embeddings without first converting every asset into a text description.
Accuracy may still vary across modality pairings. Test the searches a product needs, such as text-to-image or audio-to-video, instead of treating “multimodal” as one performance category.
The RAM Figures Describe Quantized Weights
Google reports approximately 191MB of active RAM for quantized text-only weights and 567MB for the full multimodal model on a Pixel 11 Pro.
These are useful deployment targets, though complete application-memory budgets will be larger. Media decoding, preprocessing, intermediate tensors, the vector index, and the application itself also need memory. Adding a generative model for RAG introduces another substantial component.
The Hugging Face repository lists approximately 1.49GB of BF16 weights before quantization. That measures something different from Google’s reported active RAM for quantized weights. When comparing deployment packages, distinguish download size, weight precision, resident model memory, and total process memory.
Measure peak memory during indexing as well as during an individual search. Successfully loading the model does not establish that a device can process a large batch of images or audio alongside its search index.
Quantization reduces the weight representation’s footprint; selective encoder loading avoids carrying components the application does not use. Both leave the surrounding retrieval pipeline to budget for.
The 8K Context Is a Shared Input Budget
EmbeddingGemma 2 supports an 8K-token context window, four times the original EmbeddingGemma’s announced context capacity.
Google says this accommodates up to 5.5 minutes of audio, 29 images, 58 video frames, or interleaved combinations of modalities. These are alternative capacity examples. They cannot all be added together, since mixed inputs consume the available context budget.
The larger window provides more room to represent a coherent document section or media segment, but chunking remains useful. Embedding a long recording as one item would make it difficult to return the precise moment relevant to a query, even if the content fitted within the model’s limits. A useful implementation would divide it into retrievable segments and retain timestamps.
Video needs similar care. Google’s 58-frame figure describes input capacity, with the represented duration depending on frame sampling. Segmentation and sampling rules need to preserve the events users will search for.
Longer chunks provide context; shorter chunks can provide more precise results. The 8K window leaves room to tune that tradeoff without prescribing a single indexing strategy.
Matryoshka Dimensions Shrink the Index, Not the Model
EmbeddingGemma 2 produces 768-dimensional embeddings and supports truncation to 512, 256, or 128 dimensions through Matryoshka Representation Learning. This training approach makes shorter prefixes of the embedding useful representations, so developers can choose a smaller vector size without retaining all 768 coordinates.
For one million vectors stored as float32 values, the raw vector payload would be:
| Dimensions | Raw storage | Reduction versus 768 |
|---|---|---|
| 768 | 3.072GB | Baseline |
| 512 | 2.048GB | 1.5× smaller |
| 256 | 1.024GB | 3× smaller |
| 128 | 512MB | 6× smaller |
These calculated decimal storage sizes exclude database indexes, metadata, identifiers, and the original content. Google’s “up to 6×” reduction applies to the vector payload, not necessarily the entire database or application.
Vector size does not change the model’s parameter count. It is an index-storage choice, separate from weight quantization.
Choose dimensions through retrieval testing: build a representative query set, compare the supported sizes, and measure whether relevant items still appear among the top results. A compact index that drops important matches is a poor trade for storage savings.
Keep indexed vectors and query vectors dimensionally consistent, and follow the model card’s truncation guidance. Plan to re-embed existing collections when changing models, too. Matching vector dimensions alone does not make two models’ embedding spaces interchangeable.
The Code Benchmark Gain Is Vendor-Reported
Google reports an MTEB Code score of 78.68, compared with 68.76 for the earlier EmbeddingGemma, a 9.92-point improvement.

Sources
- 740-million-parameter modelblog.google
- Hugging Face repositoryhuggingface.co





