What did Google release?

EmbeddingGemma 2 is a compact multimodal embedding model from Google DeepMind. Unlike a chat model that generates an answer, it converts an input into a numerical representation that can be compared with other representations. The release accepts text, code, images, video, audio and mixtures of those formats, mapping them into a common 768-dimensional space.

That shared space is useful when the query and the stored item are different kinds of media. A spoken note could retrieve a related video clip, or a text description could locate an image without a matching filename. Our text embeddings guide explains why these vectors measure learned similarity rather than factual correctness.

How does one model handle several media types?

The model combines a 270M-parameter text component with a 170M vision encoder and a 300M audio encoder, for 740M parameters when every component is loaded. Video is represented through sampled frames, while interleaved inputs use placeholder tokens to mark where images, video or audio appear alongside text. Mean pooling and a projection layer produce the final vector.

This design lets applications compare items that would otherwise require separate specialist pipelines. It is still an embedding system, not a complete reasoning agent: it retrieves or groups nearby items, and another component decides what to display or generate. That distinction is important when building multimodal AI workflows.

Why is it aimed at on-device use?

Google designed the encoders to be loaded selectively. A text-only setup uses 270M parameters; text plus vision uses 440M; text plus audio uses 570M; the full configuration uses 740M. Google says a quantized text-only configuration used about 191MB of active RAM on a Pixel 11 Pro, while the complete multimodal model used about 567MB.

Local inference can keep personal files on the device, reduce network round trips and support offline search. Those advantages depend on the surrounding application: indexing media still consumes compute, storage and battery, and local processing does not automatically make an app private if it later uploads queries, vectors or results.

How can developers access and use it?

Google released the model under Apache 2.0 and provides weights through Hugging Face and Kaggle. The launch lists support for Sentence Transformers, Transformers, MLX, vLLM, llama.cpp, SGLang, Ollama, LM Studio, transformers.js and WebGPU, alongside Google's MediaPipe and LiteRT tools for mobile and edge deployment.

Typical uses include semantic media search, local codebase retrieval, classification, clustering and retrieval-augmented generation. Teams planning a RAG system should separate the retrieval layer from the model that writes the final answer, as described in our RAG explainer. An embedding can return a close match without proving that the match is correct or complete.

What do the benchmarks and vector options show?

Google reports a multilingual MTEB score of 61.36 and a code MTEB score of 78.68, up from 68.76 for the earlier EmbeddingGemma on the code benchmark. The model card also reports results for image, visual-document, video and audio benchmarks. These figures cover the full-precision checkpoint and should be treated as Google-reported evaluation results, not independent guarantees for every dataset.

Matryoshka Representation Learning allows the native 768-dimension vector to be truncated to 512, 256 or 128 dimensions. Shorter vectors reduce storage and speed similarity search, but the model card shows larger losses at 128 dimensions, especially for multimodal tasks. This is a deployment tradeoff similar to model quantization: smaller representations help resource use, but quality must be checked on representative queries.

What limitations should teams test first?

Embedding quality depends on the task prefix, input preparation, language, media quality and the relevance judgments used for evaluation. Google warns that shortened vectors must be re-normalized and that queries and indexed items must use the same dimension. It also recommends bfloat16 or float32 because float16 can produce NaNs or silently degraded vectors.

Before replacing an existing search system, teams should build a labeled test set from their own documents and media, measure recall as well as latency, and inspect failure cases across languages and formats. They should also verify the library implementation, chosen precision, device thermals and storage budget. The release makes local multimodal retrieval more accessible, but it does not remove the need for application-level privacy controls, monitoring or human review.

Common questions

Is EmbeddingGemma 2 a chatbot or generative model?

No. It creates vectors for search, retrieval, classification and similarity. A separate generative model can use retrieved material, but EmbeddingGemma 2 does not write the final response itself.

Can it run without sending files to a cloud service?

Yes, Google designed it for local inference and provides edge tooling. Whether an application remains fully offline depends on how the developer handles indexing, telemetry, queries and downstream models.

Is the smallest 128-dimension output always the best choice?

No. It offers the largest storage reduction, but Google's table shows a clearer multimodal quality drop. Validate 128, 256, 512 and 768 dimensions on the intended retrieval set before choosing.

THE TAKEAWAY

What to remember

Google's new open embedding model brings text, code, image, video and audio retrieval into one on-device-oriented system. Start with a representative evaluation set, then choose encoders and vector size according to measured quality, memory and latency rather than the launch headline.

Sources & further reading

  1. EmbeddingGemma 2: an open, lightweight multimodal embedding model ↗
  2. EmbeddingGemma 2 model card ↗
How this story was made

Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.

Our editorial standards
Back to all stories