Google released EmbeddingGemma 2, an open multimodal embedding model built on Gemma 4 that maps text, images, audio and video into a single embedding space, according to the company's blog post. The model has 740 million parameters, with a 270 million parameter text-only configuration for lighter deployments, Google said.
Google released the model under the Apache 2.0 license and said the text-only weights need as little as 191 megabytes of active memory, with the full multimodal model running in about 567 megabytes on a Pixel 11 Pro. The context window is 8,000 tokens, four times larger than the first EmbeddingGemma, according to the post.
Google said the model scores 78.68 on the MTEB Code benchmark, a gain of nearly 10 points over EmbeddingGemma 1, and leads other multimodal embedding models under 1 billion parameters across the MTEB, MIEB Lite and MAEB benchmark suites. The model is available on Hugging Face and Kaggle with support for MediaPipe, LiteRT, transformers.js, vLLM, llama.cpp, Ollama and LMStudio, Google said.
Independent developer Simon Willison said the permissive license matters specifically for embedding models, because applications built on them generate large volumes of stored vectors that become expensive to regenerate if a provider changes or drops a model. He noted that OpenAI has in the past covered re-embedding costs for users after a model change, but said developers should not count on that from every vendor.
Embedding models are infrastructure: once an application stores vectors built on one model, switching later means reprocessing everything. An open license removes the single point of failure that comes with renting that infrastructure from one vendor, which is the specific risk this release is aimed at.