EmbeddingGemma 2: Google's Open Model for Your Site Search

Google DeepMind released EmbeddingGemma 2 on October 6, 2026: an open model with 740 million parameters that turns text, code, images, audio and video into comparable vectors. It ships under an Apache 2.0 license and is built to run on the device itself. It matters to anyone with a site search, a catalog or their own document library.

What we know

  • What it is: an embedding model. It does not write text or answer questions. It turns each piece of content into a list of numbers, a vector, so content can be compared by meaning. Google bills it as the most capable on-device model for multimodal embeddings.
  • What it understands: text, code, images, audio and video in one shared vector space, per the announcement. That allows searching one format with another: type a sentence, find a photo or an audio clip.
  • Size: 740 million parameters. Text-only work needs as few as 270 million; the vision encoder (170 million) and the audio encoder (300 million) are optional.
  • License: Apache 2.0, which permits commercial use, with open weights.
  • Vectors: 768 dimensions, which can be cut down to 512, 256 or 128. Google says that saves up to six times the storage.
  • How much it takes at once: about 8,000 tokens, four times the first version. Google translates that into up to 5.5 minutes of audio, 29 images or 58 video frames.
  • Memory: with quantization, on a Google Pixel 11 Pro, about 191 MB of active RAM for the text-only version and about 567 MB for the full model, per the announcement.
  • Languages: over 100 for text and code, per the documentation.
  • Performance, on Google's figures: on the MTEB Code benchmark it moves from 68.76 to 78.68 points over the first version. On multilingual text, Google says it matches its predecessor.
  • Where to get it: Hugging Face and Kaggle. Google says its enterprise platform will follow soon. The announcement lists support for transformers.js and WebGPU in the browser, for Ollama, llama.cpp and LM Studio, and for Qdrant to store the vectors.
  • Background: the first EmbeddingGemma, a text-only model, came out last year and passed 20 million downloads, according to Google.

What changes and what doesn't

What changes: there is now a small, open model that searches across formats. Until now, Google's open embedding model handled text only. Because it runs on your own hardware, there is no per-query API fee, though the hardware it runs on still costs money, and your documents do not go to a third party. Google highlights privacy, lower latency and offline use.

It is not a chatbot. To answer in sentences you need a generative model next to it; Google suggests Gemma 4. It is not Gemini Embedding either: Google says it shares technology with those models, but this one is downloaded and run on your own.

It does not improve your Google rankings by itself. It is a building block for the search inside your own site or app. That is this article's observation.

The performance figures are Google's, and independent tests are still missing. The memory figure is for one specific phone and a quantized version. And support for over 100 languages says nothing about quality in yours: test it.

"Open" does not mean "installed" either. Someone has to integrate it, and it needs a vector database and upkeep.

How to tell if your site needs it

  1. Look at what people search for on your site and don't find. If your analytics or your CMS logs internal searches, review the ones that return nothing. That is the size of the problem.
  2. Signs it could help: a catalog with many photos and short descriptions, a library of video or audio (courses, podcasts, recordings), long documentation, or customers who search in their own words and not in yours.
  3. Signs it won't: a site with a handful of pages, where a clear menu does more than any search box.
  4. Ask for a small test before deciding. A hundred products or pages and twenty real customer searches, in the language your customers use. Compare the results with the search you have today.
  5. Ask whoever integrates it three things. Where the vectors are generated and stored (the server or the visitor's browser), how many dimensions they use, and how often they are rebuilt when the catalog changes.
  6. If you handle sensitive documents, generating vectors on your own hardware avoids sending them out. Check where the vectors are stored as well.
  7. Keep the basics in place. Titles, descriptions and image alt text are still needed, for Google and for people.

Related: Mistral Large 4: 1 Trillion Parameters, Weights This Month

Sources

Updates: when independent tests of the model appear or Google brings it to more platforms, it will be added here with a link.