Google releases EmbeddingGemma 2 for local multimodal search

Illustration, not documentary evidence of the event.
Google has released EmbeddingGemma 2, an open model that turns text, images, video, audio and code into numerical vectors for finding and comparing similar content, according to the-decoder.com, citing Google. The model has 740 million parameters. Google claims it beats competing models up to twice its size on multimodal embedding benchmarks.
The report says the model runs locally without an API key and gives browser query times of about 20 to 70 milliseconds through WebGPU. It also reports roughly 191 MB of RAM usage and local vector database storage reductions of up to six times. These are reported performance figures, not results independently verified here. Weights are available on Hugging Face and Kaggle.
For developers choosing a search model, the first decision is whether the application needs multiple media types. The report also describes a 270-million-parameter option for text-only tasks. That makes the smaller version a sensible evaluation starting point for a text-only collection. For a mixed collection, test whether the larger model retrieves the right material for representative user queries across the media types the application actually uses.
The benchmark claim suggests compact models deserve consideration, but it does not settle that choice. Before replacing an existing embedding model, compare relevant results on the same documents and queries. Measure browser latency, memory use and index size on the devices intended to run the application. Treat the reported speed and storage figures as comparisons to investigate, rather than capacity assumptions.
The report says pairing EmbeddingGemma 2 with small open models such as Gemma 4 can support offline retrieval-augmented generation without sending data to external servers. For teams pursuing that setup, the useful acceptance test is the complete application: disconnect it from the network and check that ingestion, retrieval and answer generation still work. Separately inspect network activity during normal operation before concluding that the deployed application keeps its data local.