AIbaseUpdated

Google stuffs AI search into phones: EmbeddingGemma2 is open source, under 600MB, and can search text, images, audio and video even offline

Google has released EmbeddingGemma2, its first natively multimodal open-source embedding model with 740M parameters. It maps text, code, images, video and audio into a single semantic space, letting cross-modal retrieval…

Google has turned "finding things locally" into a small model that fits in a phone. Its newly released EmbeddingGemma2 is the company's first natively multimodal open-source embedding model. With a parameter count of 740M, it can nonetheless map text, code, images, video and audio all into the same semantic space, making cross-modal retrieval run this lightly on-device for the first time.

The most tangible selling points are small size and offline capability. The full multimodal model takes up less than 600MB of memory, which means it can run directly on phones, laptops and in browsers, without uploading any content to the cloud throughout the process — your photo albums, recordings and videos are all retrieved on your own device, so the privacy barrier is naturally guarded. The context window is 8K, four times that of the previous generation, allowing it to ingest longer material before matching; language coverage has also been extended to more than 100 languages, so finding things across languages is no longer an obstacle.

On performance, Google claims it clearly leads open-source rivals of the same scale on retrieval tasks involving images, video and code. Even more thoughtful is its modular design: unused parts can be loaded on demand, and the index size can also be compressed, saving both space and compute. In real-world scenarios, what it can do is very concrete — semantically dig out a particular photo or piece of media from a local photo album, jump directly to a specific segment in a long video, retrieve files scattered in various places offline, or search functions and snippets by intent in a local codebase, without first pushing the repository to some remote service to query.

When a multimodal embedding model becomes small enough to tuck into a phone and open enough for anyone to modify, the center of gravity of AI retrieval is quietly shifting from "handing data to the cloud" back to "keeping capabilities on the device." Google's move has pried open a crack for privacy-first on-device search.


Original source

AIbase

Content notes

Original publication and rights belong to the source.

Machine translation · Refer to the original