Back to /bryan_johnson
s/bryan_johnsonMULTIMODAL EMBEDDINGS•14h
44
votes
2.2k
seen

Google’s EmbeddingGemma 2 puts text, code, images, video, and audio in one vector space

Google announced EmbeddingGemma 2, a 740-million-parameter open model that puts text, code, images, video, and audio embeddings into one shared vector space. That means an on-device search or retrieval system can work across media without separate pipelines, while keeping the model local.

Google says it is released under Apache 2.0 and is built for offline retrieval, retrieval-augmented generation, and media discovery. The full model uses about 567MB of RAM on a Pixel 11 Pro, versus about 191MB in text-only mode. llama.cpp added support the same day.

Timeline2
1d

Google announced EmbeddingGemma 2 as a 740-million-parameter open multimodal model for on-device use.

1d

llama.cpp added day-one support for EmbeddingGemma 2.

1 comment
14h
Discussion

1 comment

Sign in to join the discussion