Model role and official documentation
How one representation space connects different media →
The first EmbeddingGemma focused on text.
Google DeepMind · Released 2026-10-06
Create comparable representations of text, photos, audio and video for on-device media retrieval.
Original model explanation ↗No comparable leaderboard score is included yet. Explore the evaluations and examples below.
Model role and official documentation
The first EmbeddingGemma focused on text.
Official phone-retrieval example
In Google’s phone demo, a query about a turtle eating returns several candidate moments, each with a timestamp.
Flash advances coding, multi-step reasoning and agent work at the same speed and price positioning; Flash Cyber is a separate security-focused variant.
Original release ↗The new generation is announced with access initially limited to invited testers; evaluation results do not establish public availability.
Original release ↗Image generation and conversational editing advance, emphasizing local edits, subject consistency and text rendering.
Original release ↗The retrieval model extends text representations to images, audio and video, with public weights and phone demos.
Original release ↗