All videos
0:00 / 0:00
research

⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind

Latent Space16 June 2026Watch on YouTube

Part of series

Ep. 4 · Deepmind Gemma Gift

View the series

Description

We chat with GDM's Head of Developer Experience to catch up on Gemma 4, their ultra efficient open source model featuring a new transformer architecture featuring per-layer embeddings, enabling effective parameter offloading where only a fraction (e.g., 2B of 5B parameters) needs to be loaded into the GPU for fast inference, ideal for on-device use. We also chat Gemini Nano, Native Multimodality, and Finetuning and growth trends seen at AIE Europe. Timestamps: 0:00 Introduction to Gemma 4 and team scope 0:23 Explanation of effective vs. active parameters 1:43 On-device use cases and Gemini Nano integration 3:14 Behind the scenes of a model launch and developer ecosystem 4:29 Offline vs. API usage and future model growth 6:26 Gemma 4 multimodal capabilities and limitations 8:08 Multilingual tokenizer insights 9:30 Google's Developer Experience team at AI Engineer 10:42 Introduction to research areas: diffusion models for text 13:37 Current state of fine-tuning and community trends 16:29 Trade-offs between dense and sparse architectures 18:29 Intelligence per parameter and future research 20:09 Gemma Scope and mechanistic interpretability 21:12 The intersection of research and engineering 23:59 Perspectives on "Auto-research" and agentic automation 26:06 Team expansion, global hubs, and Kaggle integration

Topics

In this video

Related reads