Speech-to-Speech Model Research at Google DeepMind
Valeria Wu Fon, Tom Ouyang
- When
- Tuesday, June 3011:10 AM – 11:30 AM · 20 min
- Where
- Track 6San Francisco, CA · imported from ai.engineer's public schedule feed
About this session
Most voice interfaces today are built as a 3-way cascade system (ASR/LLM/TTS). While functional, this cascaded approach introduces latency bottlenecks, strips away non-verbal nuance, and limits emotion-aware, multi-turn dialogue. Today, we are witnessing a profound shift toward native speech-to-speech models that process audio natively from end to end. In this session, we’ll explore the exciting paradigm at Google DeepMind to train speech-to-speech models for real-time voice agents. We will cover the high-level product and research challenges of building voice agents that feel truly conversational, optimizing for fluid turn-taking and low latency while maintaining enterprise-grade intelligence.
Speakers (2)
Product Manager, Google DeepMind
Product Manager at Google DeepMind for Gemini's speech to speech model. Previously worked across early stage companies, venture/banking, and a brief stint at a surf hostel. Valeria studied Symbolic Systems (CS, Neuroscience, and Philosophy) at Stanford, where she focused on human-centered AI. Originally from Lima, Peru, she is a single-digit golfer and a dedicated foodie.
Principal Engineer, Google DeepMind
Tom Ouyang is a Principal Engineer at Google DeepMind, where he works on research and development for Gemini Audio, focusing on real-time capabilities like natural dialog, streaming translation, and audio understanding. Previously, he spent five years as a Principal Software Engineer at Waymo, developing machine learning models for autonomous vehicle perception. Prior to Waymo, Tom spent over six years at Google working in the area of mobile text entry and language modeling. He holds a PhD in computer science from the Massachusetts Institute of Technology.
More in Voice & Realtime AI
- The New Primitives: Building AI-Native SoftwareTuesday, June 30 · 10:45 AM – 11:05 AM · Track 6
- Voice Agents Can Just Do ThingsTuesday, June 30 · 11:40 AM – 12:00 PM · Track 6
- Tolan: Voice-First AI CompanionTuesday, June 30 · 1:30 PM – 1:50 PM · Track 6
- 5 Voice Agent Failure Modes You'll Hit in Week OneTuesday, June 30 · 1:55 PM – 2:15 PM · Track 6
- I Monitored Crime Audio. Voice Agents Scare Me More.Tuesday, June 30 · 2:25 PM – 2:45 PM · Track 6
For developers: this programme is open data — JSON, iCal, schedule XML and an MCP endpoint.Show endpointsHide
- JSONEvery published session and speaker, in one request./aie-worldsfair-2026-import/feed.json
- iCalSubscribe in Google, Apple or Outlook Calendar./aie-worldsfair-2026-import/feed.ics
- Schedule XMLfrab / pentabarf — the format conference apps import./aie-worldsfair-2026-import/feed.xml
- MCP + RESTPoint Claude at the programme. OpenAPI 3.1 included./agents
No key, no signup, CORS open. Everything here is generated from the same data the organisers edit.