June 29 – July 2, 2026 · San Francisco, CA · imported from ai.engineer's public schedule feed

AI Engineer World's Fair 2026 — unofficial import demo

Unofficial demo. This programme was imported from the AI Engineer World's Fair's own public schedule feed to show vibeboard at real conference scale. Not affiliated with, or endorsed by, the organisers.

All sessions
Voice & Realtime AISession

Speech-to-Speech Model Research at Google DeepMind

Valeria Wu Fon, Tom Ouyang

When
Tuesday, June 3011:10 AM – 11:30 AM · 20 min
Where
Track 6San Francisco, CA · imported from ai.engineer's public schedule feed
Google Calendar

About this session

Most voice interfaces today are built as a 3-way cascade system (ASR/LLM/TTS). While functional, this cascaded approach introduces latency bottlenecks, strips away non-verbal nuance, and limits emotion-aware, multi-turn dialogue. Today, we are witnessing a profound shift toward native speech-to-speech models that process audio natively from end to end. In this session, we’ll explore the exciting paradigm at Google DeepMind to train speech-to-speech models for real-time voice agents. We will cover the high-level product and research challenges of building voice agents that feel truly conversational, optimizing for fluid turn-taking and low latency while maintaining enterprise-grade intelligence.

Speakers (2)

Valeria Wu Fon
Valeria Wu Fon

Product Manager, Google DeepMind

Product Manager at Google DeepMind for Gemini's speech to speech model. Previously worked across early stage companies, venture/banking, and a brief stint at a surf hostel. Valeria studied Symbolic Systems (CS, Neuroscience, and Philosophy) at Stanford, where she focused on human-centered AI. Originally from Lima, Peru, she is a single-digit golfer and a dedicated foodie.

Tom Ouyang
Tom Ouyang

Principal Engineer, Google DeepMind

Tom Ouyang is a Principal Engineer at Google DeepMind, where he works on research and development for Gemini Audio, focusing on real-time capabilities like natural dialog, streaming translation, and audio understanding. Previously, he spent five years as a Principal Software Engineer at Waymo, developing machine learning models for autonomous vehicle perception. Prior to Waymo, Tom spent over six years at Google working in the area of mobile text entry and language modeling. He holds a PhD in computer science from the Massachusetts Institute of Technology.

More in Voice & Realtime AI