June 29 – July 2, 2026 · San Francisco, CA · imported from ai.engineer's public schedule feed

AI Engineer World's Fair 2026 — unofficial import demo

Unofficial demo. This programme was imported from the AI Engineer World's Fair's own public schedule feed to show vibeboard at real conference scale. Not affiliated with, or endorsed by, the organisers.

All sessions
Voice & Realtime AISession

Voice Agents Can Just Do Things

Charlie Guo

When
Tuesday, June 3011:40 AM – 12:00 PM · 20 min
Where
Track 6San Francisco, CA · imported from ai.engineer's public schedule feed
Google Calendar

About this session

Too many voice AI integrations still treat speech as fancier chat: audio in, audio out. But we're at a point where speech can be a control plane for software, and most developers are unaware that voice has become a capability overhang. Current realtime models can understand intent, call tools, speak while work is underway, recover from corrections, and decide what the user actually needs to hear. As a result, we're seeing three practical patterns emerge: voice-to-action, systems-to-voice, and voice-to-voice. We’ll show how each pattern changes the architecture, where Realtime 2’s reasoning and tool-calling matter, and why chained STT / LLM / TTS systems start to break down as the interaction patterns become richer.

Speaker

Charlie Guo
Charlie Guo

Developer Experience Engineer, OpenAI

Charlie Guo is a Developer Experience Engineer at OpenAI, where he helps developers build with the OpenAI API. He is also the author of Artificial Ignorance, an AI publication at the intersection of engineering and intelligence. Before joining OpenAI, Charlie spent more than a decade building products and internal tools, including as a startup founder. He is based in Berkeley, California.

More in Voice & Realtime AI