Realtime Voice Agents with Frontier Intelligence
Bohan Li
- When
- Tuesday, June 302:50 PM – 3:10 PM · 20 min
- Where
- Track 6San Francisco, CA · imported from ai.engineer's public schedule feed
About this session
Dive into how the EliseAI voice agent harness orchestrates multiple models with jagged capability profiles to achieve realtime latency without sacrificing intelligence. Reduces p90 effective latency overhead of ASR, TTS, and tool calling to sub 200ms, unlocking frontier models like GPT 5.5 for voice. ### ASR: Eager Speculative Transcription We introduce speculative transcription by pairing local Whisper or Parakeet fine-tunes for speed with API models like Scribe, Nova, or Gemini Flash for accuracy. A local content match classifier operates at sub 10ms latency, allowing us to immediately trigger the downstream pipeline from the fast local transcription and dynamically replace text with the more accurate transcription if significant differences occur. This process runs on a eager 100ms VAD delay, securely releasing the generated response audio only after a fixed silence threshold has passed. ### LLM: Async background tool injection To eliminate expensive tool calling round trips, we implement system leveraging async background tool injection where the primary model makes no direct tool calls. Instead, local fine-tuned tool-calling models continuously observe the realtime transcription stream in the background. "Fake" tool call traces are then injected into the primary LLM’s context, which primes it for immediate, one-shot response generation. ### TTS: Prefix caching and infilling Many Agent responses start with the same set of 3-6 words. We can cache this audio, releasing it immediately while we infill the remaining response audio conditioned on this prefix to preserve speech prosody. With this approach, a relatively small cache can achieve a 90% hit rate across a wide range of voices, languages and model providers.
Speaker
Staff Software Engineer, EliseAi
Bo has over 10 years of experience building real time systems across databases, decentralized finance, self driving cars, and voice AI. He previously worked as an Member of Technical Staff at Cartesia and is currently at EliseAI, building AI Agents for Housing and Healthcare that improve how we live.
More in Voice & Realtime AI
- The New Primitives: Building AI-Native SoftwareTuesday, June 30 · 10:45 AM – 11:05 AM · Track 6
- Speech-to-Speech Model Research at Google DeepMindTuesday, June 30 · 11:10 AM – 11:30 AM · Track 6
- Voice Agents Can Just Do ThingsTuesday, June 30 · 11:40 AM – 12:00 PM · Track 6
- Tolan: Voice-First AI CompanionTuesday, June 30 · 1:30 PM – 1:50 PM · Track 6
- 5 Voice Agent Failure Modes You'll Hit in Week OneTuesday, June 30 · 1:55 PM – 2:15 PM · Track 6
For developers: this programme is open data — JSON, iCal, schedule XML and an MCP endpoint.Show endpointsHide
- JSONEvery published session and speaker, in one request./aie-worldsfair-2026-import/feed.json
- iCalSubscribe in Google, Apple or Outlook Calendar./aie-worldsfair-2026-import/feed.ics
- Schedule XMLfrab / pentabarf — the format conference apps import./aie-worldsfair-2026-import/feed.xml
- MCP + RESTPoint Claude at the programme. OpenAPI 3.1 included./agents
No key, no signup, CORS open. Everything here is generated from the same data the organisers edit.