"My name is... my name is...": A Linguistic Map for Building and Debugging Voice Agents
Midam Kim
- When
- Tuesday, June 303:20 PM – 3:40 PM · 20 min
- Where
- Track 6San Francisco, CA · imported from ai.engineer's public schedule feed
About this session
Every voice AI engineer has heard it: a caller repeating their name three times, getting more frustrated with each attempt. The logs look clean. Confidence scores look fine. Linguistics can help solving the mystery. By the end of this talk, you'll have a diagnostic framework for the failures that slip past standard metrics, a way to turn "the agent just didn't get it" into concrete, debuggable failure modes. The framework maps three levels of linguistic structure (sounds, words, and interactions) against the two dimensions every voice agent engineer already works in: what we hear (speech recognition) and what we speak (speech synthesis). That 3×2 grid surfaces problems your current tooling can't see, including: 1. Why your user cannot make your system understand their name 2. Why a single well-intentioned vocabulary hint can cause catastrophic drops in a non-English language 3. Why a transcript that's "cumulatively correct" can still ruin the user experience Drawing on examples from production multilingual voice AI work, I'll show where linguistic expertise connects to the engineering decisions you're already making and where it reveals failure modes that confidence scores will never warn you about. Who this is for: Voice AI engineers, ML practitioners on Voice AI pipelines, and anyone who's watched clean logs while their agent quietly fails real users.
Speaker
ML Engineer, ServiceNow
Midam Kim is an ML Engineer at ServiceNow, where she builds and evaluates a multilingual voice AI platform spanning a dozen languages. She holds a PhD in Linguistics from Northwestern University and has backgrounds in linguistics, speech science, cognitive science, machine learning, and business. Her work sits at the rare intersection of production ML engineering and speech science—translating decades of linguistic research into the engineering decisions voice AI teams are making right now.
More in Voice & Realtime AI
- The New Primitives: Building AI-Native SoftwareTuesday, June 30 · 10:45 AM – 11:05 AM · Track 6
- Speech-to-Speech Model Research at Google DeepMindTuesday, June 30 · 11:10 AM – 11:30 AM · Track 6
- Voice Agents Can Just Do ThingsTuesday, June 30 · 11:40 AM – 12:00 PM · Track 6
- Tolan: Voice-First AI CompanionTuesday, June 30 · 1:30 PM – 1:50 PM · Track 6
- 5 Voice Agent Failure Modes You'll Hit in Week OneTuesday, June 30 · 1:55 PM – 2:15 PM · Track 6
For developers: this programme is open data — JSON, iCal, schedule XML and an MCP endpoint.Show endpointsHide
- JSONEvery published session and speaker, in one request./aie-worldsfair-2026-import/feed.json
- iCalSubscribe in Google, Apple or Outlook Calendar./aie-worldsfair-2026-import/feed.ics
- Schedule XMLfrab / pentabarf — the format conference apps import./aie-worldsfair-2026-import/feed.xml
- MCP + RESTPoint Claude at the programme. OpenAPI 3.1 included./agents
No key, no signup, CORS open. Everything here is generated from the same data the organisers edit.