I Monitored Crime Audio. Voice Agents Scare Me More.
Sumanyu Sharma
- When
- Tuesday, June 302:25 PM – 2:45 PM · 20 min
- Where
- Track 6San Francisco, CA · imported from ai.engineer's public schedule feed
About this session
Bad voice-agent calls are starting to look less like QA bugs and more like incident scenes. I learned that instinct at Citizen, where noisy radio, ambiguous speech, fast-moving incidents, and real-time alerts became information people might actually act on. That work was stressful for obvious reasons. Voice agents scare me more. Not because they sound creepy. Because they sound good enough that people trust them. And now they are connected to calendars, CRMs, EHRs, reservation systems, refunds, transfers, account data, and support workflows. At Hamming, we monitor more than 10,000 voice agents and have analyzed millions of calls. The weird thing you learn at that scale is that production voice agents do not usually fail like demos. They fail quietly. The agent sounds natural, but misses a two-word answer. It handles the happy path, but loses the plot when the caller interrupts. It says the address was updated, but no tool call happened. It supports six languages, but gets worse at the switch point between two of them. This talk is about treating every bad voice-agent call like an incident scene. The evidence is there if you collect it: transcript, waveform, latency waterfall, interruption points, ASR uncertainty, tool trace, system-of-record state, and post-call outcome. At Tesla, I learned that autonomous systems need release gates and regression loops before they hit the real world. At Citizen, I learned that messy audio becomes safety-critical when people act on it. Voice agents need both instincts. The takeaway is a voice-agent forensics loop. What did the caller say? What did the agent think happened? What did the tool actually do? What does the system of record say? And how do we turn that weird production failure into a regression test before it happens 10,000 more times?
Speaker
Founder/CEO, Hamming AI
Sumanyu Sharma is the Founder & CEO of Hamming AI, a YC company that invented automated testing, monitoring, and red-teaming for AI voice agents. Hamming helps teams catch issues before their agents confidently say the wrong thing to a real customer. Before Hamming, Sumanyu worked on high-stakes systems at Tesla and Citizen across revenue, safety, and emergency response. He started in AI research at Waterloo, teaching models to search X-rays by meaning. Put simply, he has spent his career turning “the AI seems fine” into “the AI passed the test".
More in Voice & Realtime AI
- The New Primitives: Building AI-Native SoftwareTuesday, June 30 · 10:45 AM – 11:05 AM · Track 6
- Speech-to-Speech Model Research at Google DeepMindTuesday, June 30 · 11:10 AM – 11:30 AM · Track 6
- Voice Agents Can Just Do ThingsTuesday, June 30 · 11:40 AM – 12:00 PM · Track 6
- Tolan: Voice-First AI CompanionTuesday, June 30 · 1:30 PM – 1:50 PM · Track 6
- 5 Voice Agent Failure Modes You'll Hit in Week OneTuesday, June 30 · 1:55 PM – 2:15 PM · Track 6
For developers: this programme is open data — JSON, iCal, schedule XML and an MCP endpoint.Show endpointsHide
- JSONEvery published session and speaker, in one request./aie-worldsfair-2026-import/feed.json
- iCalSubscribe in Google, Apple or Outlook Calendar./aie-worldsfair-2026-import/feed.ics
- Schedule XMLfrab / pentabarf — the format conference apps import./aie-worldsfair-2026-import/feed.xml
- MCP + RESTPoint Claude at the programme. OpenAPI 3.1 included./agents
No key, no signup, CORS open. Everything here is generated from the same data the organisers edit.