Two Bugs That Hid in Plain Sight: A vLLM Debugging Detective Story
Asaf Gardin, Yuval Belfer
- When
- Thursday, July 23:20 PM – 3:40 PM · 20 min
- Where
- Track 9San Francisco, CA · imported from ai.engineer's public schedule feed
About this session
Your model generates gibberish. Once every thousand prompts. High confidence scores. No crashes. No warnings. We hit this twice while building Jamba models. First: A request gets misclassified during scheduling, loads stale state from a previous prompt cache slot, and confidently generates nonsense. Second: Logprob spikes during RL training that looked like training instability-until we noticed they tracked with rollout count, then with cache size. In this talk, we'll walk through both debugging journeys-the false starts, how we instrumented vLLM to thread request IDs through the forward pass, the search for variables that change failure structure rather than magnitude, and the lesson both share: distributed inference systems fail silently. No stack trace. No sanitizer warning. Just wrong answers with perfect confidence. You'll learn how to build comparison scripts that expose logprob divergence, force memory pressure to surface rare bugs, and shrink a distributed RL training mystery into a reproducible single-script failure. Walk away knowing how to debug vLLM when it lies to you quietly.
Speakers (2)
Senior Software Engineer/Inference Engineer, AI21
Asaf Gardin is a Senior Software Engineer on the inference team at AI21 Labs, where he works on high-performance LLM inference and the production deployment of the Jamba hybrid SSM-Transformer models. He's an active vLLM committer, contributing to quantization, scheduling, and support for Mamba-based architectures. His talk covers two production bugs in vLLM's Mamba support - a scheduler edge case that corrupted SSM state under memory pressure, and a 32-bit integer overflow in a CUDA kernel that surfaced as RL training instability - both root-caused at AI21 and fixed upstream. He also built Kernel Academy, a browser-based tutorial for learning Triton GPU programming. Previously at IBM.
Sr. Developer Advocate, AI21
Yuval is a Senior Developer Advocate at AI21 Labs, where he helps engineers go from "it works in the demo" to "it works in production." He hosts the YAAP podcast (Yet Another AI Podcast) and teaches applied GenAI on various programs. His work spans RAG, fine-tuning, agents, and evaluation (or Yuval-uation, if you're nasty).
More in Inference
- Operating Distributed Inference Systems at ScaleThursday, July 2 · 10:45 AM – 11:05 AM · Track 9
- Routing LLM Inference in Production: From Engine Signals to PolicyThursday, July 2 · 11:10 AM – 11:30 AM · Track 9
- Are LLM Performance Benchmarks Reliable?Thursday, July 2 · 11:40 AM – 12:00 PM · Track 9
- All the Things We Have to Do to Satisfy Your Insatiable Need for TokensThursday, July 2 · 11:40 AM – 12:00 PM · Leadership 1
- Vertical Mobility: Building an AI Inference Platform That Scales from MVP to Trillion-Parameter WorkloadsThursday, July 2 · 12:05 PM – 12:25 PM · Track 9
For developers: this programme is open data — JSON, iCal, schedule XML and an MCP endpoint.Show endpointsHide
- JSONEvery published session and speaker, in one request./aie-worldsfair-2026-import/feed.json
- iCalSubscribe in Google, Apple or Outlook Calendar./aie-worldsfair-2026-import/feed.ics
- Schedule XMLfrab / pentabarf — the format conference apps import./aie-worldsfair-2026-import/feed.xml
- MCP + RESTPoint Claude at the programme. OpenAPI 3.1 included./agents
No key, no signup, CORS open. Everything here is generated from the same data the organisers edit.