2 hr deep dive on LLM Inference at Scale — Part 2 of 2
Harshul Jain, Tanmay Sah
- When
- Monday, June 291:15 PM – 2:15 PM · 60 min
- Where
- Track 3San Francisco, CA · imported from ai.engineer's public schedule feed
About this session
Most engineers using LLMs can call an API. Far fewer can explain why their model is slow, why it's running out of memory, or how the inference engines powering every major LLM API actually work. This workshop walks through the full inference stack — from how a transformer generates a single token to serving billions of tokens a day with vLLM, SGLang, TensorRT-LLM, Ray, and KServe/llm-d. 60% explanation with live demos, 40% hands-on exercises. Attendees leave with a running vLLM server they benchmarked themselves. Based on the open-source practitioners handbook being built live at github.com/harshuljain13/llm-inference-at-scale
(NOTE: this is a 2 hour workshop that happens over lunch break - you should try to have lunch before or after if attending)
Speakers (2)
Senior Software Engineer - ML/AI, Audible
Harshul Jain is a Senior Software Engineer at Audible (Amazon) who builds ML and LLM infrastructure at scale — AI Search serving 10M users, a feature store processing 100K transactions per second, and LLM serving and evaluation systems powering GenAI in production. He is writing LLM Inference at Scale, a benchmark-driven handbook on GPU memory engineering, attention optimization, and production LLM serving backed by a companion repository that gained 100+ clones in its first week with zero promotion.
Senior Quantitative Modeler, Zions Bancorporation
Tanmay Sah, PhD, is a quantitative modeler and AI researcher working at the intersection of predictive modeling, model risk, AI evaluation, and agentic AI systems. His notable work includes research on AI agent verification; TanML, an open-source automated machine learning model validation toolkit; and Decoding Reddit Memes Virality. He is especially interested in the next generation of trustworthy AI systems: agents that can reason, use tools, remain auditable, and operate safely under real-world constraints.
More in Workshops Day 1
- Cooking with CodexMonday, June 29 · 9:00 AM – 11:00 AM · Track 3
- The best SDLC is the one you build yourself: Why orchestration changes everythingMonday, June 29 · 9:00 AM – 11:00 AM · Track 4
- AI Security Engineer Foundations + CertificateMonday, June 29 · 9:00 AM – 11:00 AM · Track 5
- Total Recall: Agent Memory and Harness EngineeringMonday, June 29 · 9:00 AM – 11:00 AM · Track 6
- Open-Source Inference Engineering for the Agentic EraMonday, June 29 · 9:00 AM – 11:00 AM · Track 8
For developers: this programme is open data — JSON, iCal, schedule XML and an MCP endpoint.Show endpointsHide
- JSONEvery published session and speaker, in one request./aie-worldsfair-2026-import/feed.json
- iCalSubscribe in Google, Apple or Outlook Calendar./aie-worldsfair-2026-import/feed.ics
- Schedule XMLfrab / pentabarf — the format conference apps import./aie-worldsfair-2026-import/feed.xml
- MCP + RESTPoint Claude at the programme. OpenAPI 3.1 included./agents
No key, no signup, CORS open. Everything here is generated from the same data the organisers edit.