June 29 – July 2, 2026 · San Francisco, CA · imported from ai.engineer's public schedule feed

AI Engineer World's Fair 2026 — unofficial import demo

Unofficial demo. This programme was imported from the AI Engineer World's Fair's own public schedule feed to show vibeboard at real conference scale. Not affiliated with, or endorsed by, the organisers.

All sessions
Workshops Day 1Workshop

2 hr deep dive on LLM Inference at Scale — Part 1 of 2

Harshul Jain, Tanmay Sah

When
Monday, June 2912:10 PM – 1:10 PM · 60 min
Where
Track 3San Francisco, CA · imported from ai.engineer's public schedule feed
Google Calendar

About this session

Most engineers using LLMs can call an API. Far fewer can explain why their model is slow, why it's running out of memory, or how the inference engines powering every major LLM API actually work. This workshop walks through the full inference stack — from how a transformer generates a single token to serving billions of tokens a day with vLLM, SGLang, TensorRT-LLM, Ray, and KServe/llm-d. 60% explanation with live demos, 40% hands-on exercises. Attendees leave with a running vLLM server they benchmarked themselves. Based on the open-source practitioners handbook being built live at github.com/harshuljain13/llm-inference-at-scale

(NOTE: this is a 2 hour workshop that happens over lunch break - you should try to have lunch before or after if attending)

compute kindly sponsored by Coreweave/Marimo!

Speakers (2)

Harshul Jain
Harshul Jain

Senior Software Engineer - ML/AI, Audible

Harshul Jain is a Senior Software Engineer at Audible (Amazon) who builds ML and LLM infrastructure at scale — AI Search serving 10M users, a feature store processing 100K transactions per second, and LLM serving and evaluation systems powering GenAI in production. He is writing LLM Inference at Scale, a benchmark-driven handbook on GPU memory engineering, attention optimization, and production LLM serving backed by a companion repository that gained 100+ clones in its first week with zero promotion.

Tanmay Sah
Tanmay Sah

Senior Quantitative Modeler, Zions Bancorporation

Tanmay Sah, PhD, is a quantitative modeler and AI researcher working at the intersection of predictive modeling, model risk, AI evaluation, and agentic AI systems. His notable work includes research on AI agent verification; TanML, an open-source automated machine learning model validation toolkit; and Decoding Reddit Memes Virality. He is especially interested in the next generation of trustworthy AI systems: agents that can reason, use tools, remain auditable, and operate safely under real-world constraints.

More in Workshops Day 1