All the Things We Have to Do to Satisfy Your Insatiable Need for Tokens
Daniel Kim, Michelle Nguyen
- When
- Thursday, July 211:40 AM – 12:00 PM · 20 min
- Where
- Leadership 1San Francisco, CA · imported from ai.engineer's public schedule feed
About this session
Every time the industry figures out how to serve tokens faster and cheaper, the appetite grows to match. Models get bigger, contexts get longer, agents start chaining thousands of calls together. The finish line keeps moving. This talk is a technical tour through everything the industry has done to keep up, led by two experts in high-performance inference. We'll start with the optimizations that made hardware work harder without changing the underlying architecture. Then we'll go up a level with techniques that work smarter across requests and across the model itself. And finally, a peek into the future with heterogeneous disaggregated inference, the architectural shift that splits prefill and decode across specialized hardware, and even more advanced forms of hardware specialization coming your way soon. Token demand is about to get a lot more insatiable. Let's see what the future has in store for us!
Speakers (2)
Head of Growth, Cerebras
Daniel Kim works on large-scale inference systems at Cerebras, which runs the world's fastest AI inference on the Wafer-Scale Engine (WSE-3), the largest chip ever built. More recently, Daniel has turned to building AI agents that accelerate Cerebras's own hardware and software development. Outside of work, you can find him relaxing in the park, eating spicy noodles, and recently running!
Co-Founder, Gimlet Labs
Michelle Nguyen is cofounder of Gimlet Labs, the first multi-silicon inference cloud to run agentic workloads across different types of hardware, where she leads engineering. She was the first engineer at Pixie Labs where she worked across the stack on projects ranging from Pixie's deployment mechanisms to its distributed query engine. Before Pixie, Michelle was at Trifacta helping build intuitive and interactive UIs. Michelle holds a MS and BS in EECS from UC Berkeley.
More in Inference
- Operating Distributed Inference Systems at ScaleThursday, July 2 · 10:45 AM – 11:05 AM · Track 9
- Routing LLM Inference in Production: From Engine Signals to PolicyThursday, July 2 · 11:10 AM – 11:30 AM · Track 9
- Are LLM Performance Benchmarks Reliable?Thursday, July 2 · 11:40 AM – 12:00 PM · Track 9
- Vertical Mobility: Building an AI Inference Platform That Scales from MVP to Trillion-Parameter WorkloadsThursday, July 2 · 12:05 PM – 12:25 PM · Track 9
- Stop Model Shopping: Why Ownership Beats Choice in the Agent StackThursday, July 2 · 12:05 PM – 12:25 PM · Leadership 1
For developers: this programme is open data — JSON, iCal, schedule XML and an MCP endpoint.Show endpointsHide
- JSONEvery published session and speaker, in one request./aie-worldsfair-2026-import/feed.json
- iCalSubscribe in Google, Apple or Outlook Calendar./aie-worldsfair-2026-import/feed.ics
- Schedule XMLfrab / pentabarf — the format conference apps import./aie-worldsfair-2026-import/feed.xml
- MCP + RESTPoint Claude at the programme. OpenAPI 3.1 included./agents
No key, no signup, CORS open. Everything here is generated from the same data the organisers edit.