Large clusters for small models
Daniel Svonava
- When
- Thursday, July 21:55 PM – 2:15 PM · 20 min
- Where
- Track 9San Francisco, CA · imported from ai.engineer's public schedule feed
About this session
Small task-specific models are cheaper, faster and narrowly better than the frontier. But a wide catalog of small models is tricky to serve - dedicated worker pools sit idle, top-down request routers choke up on the huge volume of small requests, your users bring 100s of LoRAs.. In this talk we show how we serve 1M tokens per second with small models, how we architect our cluster for maximum throughput AND minimum latency and how we apply autoresearch to rewrite our inference code to support 10+ new models a week.
Speaker
More in Inference
- Operating Distributed Inference Systems at ScaleThursday, July 2 · 10:45 AM – 11:05 AM · Track 9
- Routing LLM Inference in Production: From Engine Signals to PolicyThursday, July 2 · 11:10 AM – 11:30 AM · Track 9
- Are LLM Performance Benchmarks Reliable?Thursday, July 2 · 11:40 AM – 12:00 PM · Track 9
- All the Things We Have to Do to Satisfy Your Insatiable Need for TokensThursday, July 2 · 11:40 AM – 12:00 PM · Leadership 1
- Vertical Mobility: Building an AI Inference Platform That Scales from MVP to Trillion-Parameter WorkloadsThursday, July 2 · 12:05 PM – 12:25 PM · Track 9
For developers: this programme is open data — JSON, iCal, schedule XML and an MCP endpoint.Show endpointsHide
- JSONEvery published session and speaker, in one request./aie-worldsfair-2026-import/feed.json
- iCalSubscribe in Google, Apple or Outlook Calendar./aie-worldsfair-2026-import/feed.ics
- Schedule XMLfrab / pentabarf — the format conference apps import./aie-worldsfair-2026-import/feed.xml
- MCP + RESTPoint Claude at the programme. OpenAPI 3.1 included./agents
No key, no signup, CORS open. Everything here is generated from the same data the organisers edit.