June 29 – July 2, 2026 · San Francisco, CA · imported from ai.engineer's public schedule feed

AI Engineer World's Fair 2026 — unofficial import demo

Unofficial demo. This programme was imported from the AI Engineer World's Fair's own public schedule feed to show vibeboard at real conference scale. Not affiliated with, or endorsed by, the organisers.

All sessions
InferenceSession

Large clusters for small models

Daniel Svonava

When
Thursday, July 21:55 PM – 2:15 PM · 20 min
Where
Track 9San Francisco, CA · imported from ai.engineer's public schedule feed
Google Calendar

About this session

Small task-specific models are cheaper, faster and narrowly better than the frontier. But a wide catalog of small models is tricky to serve - dedicated worker pools sit idle, top-down request routers choke up on the huge volume of small requests, your users bring 100s of LoRAs.. In this talk we show how we serve 1M tokens per second with small models, how we architect our cluster for maximum throughput AND minimum latency and how we apply autoresearch to rewrite our inference code to support 10+ new models a week.

Speaker

Daniel Svonava
Daniel Svonava

CEO and co-founder, Superlinked

Daniel Svonava is CEO and co-founder of Superlinked.

More in Inference