June 29 – July 2, 2026 · San Francisco, CA · imported from ai.engineer's public schedule feed

AI Engineer World's Fair 2026 — unofficial import demo

Unofficial demo. This programme was imported from the AI Engineer World's Fair's own public schedule feed to show vibeboard at real conference scale. Not affiliated with, or endorsed by, the organisers.

All sessions
Session

Optimizing Open Models for Production Grade Inference

Sujee Maniyam, Dylan Bristot

When
Thursday, July 22:25 PM – 2:45 PM · 20 min
Where
Expo Stage 1 NESan Francisco, CA · imported from ai.engineer's public schedule feed
Google Calendar

About this session

Open-source foundation models are rapidly closing the gap with proprietary systems, enabling organizations to build powerful AI applications with greater flexibility and control. However, deploying these models in production introduces a new set of challenges: latency, throughput, scalability, and cost efficiency.In this talk, we'll explore the modern inference optimization techniques that power large-scale AI systems in production. Topics include KV cache optimization, cache-aware routing, prefill/decode disaggregation, speculative decoding, and other emerging approaches used to improve performance and reduce infrastructure costs.Through practical examples and real-world architecture patterns, attendees will gain a deeper understanding of how to run open models efficiently at scale.

Speakers (2)

Sujee Maniyam
Sujee Maniyam

Developer Advocate, Nebius

Sujee Maniyam is a developer advocate at Nebius with a background in ML, data engineering, technical training, and production inference education.

Dylan Bristot
Dylan Bristot

Product Marketing, Token Factory

Dylan Bristot works in product marketing at Token Factory and speaks on practical approaches for scaling open-model inference in production AI infrastructure.