Optimizing Open Models for Production Grade Inference
Sujee Maniyam, Dylan Bristot
- When
- Thursday, July 22:25 PM – 2:45 PM · 20 min
- Where
- Expo Stage 1 NESan Francisco, CA · imported from ai.engineer's public schedule feed
About this session
Open-source foundation models are rapidly closing the gap with proprietary systems, enabling organizations to build powerful AI applications with greater flexibility and control. However, deploying these models in production introduces a new set of challenges: latency, throughput, scalability, and cost efficiency.In this talk, we'll explore the modern inference optimization techniques that power large-scale AI systems in production. Topics include KV cache optimization, cache-aware routing, prefill/decode disaggregation, speculative decoding, and other emerging approaches used to improve performance and reduce infrastructure costs.Through practical examples and real-world architecture patterns, attendees will gain a deeper understanding of how to run open models efficiently at scale.
Speakers (2)
Developer Advocate, Nebius
Sujee Maniyam is a developer advocate at Nebius with a background in ML, data engineering, technical training, and production inference education.
Product Marketing, Token Factory
Dylan Bristot works in product marketing at Token Factory and speaks on practical approaches for scaling open-model inference in production AI infrastructure.
For developers: this programme is open data — JSON, iCal, schedule XML and an MCP endpoint.Show endpointsHide
- JSONEvery published session and speaker, in one request./aie-worldsfair-2026-import/feed.json
- iCalSubscribe in Google, Apple or Outlook Calendar./aie-worldsfair-2026-import/feed.ics
- Schedule XMLfrab / pentabarf — the format conference apps import./aie-worldsfair-2026-import/feed.xml
- MCP + RESTPoint Claude at the programme. OpenAPI 3.1 included./agents
No key, no signup, CORS open. Everything here is generated from the same data the organisers edit.