Operating Distributed Inference Systems at Scale
Nishant Gupta, Naman Ahuja
- When
- Thursday, July 210:45 AM – 11:05 AM · 20 min
- Where
- Track 9San Francisco, CA · imported from ai.engineer's public schedule feed
About this session
Inference has rapidly become one of the most important infrastructure problems in modern computing. As AI systems evolve into autonomous agents with persistent memory, tool usage, and multi-step reasoning, traditional inference architectures struggle under growing demands for latency, throughput, cost efficiency, and reliability. In this talk, I’ll share lessons from building large-scale elastic compute and AI infrastructure systems powering production workloads. We’ll explore the modern inference stack and the architectural patterns emerging to support next-generation agentic AI systems. Topics include distributed inference architectures for large-scale AI systems, GPU scheduling and elastic compute for inference workloads, multi-tenant inference infrastructure, caching, batching, latency optimization strategies, reliability and fault isolation for inference systems, observability and control loops for AI serving platforms, balancing cost, throughput, and user experience, and why inference is becoming an infrastructure orchestration problem. Attendees will gain practical insights into designing scalable, resilient, and cost-efficient inference platforms for modern AI workloads.
Speakers (2)
Software Engineer, Tech Lead, Meta
I am a Staff Software Engineer and Researcher at Meta, specializing in large-scale distributed systems, AI infrastructure, and operational resilience. Within Meta Superintelligence Labs, I build agentic infrastructure that enables AI systems to operate reliably in production through evaluation, auditing, safety controls, feedback loops, and human oversight. I previously led the development of Meta’s next-generation elastic compute infrastructure, managing roughly 30% of fleet capacity across tens of millions of servers in 20+ geo-distributed datacenters, delivering billions of dollars in infrastructure savings while shaping multi-year strategy with executive leadership. My research focuses on resource optimization, reliability, and safe AI deployment at scale. I designed and deployed Dynamic Idle Resource Leasing, a production system that safely oversubscribes datacenter capacity while preserving strict reliability guarantees. I have authored research papers with 90+ citations. I am passionate about building scalable, fault-tolerant systems and translating cutting-edge research into real-world infrastructure that delivers measurable impact.
Senior Software Engineer, Meta
Software Engineer at Meta working at the intersection of distributed systems, AI infrastructure, and hardware enablement. He focuses on adopting new hardware platforms across Meta’s datacenter fleet to support AGI-scale workloads and production AI systems. His work spans capacity management, autoscaling, reliability engineering, and datacenter-scale resource optimization. He has led initiatives to integrate cutting-edge hardware, improve compute utilization, and build reliable platforms for large-scale AI workloads and agentic workflows.
More in Inference
- Routing LLM Inference in Production: From Engine Signals to PolicyThursday, July 2 · 11:10 AM – 11:30 AM · Track 9
- Are LLM Performance Benchmarks Reliable?Thursday, July 2 · 11:40 AM – 12:00 PM · Track 9
- All the Things We Have to Do to Satisfy Your Insatiable Need for TokensThursday, July 2 · 11:40 AM – 12:00 PM · Leadership 1
- Vertical Mobility: Building an AI Inference Platform That Scales from MVP to Trillion-Parameter WorkloadsThursday, July 2 · 12:05 PM – 12:25 PM · Track 9
- Stop Model Shopping: Why Ownership Beats Choice in the Agent StackThursday, July 2 · 12:05 PM – 12:25 PM · Leadership 1
For developers: this programme is open data — JSON, iCal, schedule XML and an MCP endpoint.Show endpointsHide
- JSONEvery published session and speaker, in one request./aie-worldsfair-2026-import/feed.json
- iCalSubscribe in Google, Apple or Outlook Calendar./aie-worldsfair-2026-import/feed.ics
- Schedule XMLfrab / pentabarf — the format conference apps import./aie-worldsfair-2026-import/feed.xml
- MCP + RESTPoint Claude at the programme. OpenAPI 3.1 included./agents
No key, no signup, CORS open. Everything here is generated from the same data the organisers edit.