Routing LLM Inference in Production: From Engine Signals to Policy
Qianru Lao, Lu Zhang
- When
- Thursday, July 211:10 AM – 11:30 AM · 20 min
- Where
- Track 9San Francisco, CA · imported from ai.engineer's public schedule feed
About this session
Production LLM apps need more than a fast model: they need an inference routing layer that can choose where each request should run as engines, capacity, latency, and geography cost change. This talk shares a generalized Inference Load Balancer (ILB) proxy/controller architecture. A low-latency proxy applies routing weights and request-path signals, while a controller computes source-cluster-to-engine weights from demand, capacity/performance profiles, replica state, and geography cost. We will cover the practical debugging patterns AI engineers need: reading engine signals, explaining why a request went to one backend instead of another, handling retries and load shedding, and keeping routing behavior observable without exposing OpenAI-specific internals or non-public metrics.
Speakers (2)
Member of Technical Staff, OpenAI
Qianru Lao is a Member of Technical Staff on the Inference team at OpenAI, where she works on infrastructure for large-scale model serving. Previously, she contributed to the open-source Delta Lake project at Databricks and worked on distributed storage systems at Alibaba Cloud and infrastructure tooling at Google. She holds degrees in Computational Science and Engineering from Harvard and Computer Science from Sun Yat-sen University.
Member of Technical Staff, OpenAI
Lu is an engineer working on large-scale inference platforms, focused on making AI model serving reliable, efficient, and scalable. His work includes distributed systems, workload scheduling, performance optimization, and production reliability. Previously, Lu built and operated GPU clusters supporting large machine learning workloads.
More in Inference
- Operating Distributed Inference Systems at ScaleThursday, July 2 · 10:45 AM – 11:05 AM · Track 9
- Are LLM Performance Benchmarks Reliable?Thursday, July 2 · 11:40 AM – 12:00 PM · Track 9
- All the Things We Have to Do to Satisfy Your Insatiable Need for TokensThursday, July 2 · 11:40 AM – 12:00 PM · Leadership 1
- Vertical Mobility: Building an AI Inference Platform That Scales from MVP to Trillion-Parameter WorkloadsThursday, July 2 · 12:05 PM – 12:25 PM · Track 9
- Stop Model Shopping: Why Ownership Beats Choice in the Agent StackThursday, July 2 · 12:05 PM – 12:25 PM · Leadership 1
For developers: this programme is open data — JSON, iCal, schedule XML and an MCP endpoint.Show endpointsHide
- JSONEvery published session and speaker, in one request./aie-worldsfair-2026-import/feed.json
- iCalSubscribe in Google, Apple or Outlook Calendar./aie-worldsfair-2026-import/feed.ics
- Schedule XMLfrab / pentabarf — the format conference apps import./aie-worldsfair-2026-import/feed.xml
- MCP + RESTPoint Claude at the programme. OpenAPI 3.1 included./agents
No key, no signup, CORS open. Everything here is generated from the same data the organisers edit.