Productionizing LLM Gateways: Architecture, Tradeoffs, and Hard Lessons from the Trenches
Kanish Manuja
- When
- Tuesday, June 302:25 PM – 2:45 PM · 20 min
- Where
- Leadership 1San Francisco, CA · imported from ai.engineer's public schedule feed
About this session
As organizations scale their use of large language models, the biggest challenge is no longer prompting, it’s productionizing. This session dives deep into building and operating an LLM gateway that sits between applications and model providers, handling routing, observability, cost control, reliability, and safety at scale. Drawing from real world experience, this talk breaks down the architecture of a production LLM gateway, including model abstraction layers, request orchestration, fallback strategies, caching, rate limiting, and evaluation pipelines. We’ll explore hard tradeoffs such as latency vs. cost, quality vs. determinism, and vendor lock-in vs. flexibility. Attendees will leave with concrete design patterns, failure modes to avoid, and a mental model for turning LLM experiments into resilient, scalable systems.
Speaker
Principal Software Engineer, Twilio Inc.
Kanish Manuja is a principal AI engineer at Twilio, where he leads production LLM gateway and AI platform systems for enterprise-scale AI applications. His work focuses on building reliable, secure, and observable infrastructure for large language model adoption, including multi-tenant gateways, authentication and authorization, guardrails, audit logging, fallback strategies, and production readiness for GenAI workloads. Kanish has worked across AI platform engineering, conversational intelligence, and distributed systems, helping teams move from experimentation to production-grade LLM deployments. He has led efforts around LLM reliability, governance, tenant isolation, provider abstraction, and operational controls for high-scale customer-facing systems. In this session, Kanish will share practical lessons from designing and operating LLM gateway systems in production, including architectural tradeoffs, failure modes, platform boundaries, and what teams should consider before standardizing LLM access across an organization.
More in AI-Native Enterprises
- Building the engine while flying the plane — launching the Figma MCP serverTuesday, June 30 · 11:10 AM – 11:30 AM · Leadership 1
- Agentic SDLC at Uber: Building Blocks for Uber's Software FactoryTuesday, June 30 · 11:40 AM – 12:00 PM · Leadership 1
- Scaling Code Quality: Building uReview, Uber’s Multi-Agent Code Review EngineTuesday, June 30 · 12:05 PM – 12:25 PM · Leadership 1
- AI Evals Platform for Cross-Functional Teams at ScaleTuesday, June 30 · 1:55 PM – 2:15 PM · Leadership 1
- From AI-Assisted to AI-Native: Building a Frontier Development TeamTuesday, June 30 · 2:50 PM – 3:10 PM · Leadership 1
For developers: this programme is open data — JSON, iCal, schedule XML and an MCP endpoint.Show endpointsHide
- JSONEvery published session and speaker, in one request./aie-worldsfair-2026-import/feed.json
- iCalSubscribe in Google, Apple or Outlook Calendar./aie-worldsfair-2026-import/feed.ics
- Schedule XMLfrab / pentabarf — the format conference apps import./aie-worldsfair-2026-import/feed.xml
- MCP + RESTPoint Claude at the programme. OpenAPI 3.1 included./agents
No key, no signup, CORS open. Everything here is generated from the same data the organisers edit.