June 29 – July 2, 2026 · San Francisco, CA · imported from ai.engineer's public schedule feed

AI Engineer World's Fair 2026 — unofficial import demo

Unofficial demo. This programme was imported from the AI Engineer World's Fair's own public schedule feed to show vibeboard at real conference scale. Not affiliated with, or endorsed by, the organisers.

All sessions
AI-Native EnterprisesSession

Productionizing LLM Gateways: Architecture, Tradeoffs, and Hard Lessons from the Trenches

Kanish Manuja

When
Tuesday, June 302:25 PM – 2:45 PM · 20 min
Where
Leadership 1San Francisco, CA · imported from ai.engineer's public schedule feed
Google Calendar

About this session

As organizations scale their use of large language models, the biggest challenge is no longer prompting, it’s productionizing. This session dives deep into building and operating an LLM gateway that sits between applications and model providers, handling routing, observability, cost control, reliability, and safety at scale. Drawing from real world experience, this talk breaks down the architecture of a production LLM gateway, including model abstraction layers, request orchestration, fallback strategies, caching, rate limiting, and evaluation pipelines. We’ll explore hard tradeoffs such as latency vs. cost, quality vs. determinism, and vendor lock-in vs. flexibility. Attendees will leave with concrete design patterns, failure modes to avoid, and a mental model for turning LLM experiments into resilient, scalable systems.

Speaker

Kanish Manuja
Kanish Manuja

Principal Software Engineer, Twilio Inc.

Kanish Manuja is a principal AI engineer at Twilio, where he leads production LLM gateway and AI platform systems for enterprise-scale AI applications. His work focuses on building reliable, secure, and observable infrastructure for large language model adoption, including multi-tenant gateways, authentication and authorization, guardrails, audit logging, fallback strategies, and production readiness for GenAI workloads. Kanish has worked across AI platform engineering, conversational intelligence, and distributed systems, helping teams move from experimentation to production-grade LLM deployments. He has led efforts around LLM reliability, governance, tenant isolation, provider abstraction, and operational controls for high-scale customer-facing systems. In this session, Kanish will share practical lessons from designing and operating LLM gateway systems in production, including architectural tradeoffs, failure modes, platform boundaries, and what teams should consider before standardizing LLM access across an organization.

More in AI-Native Enterprises