Inference performance as a competitive advantage
Alex Campos, Yunmo Koo
- When
- Wednesday, July 12:50 PM – 3:10 PM · 20 min
- Where
- Expo Stage 1 NESan Francisco, CA · imported from ai.engineer's public schedule feed
About this session
Most AI teams focus on model quality, but production success often comes down to inference performance. In this session, FriendliAI will explore the optimization techniques behind high-performance LLM serving, including continuous batching, speculative decoding, smart caching, and efficient GPU utilization. Learn how leading AI teams reduce infrastructure costs, improve latency, and scale inference workloads without sacrificing performance. We'll share practical insights and deployment strategies that separate experimental AI projects from production-grade systems.Whether you're an ML engineer, platform engineer, MLOps practitioner, or technical founder, you'll leave with a better understanding of how inference optimization can become a competitive advantage for your AI applications.
Speakers (2)
Director of Sales Partnerships, FriendliAI
Alex Campos leads sales partnerships at FriendliAI, a frontier AI inference cloud focused on high-performance open-weight model serving and production inference optimization.
Founding Engineer, FriendliAI
Yunmo Koo is a founding engineer at FriendliAI focused on LLM inference optimization, distributed training, multi-cloud systems, and LLMOps. He builds production ML infrastructure for lower latency, better reliability, and improved cost efficiency.
For developers: this programme is open data — JSON, iCal, schedule XML and an MCP endpoint.Show endpointsHide
- JSONEvery published session and speaker, in one request./aie-worldsfair-2026-import/feed.json
- iCalSubscribe in Google, Apple or Outlook Calendar./aie-worldsfair-2026-import/feed.ics
- Schedule XMLfrab / pentabarf — the format conference apps import./aie-worldsfair-2026-import/feed.xml
- MCP + RESTPoint Claude at the programme. OpenAPI 3.1 included./agents
No key, no signup, CORS open. Everything here is generated from the same data the organisers edit.