June 29 – July 2, 2026 · San Francisco, CA · imported from ai.engineer's public schedule feed

AI Engineer World's Fair 2026 — unofficial import demo

Unofficial demo. This programme was imported from the AI Engineer World's Fair's own public schedule feed to show vibeboard at real conference scale. Not affiliated with, or endorsed by, the organisers.

All sessions
Session

Inference performance as a competitive advantage

Alex Campos, Yunmo Koo

When
Wednesday, July 12:50 PM – 3:10 PM · 20 min
Where
Expo Stage 1 NESan Francisco, CA · imported from ai.engineer's public schedule feed
Google Calendar

About this session

Most AI teams focus on model quality, but production success often comes down to inference performance. In this session, FriendliAI will explore the optimization techniques behind high-performance LLM serving, including continuous batching, speculative decoding, smart caching, and efficient GPU utilization. Learn how leading AI teams reduce infrastructure costs, improve latency, and scale inference workloads without sacrificing performance. We'll share practical insights and deployment strategies that separate experimental AI projects from production-grade systems.Whether you're an ML engineer, platform engineer, MLOps practitioner, or technical founder, you'll leave with a better understanding of how inference optimization can become a competitive advantage for your AI applications.

Speakers (2)

Alex Campos

Director of Sales Partnerships, FriendliAI

Alex Campos leads sales partnerships at FriendliAI, a frontier AI inference cloud focused on high-performance open-weight model serving and production inference optimization.

Yunmo Koo

Founding Engineer, FriendliAI

Yunmo Koo is a founding engineer at FriendliAI focused on LLM inference optimization, distributed training, multi-cloud systems, and LLMOps. He builds production ML infrastructure for lower latency, better reliability, and improved cost efficiency.