June 29 – July 2, 2026 · San Francisco, CA · imported from ai.engineer's public schedule feed

AI Engineer World's Fair 2026 — unofficial import demo

Unofficial demo. This programme was imported from the AI Engineer World's Fair's own public schedule feed to show vibeboard at real conference scale. Not affiliated with, or endorsed by, the organisers.

All sessions
Posttraining & MidtrainingSession

Benchmarks: The Good, the Bad, and the Ugly

Ali Khial

When
Wednesday, July 13:20 PM – 3:40 PM · 20 min
Where
Track 9San Francisco, CA · imported from ai.engineer's public schedule feed
Google Calendar

About this session

We’ll explore the good, the bad, and the ugly of AI benchmarks: where they provide useful signal, where they create false confidence, and where data quality issues like contamination, label noise, narrow task design, and leaderboard gaming can mislead teams. The goal is not to dismiss benchmarks, but to use them better: as one part of a disciplined evaluation practice that connects model performance to real-world reliability.

Speaker

Ali Khial

Head of AI/ML, G2i

Ali Khial is an engineering leader focused on building AI-native systems that work beyond the demo stage. He currently leads AI/ML at G2i, where he works across frontier AI evaluation, software engineering benchmarks, agentic workflows, and human-data quality systems. His current work centers on the gap between impressive AI prototypes and reliable production systems. He is especially interested in AI evaluation, data quality, tool-using applications, and the engineering practices needed to ship model-powered products in real-world environments.

More in Posttraining & Midtraining