Evals in AI: A Deep Dive
Tejas Kumar
- When
- Monday, June 2912:10 PM – 1:10 PM · 60 min
- Where
- Track 1San Francisco, CA · imported from ai.engineer's public schedule feed
About this session
“Our evals pass and our velocity is up, so it works.” It’s the most reassuring sentence in AI engineering and also the most dangerous. Teams are shipping more code than ever while incidents per PR and change-failure rates climb, and the instruments meant to catch this are quietly broken. This talk takes apart both halves of that false comfort. First, why velocity lies: the same AI-driven throughput that lights up your dashboard is what’s eroding quality underneath it. Then we explore four ways offline evals deceive you: LLM-as-judge bias (your grader rewards confident, wordy, wrong answers over terse correct ones), staleness, distribution shift between your golden set and real traffic, and single-score evals that hide which step of an agent actually failed. The centerpiece is a live demo. We’ll wire up an LLM judge on stage and watch it crown a confident, friendly, factually wrong answer. Then we’ll fix it live on stage with a three-line rubric change. Same model, different instrument. From there we’ll build up what to measure instead: traces and spans, production observability, probe-based evaluation, error budgets, and quality leading indicators that sit beside every velocity number. Attendees will leave with a five-line checklist they can apply Monday. No prior eval tooling required. If you’ve ever shipped something agentic and had a nagging feeling the dashboards were too kind, this is for you.
Speaker
AI Engineer, IBM
Tejas Kumar is an international keynote speaker, best selling author, and host of the developer-loved ConTejas Code podcast with an engineering background spanning 25 years, from design to frontend to backend to devops. Today, Tejas shares talks at large with developer communities worldwide, equipping them to do their best work.
More in Workshops Day 1
- Cooking with CodexMonday, June 29 · 9:00 AM – 11:00 AM · Track 3
- The best SDLC is the one you build yourself: Why orchestration changes everythingMonday, June 29 · 9:00 AM – 11:00 AM · Track 4
- AI Security Engineer Foundations + CertificateMonday, June 29 · 9:00 AM – 11:00 AM · Track 5
- Total Recall: Agent Memory and Harness EngineeringMonday, June 29 · 9:00 AM – 11:00 AM · Track 6
- Open-Source Inference Engineering for the Agentic EraMonday, June 29 · 9:00 AM – 11:00 AM · Track 8
For developers: this programme is open data — JSON, iCal, schedule XML and an MCP endpoint.Show endpointsHide
- JSONEvery published session and speaker, in one request./aie-worldsfair-2026-import/feed.json
- iCalSubscribe in Google, Apple or Outlook Calendar./aie-worldsfair-2026-import/feed.ics
- Schedule XMLfrab / pentabarf — the format conference apps import./aie-worldsfair-2026-import/feed.xml
- MCP + RESTPoint Claude at the programme. OpenAPI 3.1 included./agents
No key, no signup, CORS open. Everything here is generated from the same data the organisers edit.