All sessions
AI EngineeringLightning Talk (10 min)
Reading Eval Traces Like a Detective
Elena Vasquez
- When
- Tuesday, September 151:35 PM – 1:45 PM · 10 min
- Where
- Room 2AFort Mason Center, San Francisco
About this session
A short, practical walkthrough of how to read a failing evaluation trace: what to look at first, which signals lie, and how to turn one bad output into a reproducible test case.
Speaker
Elena Vasquez
ML Engineer, Braid Research
Elena works on retrieval and evaluation at Braid. She published an open benchmark suite after getting tired of numbers that never reproduced on her own hardware.
More in AI Engineering
- Test Data Is a Product: Fixtures for Nondeterministic SystemsTime to be announced
- Context Engineering in Production: What Actually Survives Contact With UsersTuesday, September 15 · 9:45 AM – 10:30 AM · Room 2A
- Your AI Pair Programmer Is Lying to You: Verification Patterns That ScaleTuesday, September 15 · 10:30 AM – 11:00 AM · Main Stage
- Multi-Agent Systems Are Just Distributed Systems (And That's Good News)Tuesday, September 15 · 11:15 AM – 11:45 AM · Main Stage
- Prompt Injection Is a Solved Problem (If You Do These Five Things)Tuesday, September 15 · 1:15 PM – 1:45 PM · Main Stage
For developers: this programme is open data — JSON, iCal, schedule XML and an MCP endpoint.Show endpointsHide
- JSONEvery published session and speaker, in one request./ai-builders-summit-2026/feed.json
- iCalSubscribe in Google, Apple or Outlook Calendar./ai-builders-summit-2026/feed.ics
- Schedule XMLfrab / pentabarf — the format conference apps import./ai-builders-summit-2026/feed.xml
- MCP + RESTPoint Claude at the programme. OpenAPI 3.1 included./agents
No key, no signup, CORS open. Everything here is generated from the same data the organisers edit.