Building self-learning loops for your agent
Fuad Ali
- When
- Monday, June 2911:05 AM – 12:05 PM · 60 min
- Where
- Track 1San Francisco, CA · imported from ai.engineer's public schedule feed
About this session
Building an AI demo is easy. Knowing whether it actually works — and keeping it working in production — is the hard part. Most teams ship agents on vibes: they try a few prompts, the output looks good, and they push to production with no real way to measure quality or catch regressions.
This hands-on workshop walks through the full lifecycle of shipping a real AI agent, using a working financial-analyst agent built on the Claude Agent SDK as the running example. You'll instrument it with tracing, do structured error analysis on its actual outputs, and build a layered evaluation suite — from cheap deterministic code checks to LLM-as-a-judge evaluators with custom rubrics. We'll cover the parts most tutorials skip: why agents fail in ways single LLM calls don't, the eval anti-patterns that quietly mislead you, and how to know whether you can even trust your judge (meta-evaluation). Finally, we'll close the loop: turning eval results into datasets and experiments, running evals online against production traffic, wiring them to monitors and alerts, and feeding failure explanations back to a coding agent to actually fix the underlying problems.
You'll leave with a runnable notebook and a repeatable, evaluation-driven workflow you can apply to your own agents the next day.
Speaker
Senior Product Manager, Arize AI
Fuad Ali is a Senior Product Manager at Arize AI focused on ML observability and reliable AI systems. His background spans engineering and product work at SpaceX, Tesla, Twitter and Federato, and he co-hosts The Next Iteration podcast.
More in Posttraining & Midtraining
- Bugcrowd posttraining talkTuesday, June 30 · 12:05 PM – 12:25 PM · Track 9
- What's next after RLHF?Wednesday, July 1 · 10:45 AM – 11:05 AM · Track 9
- State of DataWednesday, July 1 · 11:10 AM – 11:30 AM · Track 9
- Learning on the job: the future of post-trainingWednesday, July 1 · 12:05 PM – 12:25 PM · Track 9
- Reinforcement Learning without Verifiable RewardsWednesday, July 1 · 1:30 PM – 1:50 PM · Track 9
For developers: this programme is open data — JSON, iCal, schedule XML and an MCP endpoint.Show endpointsHide
- JSONEvery published session and speaker, in one request./aie-worldsfair-2026-import/feed.json
- iCalSubscribe in Google, Apple or Outlook Calendar./aie-worldsfair-2026-import/feed.ics
- Schedule XMLfrab / pentabarf — the format conference apps import./aie-worldsfair-2026-import/feed.xml
- MCP + RESTPoint Claude at the programme. OpenAPI 3.1 included./agents
No key, no signup, CORS open. Everything here is generated from the same data the organisers edit.