From Agent Traces to Agent Simulations: The next era of agent evaluation
Rustem Feyzkhanov
- When
- Wednesday, July 112:05 PM – 12:25 PM · 20 min
- Where
- Track 5San Francisco, CA · imported from ai.engineer's public schedule feed
About this session
Agent evaluation is moving beyond reviewing static traces after the fact. This talk explores how executable simulation environments let teams repeatedly test agents across realistic tasks, compare models and harnesses, and uncover failure modes that trace review alone misses. Drawing from Snorkel's experience building simulation datasets at scale for major labs and contributions to projects like Agents' Last Exam and Terminal-Bench, we'll cover concrete engineering patterns for building these environments: defining clear specs and requirements, implementing evaluators for simulation environments and tasks themselves, keeping environments decoupled from any single agent or model, and designing verifiers that evaluate both final outputs and agent traces. Attendees will leave with a practical mental model for creating environments that are lightweight enough to run at scale, but realistic enough to mock production systems such as databases, APIs, and tools in ways that meaningfully challenge agents.
Speaker
Senior Engineering Manager - AI Platform, Snorkel AI
Rustem Feyzkhanov is a Senior Engineering Manager on the AI Platform Engineering team at Snorkel AI, where he leads work on infrastructure and platform systems for building expert-authored datasets, simulation environments, and evaluation pipelines for frontier AI models and production agents. His work focuses on scalable agent evaluation, secure sandboxed execution, benchmark quality, and the systems needed to run large volumes of agent simulations reliably. Before Snorkel, Rustem was an ML Engineering Manager at Instrumental, applying AI to manufacturing, and an engineer at Astro Digital, building AI systems for satellite imagery. He is passionate about AI agents, evaluation infrastructure, serverless computing, and practical machine learning systems. Rustem is the author of the course and book Serverless Deep Learning with TensorFlow and AWS Lambda and Practical Deep Learning on the Cloud, and he is the main contributor to the open-source lambda-packs repository for serverless Python packages.
More in Evals
- Vending-Bench: Long-Horizon Agent Evals for a Simulated Vending BusinessWednesday, July 1 · 10:45 AM – 11:05 AM · Track 5
- From Signal to PR: Anatomy of a Self-Improving AgentWednesday, July 1 · 11:10 AM – 11:30 AM · Track 5
- Building Closed-Loop Evals for a Multimodal Agent at Uber ScaleWednesday, July 1 · 11:40 AM – 12:00 PM · Track 5
- Model Whisperers: How Evals and Prompts Shape Agent BehaviorWednesday, July 1 · 1:30 PM – 1:50 PM · Track 5
- Evaling Video SlopWednesday, July 1 · 1:55 PM – 2:15 PM · Track 5
For developers: this programme is open data — JSON, iCal, schedule XML and an MCP endpoint.Show endpointsHide
- JSONEvery published session and speaker, in one request./aie-worldsfair-2026-import/feed.json
- iCalSubscribe in Google, Apple or Outlook Calendar./aie-worldsfair-2026-import/feed.ics
- Schedule XMLfrab / pentabarf — the format conference apps import./aie-worldsfair-2026-import/feed.xml
- MCP + RESTPoint Claude at the programme. OpenAPI 3.1 included./agents
No key, no signup, CORS open. Everything here is generated from the same data the organisers edit.