Model Whisperers: How Evals and Prompts Shape Agent Behavior
Chris Souza, Preetika Bhateja, Daniel Bump
- When
- Wednesday, July 11:30 PM – 1:50 PM · 20 min
- Where
- Track 5San Francisco, CA · imported from ai.engineer's public schedule feed
About this session
Getting an AI agent to behave the way you want isn’t just about writing better prompts. In real systems, behavior emerges from a loop: prompts->evals->iteration->feedback. Small changes in any part of that loop can completely change outcomes. We saw this while building a seed asset agent - a system that turns messy, real-world advertising creatives (low quality images, cluttered visuals, heavy text overlays) into clean, reusable assets for downstream Gen AI tools. The agent acts like an editor, simplifying visuals, removing unnecessary elements, and isolating core content so that additional context (like text or CTAs) can be added back in a more controlled, brand-safe way. But the real challenge wasn’t just building the agent - it was making it reliable. And prompting alone wasn’t enough. What actually moved the system forward was how we defined success—and how we used evals to reinforce it. Over time, evals stopped being just a way to measure quality. They became part of how the agent learned what “good” looks like. In this talk, we’ll cover: Why prompting alone doesn’t give you stable agent behavior How evals act like feedback signals, not just scorecards How we built evals sets that reflect the real-world Using agent trace logs to understand why things fail (not just that they fail) How to iterate without breaking things you already fixed By the end, you’ll have a set of patterns you can apply to any system dealing with messy/continuously changing data and how to tweak your prompt and evals to accommodate such changes.
Speakers (3)
Product Manager, Google
Product Manager at Google/YouTube working on ads, evals, agents, llm-as-judge systems. Before PM, data engineer at google cloud
Engineer, Google
Engineer at Google. Focus area: image/video generation and computer vision.
More in Evals
- Vending-Bench: Long-Horizon Agent Evals for a Simulated Vending BusinessWednesday, July 1 · 10:45 AM – 11:05 AM · Track 5
- From Signal to PR: Anatomy of a Self-Improving AgentWednesday, July 1 · 11:10 AM – 11:30 AM · Track 5
- Building Closed-Loop Evals for a Multimodal Agent at Uber ScaleWednesday, July 1 · 11:40 AM – 12:00 PM · Track 5
- From Agent Traces to Agent Simulations: The next era of agent evaluationWednesday, July 1 · 12:05 PM – 12:25 PM · Track 5
- Evaling Video SlopWednesday, July 1 · 1:55 PM – 2:15 PM · Track 5
For developers: this programme is open data — JSON, iCal, schedule XML and an MCP endpoint.Show endpointsHide
- JSONEvery published session and speaker, in one request./aie-worldsfair-2026-import/feed.json
- iCalSubscribe in Google, Apple or Outlook Calendar./aie-worldsfair-2026-import/feed.ics
- Schedule XMLfrab / pentabarf — the format conference apps import./aie-worldsfair-2026-import/feed.xml
- MCP + RESTPoint Claude at the programme. OpenAPI 3.1 included./agents
No key, no signup, CORS open. Everything here is generated from the same data the organisers edit.