Building Closed-Loop Evals for a Multimodal Agent at Uber Scale
Soumya Gupta, Jai Chopra
- When
- Wednesday, July 111:40 AM – 12:00 PM · 20 min
- Where
- Track 5San Francisco, CA · imported from ai.engineer's public schedule feed
About this session
This talk covers how we designed evals for Uber's food enhancement agent—which edits food photography to better present dishes for smaller, independent Uber Eats merchants—along with the pitfalls and lessons learned along the way.
The problem is uniquely hard: we must stay faithful to the original dish, preserve each merchant's brand and packaging, and avoid homogenizing the marketplace—all without an existing playbook for multimodal evals in a narrow domain. We'll dig into what we learned navigating reward hacking, where the agent figured out how to game the eval loop, and how we built a closed feedback loop incorporating offline and online signals for continuous improvement—all while balancing creativity against rigid safety guardrails at scale.
If you're an ML or applied AI practitioner working on multimodal systems, agentic pipelines, or eval design—especially building generative features under tight safety or quality constraints—you'll walk away with practical strategies for designing multimodal evals in a narrow domain, recognizing and countering reward hacking, and building offline/online feedback loops that keep a generative agent improving in production.
Speakers (2)
ML Engineer, Uber
Soumya Gupta is a Tech Lead and Applied AI Engineer at Uber, where she architects and scales production-grade Generative AI and Computer Vision solutions. Her work focuses on deploying agentic orchestration, multimodal modeling, and core machine learning primitives at global scale.
Product Manager, Uber
Product Lead in the Applied AI team at Uber. Previously worked at Cruise and various startups.
More in Evals
- Vending-Bench: Long-Horizon Agent Evals for a Simulated Vending BusinessWednesday, July 1 · 10:45 AM – 11:05 AM · Track 5
- From Signal to PR: Anatomy of a Self-Improving AgentWednesday, July 1 · 11:10 AM – 11:30 AM · Track 5
- From Agent Traces to Agent Simulations: The next era of agent evaluationWednesday, July 1 · 12:05 PM – 12:25 PM · Track 5
- Model Whisperers: How Evals and Prompts Shape Agent BehaviorWednesday, July 1 · 1:30 PM – 1:50 PM · Track 5
- Evaling Video SlopWednesday, July 1 · 1:55 PM – 2:15 PM · Track 5
For developers: this programme is open data — JSON, iCal, schedule XML and an MCP endpoint.Show endpointsHide
- JSONEvery published session and speaker, in one request./aie-worldsfair-2026-import/feed.json
- iCalSubscribe in Google, Apple or Outlook Calendar./aie-worldsfair-2026-import/feed.ics
- Schedule XMLfrab / pentabarf — the format conference apps import./aie-worldsfair-2026-import/feed.xml
- MCP + RESTPoint Claude at the programme. OpenAPI 3.1 included./agents
No key, no signup, CORS open. Everything here is generated from the same data the organisers edit.