AI Evals Platform for Cross-Functional Teams at Scale
Nachiket Paranjape, Swaroop Chitlur Haridas
- When
- Tuesday, June 301:55 PM – 2:15 PM · 20 min
- Where
- Leadership 1San Francisco, CA · imported from ai.engineer's public schedule feed
About this session
DoorDash's Evals Platform is designed for more than just engineers. It brings human review, automated judges, and online experimentation into a single calibration loop so engineering, product managers, and strategy and operations teams can all contribute to improving AI quality. Engineers can instrument, trace, and evaluate agent behavior, while cross-functional teams can review outputs, curate trusted examples, and provide structured feedback that improves how automated judges behave over time. By combining experimentation, fully customized annotation workflows, calibration, and analytics in one system, the platform turns AI quality from a fragmented technical exercise into a shared operating model for continuously improving agent performance and making rollout decisions with confidence. While vendor platforms offer pieces of this workflow, we needed something broader: a unified system that lets engineers, product managers, and Strategy & Ops all participate directly in improving AI quality. Our goal is not just to run evals, but to enable cross-functional teams to review outputs, calibrate judges, run experiments, and make rollout decisions without being blocked on engineering. That requirement, along with tighter integration into our internal workflows and operating model, is why we are building this platform in-house.
Speakers (2)
Software Engineer, DoorDash
Software Engineer at DoorDash's AI Platform Team. Currently leading the AI Evals initiative. Previously Engineering Lead at Galileo AI (acquired by Cisco).
DoorDash
More in AI-Native Enterprises
- Building the engine while flying the plane — launching the Figma MCP serverTuesday, June 30 · 11:10 AM – 11:30 AM · Leadership 1
- Agentic SDLC at Uber: Building Blocks for Uber's Software FactoryTuesday, June 30 · 11:40 AM – 12:00 PM · Leadership 1
- Scaling Code Quality: Building uReview, Uber’s Multi-Agent Code Review EngineTuesday, June 30 · 12:05 PM – 12:25 PM · Leadership 1
- Productionizing LLM Gateways: Architecture, Tradeoffs, and Hard Lessons from the TrenchesTuesday, June 30 · 2:25 PM – 2:45 PM · Leadership 1
- From AI-Assisted to AI-Native: Building a Frontier Development TeamTuesday, June 30 · 2:50 PM – 3:10 PM · Leadership 1
For developers: this programme is open data — JSON, iCal, schedule XML and an MCP endpoint.Show endpointsHide
- JSONEvery published session and speaker, in one request./aie-worldsfair-2026-import/feed.json
- iCalSubscribe in Google, Apple or Outlook Calendar./aie-worldsfair-2026-import/feed.ics
- Schedule XMLfrab / pentabarf — the format conference apps import./aie-worldsfair-2026-import/feed.xml
- MCP + RESTPoint Claude at the programme. OpenAPI 3.1 included./agents
No key, no signup, CORS open. Everything here is generated from the same data the organisers edit.