Why building building agent quality platforms is hard.
Hossein Niazmandi
- When
- Tuesday, June 3012:05 PM – 12:25 PM · 20 min
- Where
- Expo Stage 2 NWSan Francisco, CA · imported from ai.engineer's public schedule feed
About this session
An eval platform is not just a test runner. You are building shared definitions of good, reliable data pipelines, labeling workflows, versioning, and trust in results across many teams and model changes. This session breaks down the hidden complexity, the common failure modes, and the design principles that make evals credible and usable in day-to-day engineering.
Speaker
Solutions, Braintrust
Hossein Niazmandi works on Solutions at Braintrust, with prior experience at Databricks and Salesforce. His WF26 session focuses on why building agent quality platforms is hard.
For developers: this programme is open data — JSON, iCal, schedule XML and an MCP endpoint.Show endpointsHide
- JSONEvery published session and speaker, in one request./aie-worldsfair-2026-import/feed.json
- iCalSubscribe in Google, Apple or Outlook Calendar./aie-worldsfair-2026-import/feed.ics
- Schedule XMLfrab / pentabarf — the format conference apps import./aie-worldsfair-2026-import/feed.xml
- MCP + RESTPoint Claude at the programme. OpenAPI 3.1 included./agents
No key, no signup, CORS open. Everything here is generated from the same data the organisers edit.