The Art of Building Verifiers for Computer Use Agents
Miguel González Fernández, Corby Rosset
- When
- Thursday, July 211:40 AM – 12:00 PM · 20 min
- Where
- Expo Stage 1 NESan Francisco, CA · imported from ai.engineer's public schedule feed
About this session
Every team building browser agents has the same problem: you can't trust your own evals. Browser tasks are too open-ended for deterministic checks, so teams use LLM verifiers as judges, and the judges are wrong constantly. WebVoyager misses 45% of failures. WebJudge misses 22%. Used as RL reward, you're not training a better agent, you're training a more confident liar. This talk walks through the Universal Verifier, open-sourced with Microsoft Research: false positive rate near zero, Cohen's κ matching human-human agreement. Four design principles, one open benchmark, and an honest account of where auto-research worked and where it plateaued.
Speakers (2)
Tech Lead, Browserbase
Miguel González Fernández is a Tech Lead at Browserbase and co-author of the Microsoft Research/Browserbase Universal Verifier work for computer-use agents.
Senior Researcher, Microsoft Research
Corby Rosset is a Senior Researcher at Microsoft Research studying the intersection of large language models and search/retrieval systems. His work includes conversational search, Bing's People Also Ask feature, and recent verifier research for computer-use agents.
For developers: this programme is open data — JSON, iCal, schedule XML and an MCP endpoint.Show endpointsHide
- JSONEvery published session and speaker, in one request./aie-worldsfair-2026-import/feed.json
- iCalSubscribe in Google, Apple or Outlook Calendar./aie-worldsfair-2026-import/feed.ics
- Schedule XMLfrab / pentabarf — the format conference apps import./aie-worldsfair-2026-import/feed.xml
- MCP + RESTPoint Claude at the programme. OpenAPI 3.1 included./agents
No key, no signup, CORS open. Everything here is generated from the same data the organisers edit.