Computer Use at the Edge of the Statistical Precipice
Pierluca D'Oro
- When
- Wednesday, July 111:10 AM – 11:30 AM · 20 min
- Where
- Track 7San Francisco, CA · imported from ai.engineer's public schedule feed
About this session
Evaluating Computer Use Agents (CUAs) on interactive environments is fraught with methodological pitfalls that the field has yet to systematically address. We show that a 1MB replay script that blindly executes a recorded action sequence without ever observing the screen outperforms frontier models on prominent static benchmarks, and prove that its expected success rate is exactly equal to the source agent's pass@k in deterministic environments. We trace this and other failures to two root causes: non-principled environment design (static, unsandboxed, or unreliably verified environments) and non-principled evaluation methodology (naive aggregation and misuse of pass@k for stateful UI interactions). To address the first, we propose PRISM, five design principles for CUA environments and instantiate them in DigiWorld, a benchmark of 15 realistic sandboxed mobile applications able to evaluate agents in over 3.2 million verified unique configurations. To address the second, we develop an aggregation framework that correctly accounts for the nested structure of CUA benchmarks. All together, we show that principled environment design and rigorous evaluation methodology are not optional refinements but prerequisites for meaningful CUA research.
Speaker
Founder, Programma Labs
Pierluca D’Oro is founder at a stealth startup revolutionizing how humans interact with AI-generated software. At Mila, he pioneered two early ideas that now sit at the center of agent development: making reinforcement learning scale through simple recipes, and using LLMs as feedback systems to train agents. At Meta Superintelligence Labs, he worked on frontier model development and led environment generation for mobile computer use agents.
More in Computer Use
- Computer-use models will agentify the web, not APIsWednesday, July 1 · 10:45 AM – 11:05 AM · Track 7
- Bringing agents onto the world wide webWednesday, July 1 · 11:40 AM – 12:00 PM · Track 7
- The Dark Arts of Web Automation: Teaching Agents to Use Websites Like HumansWednesday, July 1 · 12:05 PM – 12:25 PM · Track 7
- From RL to IRLWednesday, July 1 · 1:30 PM – 1:50 PM · Track 7
- The Rise of CaaS: Context-as-a-Service for Agentic AIWednesday, July 1 · 1:55 PM – 2:15 PM · Track 7
For developers: this programme is open data — JSON, iCal, schedule XML and an MCP endpoint.Show endpointsHide
- JSONEvery published session and speaker, in one request./aie-worldsfair-2026-import/feed.json
- iCalSubscribe in Google, Apple or Outlook Calendar./aie-worldsfair-2026-import/feed.ics
- Schedule XMLfrab / pentabarf — the format conference apps import./aie-worldsfair-2026-import/feed.xml
- MCP + RESTPoint Claude at the programme. OpenAPI 3.1 included./agents
No key, no signup, CORS open. Everything here is generated from the same data the organisers edit.