June 29 – July 2, 2026 · San Francisco, CA · imported from ai.engineer's public schedule feed

AI Engineer World's Fair 2026 — unofficial import demo

Unofficial demo. This programme was imported from the AI Engineer World's Fair's own public schedule feed to show vibeboard at real conference scale. Not affiliated with, or endorsed by, the organisers.

All sessions
Computer UseSession

Computer Use at the Edge of the Statistical Precipice

Pierluca D'Oro

When
Wednesday, July 111:10 AM – 11:30 AM · 20 min
Where
Track 7San Francisco, CA · imported from ai.engineer's public schedule feed
Google Calendar

About this session

Evaluating Computer Use Agents (CUAs) on interactive environments is fraught with methodological pitfalls that the field has yet to systematically address. We show that a 1MB replay script that blindly executes a recorded action sequence without ever observing the screen outperforms frontier models on prominent static benchmarks, and prove that its expected success rate is exactly equal to the source agent's pass@k in deterministic environments. We trace this and other failures to two root causes: non-principled environment design (static, unsandboxed, or unreliably verified environments) and non-principled evaluation methodology (naive aggregation and misuse of pass@k for stateful UI interactions). To address the first, we propose PRISM, five design principles for CUA environments and instantiate them in DigiWorld, a benchmark of 15 realistic sandboxed mobile applications able to evaluate agents in over 3.2 million verified unique configurations. To address the second, we develop an aggregation framework that correctly accounts for the nested structure of CUA benchmarks. All together, we show that principled environment design and rigorous evaluation methodology are not optional refinements but prerequisites for meaningful CUA research.

Speaker

Pierluca D'Oro
Pierluca D'Oro

Founder, Programma Labs

Pierluca D’Oro is founder at a stealth startup revolutionizing how humans interact with AI-generated software. At Mila, he pioneered two early ideas that now sit at the center of agent development: making reinforcement learning scale through simple recipes, and using LLMs as feedback systems to train agents. At Meta Superintelligence Labs, he worked on frontier model development and led environment generation for mobile computer use agents.

More in Computer Use