The Best Models Still Reason Like Toddlers
Andrew Dai
- When
- Tuesday, June 301:55 PM – 2:15 PM · 20 min
- Where
- Track 2San Francisco, CA · imported from ai.engineer's public schedule feed
About this session
Frontier AI models score 80–90% on standard benchmarks like RKGI, yet when tested on visual tasks any 3-year-old handles effortlessly (like counting objects in an image), those same models fall to pieces. I watched this gap widen firsthand during my 14 years at Google Brain and DeepMind, where I co-led development on GLaM, PaLM 2, and Gemini. The problem is that most models hit high RKGI scores not through genuine visual understanding, but by coding – a workaround that scores well and reveals little. Strip that away and you're left with systems that struggle to solve a simple crossword puzzle, identify what's the same or different across two images, or navigate a basic 3D view. These tasks are essential to achieve human-level reasoning capability. And the current benchmark ecosystem wasn’t built to evaluate for it, leaving us with top scoring models that can’t even follow along with Count Von Count. In this talk I'll dig into why the current eval landscape systematically overstates capability, the structural reasons it does so, and how we got here from the viewpoint of someone who was inside a leading frontier lab. I'll close with what I believe a more rigorous, consensus-driven eval framework needs to look like, and why the field needs to build one before the next generation of visual systems ships into the real world. Fixing visual reasoning starts with fixing how we measure it. For engineers building on top of these models today, whether that's document understanding, robotic perception, medical imaging, or any system where visual perception context matters, the cost of getting this wrong is already showing up in production.
Speaker
Co-founder and CEO, Elorian
Andrew Dai spent 12 years as a Research Scientist at Google Brain and DeepMind. He wrote the 2015 paper that OpenAI later cited as the original recipe for ChatGPT, was a core Lead on Gemini, GLaM, and PaLM 2, and his published research has accumulated over 67,000 citations. Now, he leads Elorian, a company building AI systems that understand the visual medium and apply reasoning the way humans do. Elorian recently launched with $55M at a $300M valuation, backed by Menlo Ventures, Altimeter, Striker Venture Partners, NVIDIA and Jeff Dean.
More in Vision & OCR
- The State of VisionTuesday, June 30 · 10:45 AM – 11:05 AM · Track 2
- Building the Document Context Layer for AI AgentsTuesday, June 30 · 11:10 AM – 11:30 AM · Track 2
- Skill issue: stop deploying vision language models, use them with Skills to build e2e vision apps on edgeTuesday, June 30 · 11:40 AM – 12:00 PM · Track 2
- Modality Misalignment and Originality Attribution in Short-Form Video: A Multi-Agent Approach at Platform ScaleTuesday, June 30 · 12:05 PM – 12:25 PM · Track 2
- From Ingestion to Agents: How Leading AI Teams Build on Document IntelligenceTuesday, June 30 · 1:30 PM – 1:50 PM · Track 2
For developers: this programme is open data — JSON, iCal, schedule XML and an MCP endpoint.Show endpointsHide
- JSONEvery published session and speaker, in one request./aie-worldsfair-2026-import/feed.json
- iCalSubscribe in Google, Apple or Outlook Calendar./aie-worldsfair-2026-import/feed.ics
- Schedule XMLfrab / pentabarf — the format conference apps import./aie-worldsfair-2026-import/feed.xml
- MCP + RESTPoint Claude at the programme. OpenAPI 3.1 included./agents
No key, no signup, CORS open. Everything here is generated from the same data the organisers edit.