Data Quality is the Compute Multiplier
Ari Morcos
- When
- Tuesday, June 3010:45 AM – 11:05 AM · 20 min
- Where
- Track 9San Francisco, CA · imported from ai.engineer's public schedule feed
About this session
Better data quality is the highest-leverage and most underinvested part of building a model: it produces a better model for the same compute, whether you're mid-training on an open base or pre-training from scratch.
This session is a practical look at data curation, covering what data quality actually means, the stages of a modern curation pipeline (cleaning, filtering, deduplication, synthetic data generation, algorithmic mixing, and multi-stage composition), and which steps matter most in practice. It draws on DatologyAI's frontier data research and customer results, including Thomson Reuters' mid-training gains on proprietary legal domain data and Arcee's Trinity model reaching the open frontier on public data alone. You'll leave with a concrete sense of where better data quality pays off and how data curation is shaping the future of model training.
Speaker
Co-founder, CEO, DatologyAI
Ari Morcos is co-founder and CEO of DatologyAI, building a self-service data curation platform for AI teams. Prior to founding Datology, Ari spent five years at FAIR (Meta AI), most recently as a Senior Staff Research Scientist, where his research on data curation and self-supervised learning received Outstanding Paper Awards at NeurIPS 2022 ("Beyond neural scaling laws: beating power law scaling via data pruning") and ICLR 2023. Before Meta, he was a Research Scientist at DeepMind, applying tools from neuroscience to understand generalization, representation learning, and the dynamics of training in deep networks. He holds a PhD in Neuroscience from Harvard and a BS in Neuroscience from UC San Diego.
More in Data Quality
- The Messy Reality of Scale: Synthetic Data and Pre-Training at PoolsideTuesday, June 30 · 11:10 AM – 11:30 AM · Track 9
- Rethinking Environments for Long Horizon WorkTuesday, June 30 · 11:40 AM – 12:00 PM · Track 9
- The Base Model is DeadTuesday, June 30 · 1:30 PM – 1:50 PM · Track 9
- Ending AI SlopTuesday, June 30 · 1:55 PM – 2:15 PM · Track 9
- Scaling to Long-Horizons: Algorithms, Environments, ComputeTuesday, June 30 · 2:25 PM – 2:45 PM · Track 9
For developers: this programme is open data — JSON, iCal, schedule XML and an MCP endpoint.Show endpointsHide
- JSONEvery published session and speaker, in one request./aie-worldsfair-2026-import/feed.json
- iCalSubscribe in Google, Apple or Outlook Calendar./aie-worldsfair-2026-import/feed.ics
- Schedule XMLfrab / pentabarf — the format conference apps import./aie-worldsfair-2026-import/feed.xml
- MCP + RESTPoint Claude at the programme. OpenAPI 3.1 included./agents
No key, no signup, CORS open. Everything here is generated from the same data the organisers edit.