June 29 – July 2, 2026 · San Francisco, CA · imported from ai.engineer's public schedule feed

AI Engineer World's Fair 2026 — unofficial import demo

Unofficial demo. This programme was imported from the AI Engineer World's Fair's own public schedule feed to show vibeboard at real conference scale. Not affiliated with, or endorsed by, the organisers.

All sessions
Data QualitySession

Data Quality is the Compute Multiplier

Ari Morcos

When
Tuesday, June 3010:45 AM – 11:05 AM · 20 min
Where
Track 9San Francisco, CA · imported from ai.engineer's public schedule feed
Google Calendar

About this session

Better data quality is the highest-leverage and most underinvested part of building a model: it produces a better model for the same compute, whether you're mid-training on an open base or pre-training from scratch.
This session is a practical look at data curation, covering what data quality actually means, the stages of a modern curation pipeline (cleaning, filtering, deduplication, synthetic data generation, algorithmic mixing, and multi-stage composition), and which steps matter most in practice. It draws on DatologyAI's frontier data research and customer results, including Thomson Reuters' mid-training gains on proprietary legal domain data and Arcee's Trinity model reaching the open frontier on public data alone. You'll leave with a concrete sense of where better data quality pays off and how data curation is shaping the future of model training.

Speaker

Ari Morcos
Ari Morcos

Co-founder, CEO, DatologyAI

Ari Morcos is co-founder and CEO of DatologyAI, building a self-service data curation platform for AI teams. Prior to founding Datology, Ari spent five years at FAIR (Meta AI), most recently as a Senior Staff Research Scientist, where his research on data curation and self-supervised learning received Outstanding Paper Awards at NeurIPS 2022 ("Beyond neural scaling laws: beating power law scaling via data pruning") and ICLR 2023. Before Meta, he was a Research Scientist at DeepMind, applying tools from neuroscience to understand generalization, representation learning, and the dynamics of training in deep networks. He holds a PhD in Neuroscience from Harvard and a BS in Neuroscience from UC San Diego.

More in Data Quality