June 29 – July 2, 2026 · San Francisco, CA · imported from ai.engineer's public schedule feed

AI Engineer World's Fair 2026 — unofficial import demo

Unofficial demo. This programme was imported from the AI Engineer World's Fair's own public schedule feed to show vibeboard at real conference scale. Not affiliated with, or endorsed by, the organisers.

All sessions
Data QualitySession

The Messy Reality of Scale: Synthetic Data and Pre-Training at Poolside

Robert McHardy, Marah Abdin

When
Tuesday, June 3011:10 AM – 11:30 AM · 20 min
Where
Track 9San Francisco, CA · imported from ai.engineer's public schedule feed
Google Calendar

About this session

TBD — focus on data quality considerations for LLM pretraining and code generation.

Speakers (2)

Robert McHardy
Robert McHardy

Pre-training Lead, poolside

Team and tech lead for pre-training at poolside, where he trains large language models for code. Recently led the pre-training of Laguna XS.2 and M.1, poolside's first two public open-weight models. Before that, Robert worked as a Senior Researcher at AssemblyAI where he trained multilingual speech models, and previously built AI for cancer and infectious-disease research at InstaDeep and BioNTech's joint lab. MSc in Machine Learning from UCL.

Marah Abdin
Marah Abdin

Team Lead - Synthetic Data, poolside

Synthetic data Lead at Poolside, building the *Laguna* models (pre-training/post-training). Previously at Microsoft AI/Research, building the *Phi* models.

More in Data Quality