The Messy Reality of Scale: Synthetic Data and Pre-Training at Poolside
Robert McHardy, Marah Abdin
- When
- Tuesday, June 3011:10 AM – 11:30 AM · 20 min
- Where
- Track 9San Francisco, CA · imported from ai.engineer's public schedule feed
About this session
TBD — focus on data quality considerations for LLM pretraining and code generation.
Speakers (2)
Pre-training Lead, poolside
Team and tech lead for pre-training at poolside, where he trains large language models for code. Recently led the pre-training of Laguna XS.2 and M.1, poolside's first two public open-weight models. Before that, Robert worked as a Senior Researcher at AssemblyAI where he trained multilingual speech models, and previously built AI for cancer and infectious-disease research at InstaDeep and BioNTech's joint lab. MSc in Machine Learning from UCL.
Team Lead - Synthetic Data, poolside
Synthetic data Lead at Poolside, building the *Laguna* models (pre-training/post-training). Previously at Microsoft AI/Research, building the *Phi* models.
More in Data Quality
- Data Quality is the Compute MultiplierTuesday, June 30 · 10:45 AM – 11:05 AM · Track 9
- Rethinking Environments for Long Horizon WorkTuesday, June 30 · 11:40 AM – 12:00 PM · Track 9
- The Base Model is DeadTuesday, June 30 · 1:30 PM – 1:50 PM · Track 9
- Ending AI SlopTuesday, June 30 · 1:55 PM – 2:15 PM · Track 9
- Scaling to Long-Horizons: Algorithms, Environments, ComputeTuesday, June 30 · 2:25 PM – 2:45 PM · Track 9
For developers: this programme is open data — JSON, iCal, schedule XML and an MCP endpoint.Show endpointsHide
- JSONEvery published session and speaker, in one request./aie-worldsfair-2026-import/feed.json
- iCalSubscribe in Google, Apple or Outlook Calendar./aie-worldsfair-2026-import/feed.ics
- Schedule XMLfrab / pentabarf — the format conference apps import./aie-worldsfair-2026-import/feed.xml
- MCP + RESTPoint Claude at the programme. OpenAPI 3.1 included./agents
No key, no signup, CORS open. Everything here is generated from the same data the organisers edit.