The Base Model is Dead
Varun Singh
- When
- Tuesday, June 301:30 PM – 1:50 PM · 20 min
- Where
- Track 9San Francisco, CA · imported from ai.engineer's public schedule feed
About this session
It's a common belief that large language models are trained to be a good model of human web-text, and thus base models are "mirrors" of what we see on the internet. Historically, this was largely true, but no modern base model truly reflects the internet in the way that GPT-3 once did. Instruction data along with synthetic reasoning traces are moving earlier and earlier into the training pipeline, and "mid-training" has emerged as a new stage to accommodate longer datapoints that more concretely resemble downstream capabilities. As a result, pre-training no longer has the goal of creating a linguistic prior, but instead has the additional goals of baking in behavior and more atomic skills into the trained "base" model. Between this shift in what a base model is and the blurring of the lines between the different stages of model training, it's an open question as to what the best approach is here (at least outside the walls of the big labs). But I believe that the role we view the base model playing will continue to shift as we're pulled forward through new phases of model capabilities.
Speaker
Pre-Training Lead, Arcee AI
Pre-training lead at Arcee AI working on end to end pre-training of large language models, with a strong interest in architecture and optimization. Led the pre-training of Arcee's Trinity series of models, ranging from a 6B mixture-of-experts to a 400B mixture-of-experts model.
More in Data Quality
- Data Quality is the Compute MultiplierTuesday, June 30 · 10:45 AM – 11:05 AM · Track 9
- The Messy Reality of Scale: Synthetic Data and Pre-Training at PoolsideTuesday, June 30 · 11:10 AM – 11:30 AM · Track 9
- Rethinking Environments for Long Horizon WorkTuesday, June 30 · 11:40 AM – 12:00 PM · Track 9
- Ending AI SlopTuesday, June 30 · 1:55 PM – 2:15 PM · Track 9
- Scaling to Long-Horizons: Algorithms, Environments, ComputeTuesday, June 30 · 2:25 PM – 2:45 PM · Track 9
For developers: this programme is open data — JSON, iCal, schedule XML and an MCP endpoint.Show endpointsHide
- JSONEvery published session and speaker, in one request./aie-worldsfair-2026-import/feed.json
- iCalSubscribe in Google, Apple or Outlook Calendar./aie-worldsfair-2026-import/feed.ics
- Schedule XMLfrab / pentabarf — the format conference apps import./aie-worldsfair-2026-import/feed.xml
- MCP + RESTPoint Claude at the programme. OpenAPI 3.1 included./agents
No key, no signup, CORS open. Everything here is generated from the same data the organisers edit.