Reinforcement Learning without Verifiable Rewards
Will Brown
- When
- Wednesday, July 11:30 PM – 1:50 PM · 20 min
- Where
- Track 9San Francisco, CA · imported from ai.engineer's public schedule feed
About this session
Verifiable rewards are the gold standard for RL training, but real-world agent tasks frequently lack clean deterministic evaluation objectives. This talk surveys our efforts to scale RL in non-verifiable settings -- including task synthesis, unsupervised environment design, and automatic judge calibration -- to ultimately enable self-improvement in production, grounded in real-world agent traces and domain-specific context.
Speaker
Researcher, Prime Intellect
Will Brown leads Applied Research at Prime Intellect and builds open research infrastructure to enable every company to train, deploy, and self-improve their own frontier agentic models. He holds a PhD in Computer Science from Columbia University.
More in Posttraining & Midtraining
- Building self-learning loops for your agentMonday, June 29 · 11:05 AM – 12:05 PM · Track 1
- Bugcrowd posttraining talkTuesday, June 30 · 12:05 PM – 12:25 PM · Track 9
- What's next after RLHF?Wednesday, July 1 · 10:45 AM – 11:05 AM · Track 9
- State of DataWednesday, July 1 · 11:10 AM – 11:30 AM · Track 9
- Learning on the job: the future of post-trainingWednesday, July 1 · 12:05 PM – 12:25 PM · Track 9
For developers: this programme is open data — JSON, iCal, schedule XML and an MCP endpoint.Show endpointsHide
- JSONEvery published session and speaker, in one request./aie-worldsfair-2026-import/feed.json
- iCalSubscribe in Google, Apple or Outlook Calendar./aie-worldsfair-2026-import/feed.ics
- Schedule XMLfrab / pentabarf — the format conference apps import./aie-worldsfair-2026-import/feed.xml
- MCP + RESTPoint Claude at the programme. OpenAPI 3.1 included./agents
No key, no signup, CORS open. Everything here is generated from the same data the organisers edit.