Beyond Static Intelligence: Evaluating Continual Learning
Parth Asawa
- When
- Wednesday, July 110:45 AM – 11:05 AM · 20 min
- Where
- Track 3San Francisco, CA · imported from ai.engineer's public schedule feed
About this session
Continual learning, the ability of AI systems to improve through sequential experience, has attracted substantial interest, but no high-quality benchmark exists to evaluate it. We introduce Continual Learning Bench (CL-Bench), the first difficult, expert-validated benchmark designed to measure whether LLM-based systems genuinely improve with experience. CL-Bench spans six diverse domains (software engineering, signal processing, disease outbreak forecasting, database querying, strategic game-playing, and demand forecasting), each validated by domain experts and designed so that tasks share a learnable latent structure (codebase layout, disease outbreak dynamics, opponent strategies) that a stateful system can discover online but a stateless one cannot. We evaluate frontier models across several agent architectures, from naive in-context learning (ICL) to dedicated memory systems, introducing a gain metric to isolate learning from prior capabilities. We find that these systems leave headroom for improved continual learning: agents frequently overfit to immediate observations or fail to reuse knowledge across instances, and dedicated memory systems do not fix this---in fact, naive ICL outperforms systems dedicated to memory management. CL-Bench is the first benchmark to evaluate continual learning across diverse real-world domains with expert-validated tasks and isolate online learning from underlying model capability, showing a need for better continual learning systems.
Speaker
CS PhD student, UC Berkeley
Parth Asawa is a PhD student at UC Berkeley advised by Professor Matei Zaharia and Professor Joey Gonzalez. Parth's research is on continual learning, studying how to enable models to stably learn from streams of experiences over time. His work focuses on sample-efficient learning and spans the stack of data, learning algorithms, architectures, and evaluation.
More in Memory & Continual Learning
- Scaling up Continual LearningWednesday, July 1 · 11:10 AM – 11:30 AM · Track 3
- Memory Harnesses for Long-Running Research AgentsWednesday, July 1 · 11:40 AM – 12:00 PM · Main Stage
- Scaling Compute on ContextWednesday, July 1 · 11:40 AM – 12:00 PM · Track 3
- Intelligence + Continual Learning = ExpertiseWednesday, July 1 · 12:05 PM – 12:25 PM · Track 3
- Adaption Labs — Gradient-Free Continual LearningWednesday, July 1 · 1:30 PM – 1:50 PM · Track 3
For developers: this programme is open data — JSON, iCal, schedule XML and an MCP endpoint.Show endpointsHide
- JSONEvery published session and speaker, in one request./aie-worldsfair-2026-import/feed.json
- iCalSubscribe in Google, Apple or Outlook Calendar./aie-worldsfair-2026-import/feed.ics
- Schedule XMLfrab / pentabarf — the format conference apps import./aie-worldsfair-2026-import/feed.xml
- MCP + RESTPoint Claude at the programme. OpenAPI 3.1 included./agents
No key, no signup, CORS open. Everything here is generated from the same data the organisers edit.