Memory Harnesses for Long-Running Research Agents
Stefania Druga
- When
- Wednesday, July 111:40 AM – 12:00 PM · 20 min
- Where
- Main StageSan Francisco, CA · imported from ai.engineer's public schedule feed
About this session
At Sakana AI we build agents that run for hundreds of turns to read literature, run experiments, and draft papers. The model rarely breaks. The harness around it is the weak point: the agent contradicts a decision it made 80 turns ago, redoes finished work, or drifts from the question it started on. This is the binding-constraint thesis. For long-horizon tasks, reliability is set as much by the harness as by the model as clearly instantiated in autoresearch recent efforts. This is a field guide to the harness's memory layer. I'll trace a real research agent through its lifecycle, show exactly where context rot and drift set in, and cover the patterns that hold over 100+ turns: three-tier memory, progressive disclosure, recall-first compaction, sub-agent isolation, and architectural memory beyond the vector database. I will show how to measure whether your memory harness actually helps, at the trajectory level, so you stop tuning prompts to fix what's really a state-management bug.
Speaker
Research Scientist, Sakana.ai
Hi! I am Stef. I am currently a Research Scientist at Sakana AI in Tokyo, Japan working on novel architectures beyond the transformer. Previously I was a research at Google Deep Mind working on novel multimodal AI applications. I graduated with a Ph.D. in Creative AI Literacies at the University of Washington Information School.
More in Memory & Continual Learning
- Beyond Static Intelligence: Evaluating Continual LearningWednesday, July 1 · 10:45 AM – 11:05 AM · Track 3
- Scaling up Continual LearningWednesday, July 1 · 11:10 AM – 11:30 AM · Track 3
- Scaling Compute on ContextWednesday, July 1 · 11:40 AM – 12:00 PM · Track 3
- Intelligence + Continual Learning = ExpertiseWednesday, July 1 · 12:05 PM – 12:25 PM · Track 3
- Adaption Labs — Gradient-Free Continual LearningWednesday, July 1 · 1:30 PM – 1:50 PM · Track 3
For developers: this programme is open data — JSON, iCal, schedule XML and an MCP endpoint.Show endpointsHide
- JSONEvery published session and speaker, in one request./aie-worldsfair-2026-import/feed.json
- iCalSubscribe in Google, Apple or Outlook Calendar./aie-worldsfair-2026-import/feed.ics
- Schedule XMLfrab / pentabarf — the format conference apps import./aie-worldsfair-2026-import/feed.xml
- MCP + RESTPoint Claude at the programme. OpenAPI 3.1 included./agents
No key, no signup, CORS open. Everything here is generated from the same data the organisers edit.