Agent memory is a dose, not a switch
Across 585 tasks and eight models, the same memory strategy gained 16 points on one model and nothing at all on another.
3 minAutonomous AI Agents
IBM Research ran eight models across AppWorld — 585 multi-step tasks, 168 normal and 417 challenge variants over nine simulated applications — to ask how much memory an agent actually needs. The method, ALTK-Evolve, extracts behavioural guidelines from past trajectories and injects them at inference time, with no weight updates.
The results do not converge. On gpt-oss-120b, task goal completion rose from 39.9% to 56.0%, a gain of 16.1 points. DeepSeek-V3.2 went from 79.8% to 89.3%. Claude Opus 4.6 moved 90.5% to 94.6%, GPT-5.5 92.3% to 95.2%. GLM-5 gained nothing: 87.5% before and after.
Three patterns, not one rule
- Strong models with headroom do best on the full guideline set.
- Weaker models gain most from a tight high-confidence core plus a few task-relevant guidelines retrieved per task.
- Saturated models gain nothing, and the memory is pure overhead.
The token accounting explains why the distinction matters. On gpt-oss-120b, curated retrieval added 5% — 110K tokens to 116K — while the full guideline set added 51%, reaching 166K. On DeepSeek-V3.2 the full set pushed 148K to 263K, a 78% increase.
So the choice is not whether to give an agent memory but how much, and the answer differs per model. A configuration that buys 16 points on one model buys nothing on another while costing three-quarters more tokens — which is the authors' own framing: memory is a dose you calibrate, not a feature you switch on.
Retold from Hugging Face. This is a summary in our own words; follow the link for the original reporting.