Skip to content
BenchmarkAG-2026-0166

Agent memory is a dose, not a switch

Across 585 tasks and eight models, the same memory strategy gained 16 points on one model and nothing at all on another.

3 minAutonomous AI Agents

IBM Research ran eight models across AppWorld — 585 multi-step tasks, 168 normal and 417 challenge variants over nine simulated applications — to ask how much memory an agent actually needs. The method, ALTK-Evolve, extracts behavioural guidelines from past trajectories and injects them at inference time, with no weight updates.

The results do not converge. On gpt-oss-120b, task goal completion rose from 39.9% to 56.0%, a gain of 16.1 points. DeepSeek-V3.2 went from 79.8% to 89.3%. Claude Opus 4.6 moved 90.5% to 94.6%, GPT-5.5 92.3% to 95.2%. GLM-5 gained nothing: 87.5% before and after.

Three patterns, not one rule

  • Strong models with headroom do best on the full guideline set.
  • Weaker models gain most from a tight high-confidence core plus a few task-relevant guidelines retrieved per task.
  • Saturated models gain nothing, and the memory is pure overhead.

The token accounting explains why the distinction matters. On gpt-oss-120b, curated retrieval added 5% — 110K tokens to 116K — while the full guideline set added 51%, reaching 166K. On DeepSeek-V3.2 the full set pushed 148K to 263K, a 78% increase.

So the choice is not whether to give an agent memory but how much, and the answer differs per model. A configuration that buys 16 points on one model buys nothing on another while costing three-quarters more tokens — which is the authors' own framing: memory is a dose you calibrate, not a feature you switch on.

Retold from Hugging Face. This is a summary in our own words; follow the link for the original reporting.

Read next

Across the network

Desks that share a zone with this one on the BITBRIEF coverage map.

Terms defined