Long-running agents need checkpoints, not longer timeouts
Teams keep raising the ceiling instead of making runs resumable.
7 minAutonomous AI Agents
The common response to an agent that fails at minute forty is to allow it sixty. This works until the failure rate compounds, which it does, because a run that takes an hour touches more things that can break.
Checkpointing after each externally visible action costs little and turns a failed run into a resumed one. The teams that adopted it early report the same thing: their reliability numbers improved without any change to the model.