RippleCP: Measuring Counterfactual Checkpoint Advantage in Coding Agents

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the miscalculation of recovery costs in checkpointing systems for coding agents, which stems from the inability to distinguish valid safety boundaries. We propose a counterfactual checkpoint advantage framework that drives dual branches toward identical logical failures and matches conditional recoveries to quantify the performance differential between saving and skipping states. This approach reveals that recovery fundamentally entails re-derivation rather than simple replay, demonstrating that traditional evaluation metrics based on preserved workload introduce severe bias. Controlled experiments on the SWE-bench Verified benchmark show that this strategy saves an average of 49.4 seconds across twelve tasks. Furthermore, the observed first-checkpoint transition rate is 0.64, whereas subsequent checkpoints yield negative values, thereby establishing a performance baseline for optimal checkpoint placement.
📝 Abstract
Agent checkpoint systems decide what state is recovery-relevant, how to snapshot it, and whether rollback is admissible. None decides which of the safe boundaries they expose are worth materializing. We formulate this as counterfactual checkpoint advantage, the reduction in future recovery cost obtained by checkpointing a candidate rather than skipping it, and measure it by driving a CP branch and a SKIP branch to the same logical failure and recovering both under matched model, tool, verifier, and stopping conditions. On a frozen pilot of 12 SWE-bench Verified tasks and 106 real recovery branches, checkpointing saves 49.4 s per task, and that figure resolves into two regimes two orders of magnitude apart. The first checkpoint returns 100.0 s on 156.5 s of protected work, a conversion of 0.64; a second one step later returns-1.1 s on 41.8 s, a conversion of -0.03. Recovery is a re-derivation rather than a replay, so preserved work is a poor guide to saved work, and the classical elapsed-work rule misprices the second checkpoint by its full nominal cost. We identify where placement can pay, and set the bar a placement policy must clear.
Problem

Research questions and friction points this paper is trying to address.

checkpoint advantage
coding agents
counterfactual evaluation
recovery cost
checkpoint placement
Innovation

Methods, ideas, or system contributions that make the work stand out.

Counterfactual Checkpoint Advantage
Coding Agents
Checkpoint Placement
Recovery Cost
RippleCP
🔎 Similar Papers
2024-02-08International Conference on Machine LearningCitations: 6