RolloutFaith: Auditing Persistent Internal Interventions in Visual World Model

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that internal interventions in visual world models struggle to sustain their effects during long-horizon autonomous prediction. To tackle this, it proposes the RolloutFaith framework, which introduces a novel metric for evaluating intervention persistence and benchmarks reference activation patching combined with low-rank editing in environments such as Crafter. Building upon these findings, the work further presents Delayed LoReFT, an optimized intervention method. The study reveals that current editors exhibit limited long-term efficacy, identifying newly generated frames and memory states as critical carriers of persistent effects. By validating the training hypothesis that rewarding future consequences enhances intervention durability, the proposed approach significantly improves persistence, thereby establishing a new paradigm for the reliable editing of world models.
📝 Abstract
Interpretability methods such as probes, activation patches and learned editors are designed to reveal or modify a model's current computation. World models pose a harder requirement: because their predictions become inputs to later predictions, a useful internal correction must survive after editing stops. We therefore propose RolloutFaith, a framework that measures semantic improvement both in the prediction produced at intervention time and over later autonomous predictions under fixed events, actions, noise, and information budgets. We evaluate ten fitted editors on three world models across Crafter, Cartpole, and CoinRun. We also use Reference Activation Patching, which replaces a model activation with the paired activation computed from the real observation, to measure the correction available at the chosen interface. This reference intervention improves later predictions in all nine model and task combinations and outperforms the best fitted editor in eight, yet its sustained gain decreases with horizon in five of nine combinations. Current fitted editors recover only limited and inconsistent long term effects. By restoring individual state components to their untouched values, we find that persistent effects travel through the newest generated frame in DIAMOND, recurrent memory in DreamerV3, and both in STORM. These findings suggest that training should reward future consequences. To test this hypothesis, we propose Delayed LoReFT, which optimizes the same low rank intervention through four frozen future transitions and improves sustained intervention effects to some extent.
Problem

Research questions and friction points this paper is trying to address.

Visual World Model
Internal Intervention
Interpretability
Persistent Effect
Learned Editor
Innovation

Methods, ideas, or system contributions that make the work stand out.

RolloutFaith
World Models
Interpretability
Activation Patching
Delayed LoReFT