LatentHarness: Learning Latent Actions for Memory and Reasoning via Counterfactual Policy Distillation

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the bottleneck in long-context reasoning where evidence retention and multi-step computation are difficult to reconcile. It pioneers a sequential latent action selection framework that unifies external memory access with internal latent reasoning. By incorporating a fast-weight memory mechanism and employing counterfactual policy distillation, the model is trained to dynamically decide among thinking, recalling, and exiting, thereby effectively optimizing memory content selection. Experimental results demonstrate that a model with only 1.4B parameters achieves substantial performance improvements across multiple benchmarks while attaining an inference speed 5.9 times faster than the strongest baseline. This work thus realizes a dual breakthrough in both computational efficiency and predictive accuracy for long-context reasoning tasks.
📝 Abstract
Long-context reasoning faces two complementary bottlenecks: retaining evidence across long inputs and sustaining computation across many reasoning steps. Existing approaches largely address them separately, with external memory extending access to distant evidence and latent reasoning compressing multi-step computation. We introduce LatentHarness, which unifies memory access and latent reasoning as sequential latent action selection. At each internal step, the model chooses THINK for further computation, RECALL from a fast-weight memory of input evidence and intermediate reasoning states, or EXIT to emit the next token. We train this policy with counterfactual policy distillation, which branches every action for one step and scores its effect on the emitted token. These gains teach the policy when memory is more useful than further reasoning, while gradients through counterfactual recall teach which intermediate states should be retained in memory for future use. Across six general and long-context reasoning benchmarks, LatentHarness at 1.4B improves on the strongest baselines by 2.8% and 10.0% relative, respectively, and runs 5.9x faster than the strongest long-context baseline.
Problem

Research questions and friction points this paper is trying to address.

long-context reasoning
memory retention
latent reasoning
multi-step computation
evidence access
Innovation

Methods, ideas, or system contributions that make the work stand out.

Latent Actions
Counterfactual Policy Distillation
Long-context Reasoning
Fast-weight Memory
Latent Reasoning
🔎 Similar Papers
No similar papers found.