🤖 AI Summary
This study addresses the limitation that deep computations in Transformers cannot refine early-layer representations, as well as the high cost and poor adaptability of existing feedback mechanisms. To this end, we propose ReFlux, a learnable feedback graph. This method innovatively introduces a depth increment Δ as a feedback signal, combining residual stream state write-back with streaming processing to construct dynamic composite routing for efficient synchronous and streaming feedback. It fully unleashes the model's latent computational capacity without increasing theoretical FLOPs. Experiments demonstrate that ReFlux significantly reduces perplexity, improving overall accuracy by 2.1–2.3 points and multi-hop reasoning performance by 4.7 points, thereby effectively enhancing the model's complex reasoning capabilities.
📝 Abstract
Transformers process information strictly forward through depth, preventing deeper computation from revisiting and refining earlier representations. To augment the standard forward pass, existing approaches either re-execute depth, incurring additional computation, or modify the residual stream using predefined directions, limiting their instance-level adaptation. Recently, inference-time feedback offers a direct mechanism for recycling endogenously produced computation by writing deeper residual states back to earlier layers, yet what should be fed back remains unclear. We argue that the depth increment Delta, capturing newly accumulated computation between two layers, provides a more effective, composable, and scalable feedback signal than the full state. Building on this observation, we introduce ReFlux, a learnable feedback graph that dynamically selects and composes increment-carrying routes. ReFlux supports synchronous feedback to the same token and streaming feedback to subsequent tokens. Extensive experiments across various models, corpora, and benchmarks show that synchronous ReFlux consistently reduces perplexity across ten language-modeling corpora, and improves accuracy by 2.1-2.3 points, with gains reaching 4.7 points on multi-hop reasoning. Streaming ReFlux further retains most of these gains while preserving the base model's 1x theoretical backbone FLOPs. These results establish ReFlux as an efficient paradigm for unlocking the latent computational potential of LLMs, allowing them to revisit the same tokens with fresh representations. Code implementation can be found at https://github.com/gooogleshanghai/reflux.