From Chain-of-Thought to Loops: Non-Autoregressive Latent Reasoning via Looped Transformers

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high latency inherent in autoregressive generation of explicit chain-of-thought reasoning and the persistent serial dependencies that limit existing latent reasoning approaches. To overcome these challenges, this work proposes LLoCoT, a framework for non-autoregressive latent reasoning. Rather than generating tokens sequentially, LLoCoT employs a recurrent Transformer to iteratively refine a compact latent workspace, while utilizing a probabilistic head to sample latent tokens in parallel to facilitate decoding. Experimental evaluations on HumanEval and MBPP demonstrate that the proposed framework achieves accuracy comparable to supervised fine-tuning baselines while reducing time-to-first-token latency by 36× and improving throughput by 9.2%. These results indicate that LLoCoT effectively balances inference efficiency with generation quality.
📝 Abstract
Chain-of-thought (CoT) reasoning often improves language-model performance by giving models additional computation before answering. However, explicit CoT expresses this computation as a sequence of autoregressively generated tokens. Latent reasoning replaces these tokens with compact continuous states, but most autoregressive latent-reasoning methods retain a left-to-right dependency among latent vectors. We introduce LLoCoT: a looped latent-reasoning framework that replaces left-to-right latent generation with iterative refinement of a compact latent workspace. A shared transformer is reapplied for a small number of refinement iterations, jointly updating the latent slots based on the prompt and the evolving workspace state. Using the refined state, a probabilistic head predicts a distribution from which latent tokens are sampled in parallel and used to condition an autoregressive decoder for answer generation. Training uses continuous representations derived from explicit CoT together with a final-answer prediction loss and likelihood-based supervision of the latent states. Across HumanEval and MBPP, LLoCoT achieves the highest mean among the evaluated methods, performing on par in accuracy with Reasoning SFT, our explicit-CoT baseline, while outperforming the base model, answer-only SFT and NF-CoT. Relative to Reasoning SFT, LLoCoT reduces time to the first answer token by approximately $36\times$ and reasoning-phase latency by approximately $42\times$, while increasing end-to-end throughput by $9.2\%$. This design replaces serial thought generation with parallel latent-slot refinement while retaining probabilistic latent modeling and autoregressive answer decoding.
Problem

Research questions and friction points this paper is trying to address.

Chain-of-Thought reasoning
Latent reasoning
Autoregressive generation
Inference latency
Non-autoregressive decoding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Looped Transformers
Non-Autoregressive Latent Reasoning
Iterative Refinement
Chain-of-Thought
Parallel Latent Slots
🔎 Similar Papers
2024-02-26Annual Meeting of the Association for Computational LinguisticsCitations: 97
G
Gerard Grau García
Qualcomm AI Research, MIT
A
Arnau Padrés Masdemont
Qualcomm AI Research
Niccolò Grillo
Niccolò Grillo
Politecnico di Milano
Geometric Deep LearningReinforcement LearningAlgorithmic Reasoning
J
Jordi Ros-Giralt
Qualcomm AI Research
A
Arash Behboodi
Qualcomm AI Research
V
Victor Conchello Vendrell
Qualcomm AI Research, Harvard University