SofT-GRPO: Surpassing Discrete-Token LLM Reinforcement Learning via Gumbel-Reparameterized Soft-Thinking Policy Optimization

πŸ“… 2025-11-09
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
To address the challenges of insufficient stochasticity injection, unstable policy updates, and semantic drift in reinforcement learning under soft-thought reasoning paradigms, this paper proposes SofT-GRPOβ€”the first reinforcement learning algorithm tailored for soft-thought policy optimization in large language models (LLMs). Its core innovation lies in the first integration of Gumbel-Softmax with reparameterization techniques, enabling differentiable policy gradient estimation over continuous thought representations and thereby overcoming inherent limitations of discrete chain-of-thought reasoning in stochastic modeling and gradient propagation. SofT-GRPO demonstrates training stability across LLMs ranging from 1.5B to 7B parameters and achieves significant improvements over discrete-token GRPO on multiple reasoning benchmarks: +0.13% in Pass@1 and a substantial +2.19% in Pass@32, empirically validating the effectiveness and generalization advantage of soft-thought policy optimization.

Technology Category

Reasoning under Uncertainty: Stochastic OptimizationSearch and Optimization: Learning to SearchMachine Learning: Reinforcement Learning

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendation
πŸ“ Abstract
The soft-thinking paradigm for Large Language Model (LLM) reasoning can outperform the conventional discrete-token Chain-of-Thought (CoT) reasoning in some scenarios, underscoring its research and application value. However, while the discrete-token CoT reasoning pattern can be reinforced through policy optimization algorithms such as group relative policy optimization (GRPO), extending the soft-thinking pattern with Reinforcement Learning (RL) remains challenging. This difficulty stems from the complexities of injecting stochasticity into soft-thinking tokens and updating soft-thinking policies accordingly. As a result, previous attempts to combine soft-thinking with GRPO typically underperform their discrete-token GRPO counterparts. To fully unlock the potential of soft-thinking, this paper presents a novel policy optimization algorithm, SofT-GRPO, to reinforce LLMs under the soft-thinking reasoning pattern. SofT-GRPO injects the Gumbel noise into logits, employs the Gumbel-Softmax technique to avoid soft-thinking tokens outside the pre-trained embedding space, and leverages the reparameterization trick in policy gradient. We conduct experiments across base LLMs ranging from 1.5B to 7B parameters, and results demonstrate that SofT-GRPO enables soft-thinking LLMs to slightly outperform discrete-token GRPO on Pass@1 (+0.13% on average accuracy), while exhibiting a substantial uplift on Pass@32 (+2.19% on average accuracy). Codes and weights are available on https://github.com/zz1358m/SofT-GRPO-master
Problem

Research questions and friction points this paper is trying to address.

Extending soft-thinking LLM reasoning with Reinforcement Learning remains challenging
Injecting stochasticity into soft-thinking tokens complicates policy optimization updates
Previous soft-thinking GRPO implementations underperform discrete-token reinforcement learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uses Gumbel noise injection for soft-thinking tokens
Applies Gumbel-Softmax to maintain embedding space integrity
Leverages reparameterization trick in policy gradient optimization
πŸ”Ž Similar Papers