Tropical Reinforcement Learning

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of traditional reinforcement learning in supporting compositional reasoning for large models, particularly its susceptibility to catastrophic forgetting and the inability to reuse successful trajectories. To overcome these bottlenecks, this work innovatively introduces tropical semiring algebra to reconstruct the underlying mathematical structure of RL, replacing probabilistic summation with max-plus operations. Building upon this formulation, we propose TROPIC, an algorithm that enables explicit path recombination based on optimal solutions alongside a replayable and composable value evaluation mechanism. Extensive experiments across four categories of agent tasks demonstrate that TROPIC outperforms the strongest baseline by up to 16 percentage points. By facilitating the systematic reuse and composition of previously acquired reasoning paths, this approach establishes a novel paradigm for complex reasoning in large language models.
📝 Abstract
Reinforcement learning for large language models typically maximizes expected return, adding up the probabilities of all successful trajectories. However, the classical sum formulation can only report how often the model policy succeeds, not which solution actually worked, and because probabilities sum to one, reinforcing one solution can make the model forget another that was never shown to be wrong. This makes expected return a poor fit for compositional reasoning, where a solution must be assembled from reasoning steps that the model produces in separate, often failed, attempts but rarely produces together. To address this, we propose Tropical Reinforcement Learning, which rests on a simple change of algebra: instead of adding the probabilities of alternative solutions, we take their maximum, which yields the tropical semiring. The value of a state then becomes the log-probability of its most likely verified solution, together with an explicit path that can be replayed and reused. This enables true composition, since the best prefix and the best suffix meeting at a shared state can be joined even when they come from different rollouts. To put this into practice, we introduce TROPIC, a training algorithm for deterministic, resettable environments with verifiable outcomes. On four agentic tasks (Sokoban, Countdown, FrozenLake, WebShop), TROPIC outperforms the strongest on-policy baselines by up to 16 percentage points. Changing the algebra of reinforcement learning, not just its estimators, can thus substantially improve compositional reasoning in language models
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning
Compositional Reasoning
Expected Return
Large Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Tropical Reinforcement Learning
Tropical Semiring
Compositional Reasoning
TROPIC Algorithm
Large Language Models
🔎 Similar Papers
No similar papers found.