🤖 AI Summary
This study addresses the challenge in large language model reasoning where parameter-efficient reinforcement learning struggles to balance computational reuse with performance gains. To this end, it proposes a thalamus routing mechanism that couples an addressable library of compressed computation blocks with a recursive controller, enabling dynamic selection and scaling of early computation blocks for controlled reuse. During training, only interface parameters are updated while the backbone remains frozen, combining correctness-rewarded reinforcement learning with tool-mediated reasoning to enhance capabilities. Evaluated on an 8.95B-parameter model, the approach achieves a MathAvg score of 83.64 (with 60.56 accuracy on AIME) using merely 0.466% additional parameters, significantly outperforming full-parameter GRPO and LoRA baselines.
📝 Abstract
Parameter-efficient reinforcement learning aims to improve reasoning with a compact trainable interface to a pretrained model. We introduce the Thalamic Router (T-Router), which concentrates adaptation on the reuse of completed computations. A compressed, addressable bank preserves block changes; a depth-recurrent controller conditions their selection and relative-scale writeback. This coupling gives thalamic context-dependent routing a concrete computational form: learn which earlier contributions a receiving layer uses, and with what influence. Correctness rewards train the interface while preserving backbone parameters and layer order. On an 8.95B-parameter backbone, T-Router allocates 41.73M parameters (0.466% of the backbone) and achieves 83.64 +/- 1.16 MathAvg after GSM8K RL, compared with 73.79 +/- 1.83 for full-parameter GRPO across three evaluation rounds. At a comparable parameter budget and with matched retries, it exceeds LoRA's 77.28 +/- 1.95 MathAvg, improving all three task families and raising mean AIME accuracy from 48.33 to 60.56. Capacity-controlled comparisons favor addressable block changes and recurrent context; separate search training extends the interface to tool-mediated reasoning. These results establish controlled computation reuse as an effective route to parameter-efficient reasoning reinforcement learning.