Compositional Reasoning in Language Models under Reinforcement Learning Post-Training

📅 2026-09-16
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究通过强化学习后训练方法提升语言模型的组合推理能力,提出依赖图框架以形式化组合推理,并发现分解技能到组合任务的不对称性。
📝 Abstract
Compositional reasoning is critical for real-world problem solving: since training data is necessarily limited, models must generalize by composing learned skills in new ways. While post-training methods such as reinforcement learning (RL) have substantially improved the reasoning abilities of language models (LMs), their effects on compositional reasoning remain less well understood. We propose a dependency-graph framework to formalize compositional reasoning, yielding three levels of compositionality with increasing complexity. Empirically, we instantiate this framework with data-structure tasks, which provide deterministic reward computation and clear compositional structure. We find a consistent decomposed-to-composed asymmetry: decomposed-skill training does not reliably transfer to composed tasks, whereas composed-task training transfers more readily back to decomposed tasks. We provide theoretical explanation for this asymmetry, and further evaluate compositional generalization under length extrapolation, structural distribution shift, and transfer to tasks requiring unseen skills. Finally, we present a pilot study on real-world tool-calling benchmarks, showing preliminary evidence that the decomposed-to-composed asymmetry can extend to practical settings.
Problem

Research questions and friction points this paper is trying to address.

Compositional Reasoning
Reinforcement Learning
Language Models
Post-Training
Innovation

Methods, ideas, or system contributions that make the work stand out.

dependency-graph framework
compositional reasoning
reinforcement learning
asymmetry in skill transfer
🔎 Similar Papers
No similar papers found.
Y
Yu He
Stanford University
Y
Yingxi Li
Stanford University
Y
Yifei Wang
Amazon AGI Labs
Ellen Vitercik
Ellen Vitercik
Stanford University