Ladders of Thought: A Self-Evolving Curriculum of Progressively Simplified Reasoning Traces

📅 2026-09-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文提出Ladders-of-Thought框架,通过逐步简化问题和自适应课程来提高小规模语言模型的推理能力,优于知识蒸馏方法。
📝 Abstract
Large language models (LLMs) excel at reasoning when scaled to hundreds of billions of parameters, but small- and mid-scale models remain brittle reasoners even with knowledge distillation (KD). We present Ladders-of-Thought (LoT), a framework that improves reasoning by combining progressive question rewrites with a self-evolving curriculum. LoT automatically generates semantically faithful but easier variants of reasoning problems, organizes them into difficulty buckets using step-based measures, and employs a self-evolving bandit scheduler to allocate training adaptively. Evaluated on two reasoning domains, math and multi-hop reasoning, across 1-8B models from different families, LoT consistently improves over KD. It delivers large gains on arithmetic tasks (e.g., +32 percentage points on AddSub, +25pp on SVAMP), +2-8pp improvements on in-domain test splits, and strong though dataset-dependent benefits on multi-hop reasoning (e.g., +16pp on QASC, +25pp on StrategyQA). LoT also converges faster than staged curricula, highlighting the value of adaptive progression. These results show that progressive rewrites coupled with adaptive curricula provide a simple yet effective recipe for strengthening reasoning in smaller LLMs.
Problem

Research questions and friction points this paper is trying to address.

language models
reasoning
knowledge distillation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Ladders-of-Thought
progressive question rewrites
self-evolving curriculum
adaptive training
reasoning improvement
🔎 Similar Papers
No similar papers found.