Guiding Through Complexity: What Makes Good Supervision for Hard Reasoning Tasks?

📅 2024-10-27
🏛️ arXiv.org
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work investigates how to effectively enhance large language models’ performance on high-difficulty reasoning tasks—such as mathematical proof generation and complex logical deduction—using weak supervision sources (e.g., non-expert annotators or off-the-shelf AI systems). Addressing the trade-off between supervision quality and task difficulty, we establish a key empirical finding: supervision for hard tasks with high step-level error rates yields superior model performance compared to error-free supervision on easy tasks; critically, step-level error rate proves a more informative training signal than final-answer accuracy. Building on this insight, we propose a collaborative supervision paradigm that jointly leverages subtasks and hard tasks, integrated with supervised data mixing, dynamic reweighting, and empirically grounded evaluation. Our approach achieves up to 30% absolute accuracy gains on challenging benchmarks such as MATH. All code and datasets are publicly released.

Technology Category

Machine Learning: Calibration & Uncertainty QuantificationKnowledge Representation and Reasoning: Computational Complexity of ReasoningNatural Language Processing: (Large) Language Models

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsEconomics, Online Markets and Human Computation: Data quality aspects of human-annotated datasetsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for ranking
📝 Abstract
How can"weak teacher models"such as average human annotators or existing AI systems, effectively supervise LLMs to improve performance on hard reasoning tasks, especially those that challenge and requires expertise or daily practice from the teacher models? In this paper, we seek for empirical answers to this question by investigating various data-driven strategies that offer supervision data at different quality levels upon tasks of varying complexity. Two intuitive strategies emerge for teacher models to provide supervision during alignment training: 1) using lower-quality supervision from complete tasks that match the difficulty of the target reasoning tasks, and 2) leveraging higher-quality supervision from easier subtasks that are less challenging. Interestingly, we find that even when the outcome error rate for hard task supervision is high (e.g., 90%), training on such data can outperform perfectly correct supervision of easier subtasks on multiple hard math benchmarks. We further identify a more critical factor influencing training performance: step-wise error rates, which indicate the severity of errors in solutions. Specifically, training on hard task supervision with the same outcome error rates but disparate step-wise error rates can lead to a 30% accuracy gap on MATH benchmark. Our results also reveal that supplementing hard task supervision with the corresponding subtask supervision can yield notable performance improvements than simply combining rephrased hard full task supervision, suggesting new avenues for data augmentation. Data and code are released at https://github.com/hexuan21/Weak-to-Strong.
Problem

Research questions and friction points this paper is trying to address.

Effective supervision by weak teacher models
Improving LLMs on hard reasoning tasks
Impact of step-wise error rates
Innovation

Methods, ideas, or system contributions that make the work stand out.

Utilizing low-quality complete task supervision
Leveraging high-quality easier subtask supervision
Focusing on step-wise error rates impact
🔎 Similar Papers
No similar papers found.
Tsinghua University | University of California, Los Angeles
X
Xuan He
Tsinghua University
Da Yin
Da Yin
Meta FAIR
Natural Language Processing
N
Nanyun Peng
University of California, Los Angeles