Escaping the Verifier: Learning to Reason via Demonstrations

📅 2025-11-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Large language models (LLMs) struggle to learn complex reasoning from expert demonstrations when no task-specific verifier is available. Method: This paper proposes RARO—a novel inverse reinforcement learning (IRL) framework that introduces a relative adversarial mechanism to implicitly model reasoning rewards solely from expert demonstrations. RARO jointly optimizes a generative policy and a relative discriminator via adversarial interaction and stabilization techniques, enabling end-to-end reasoning improvement without explicit verification signals or hand-crafted reward annotations. Contribution/Results: RARO eliminates the conventional IRL dependency on task-specific verifiers. Empirical evaluation demonstrates that RARO significantly outperforms verifier-free baselines on unverifiable tasks—including Countdown, DeepMath, and poetic composition—and achieves robust scalability comparable to that observed on verifiable tasks.

Technology Category

Machine Learning: Imitation Learning & Inverse Reinforcement LearningMultiagent Systems: Adversarial AgentsKnowledge Representation and Reasoning: Argumentation

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationEconomics, Online Markets and Human Computation: LLM based quality controls for crowd work
📝 Abstract
Training Large Language Models (LLMs) to reason often relies on Reinforcement Learning (RL) with task-specific verifiers. However, many real-world reasoning-intensive tasks lack verifiers, despite offering abundant expert demonstrations that remain under-utilized for reasoning-focused training. We introduce RARO (Relativistic Adversarial Reasoning Optimization) that learns strong reasoning capabilities from only expert demonstrations via Inverse Reinforcement Learning. Our method sets up an adversarial interaction between a policy (generator) and a relativistic critic (discriminator): the policy learns to mimic expert answers, while the critic learns to compare and distinguish between policy and expert answers. Our method trains both the policy and the critic jointly and continuously via RL, and we identify the key stabilization techniques required for robust learning. Empirically, RARO significantly outperforms strong verifier-free baselines on all of our evaluation tasks -- Countdown, DeepMath, and Poetry Writing -- and enjoys the same robust scaling trends as RL on verifiable tasks. These results demonstrate that our method effectively elicits strong reasoning performance from expert demonstrations alone, enabling robust reasoning learning even when task-specific verifiers are unavailable.
Problem

Research questions and friction points this paper is trying to address.

Learning reasoning from expert demonstrations without verifiers
Using adversarial interaction to mimic expert answers
Achieving robust reasoning performance on verifier-free tasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uses Inverse Reinforcement Learning from expert demonstrations
Adversarial policy-critic interaction for reasoning optimization
Joint continuous training with stabilization techniques
🔎 Similar Papers
No similar papers found.