Learning What to Investigate Next: Meta-Reasoning for Long-Horizon Research Agents

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the decision-making challenges and credit assignment difficulties under sparse signals encountered by long-horizon research agents during evidence accumulation. To this end, we propose MIRA, a hierarchical architecture that decouples research planning from execution. This approach explicitly formulates meta-reasoning as a learnable policy for the first time, employing a generative critic to evaluate the value of partial progress while training an Actor-Critic model to optimize decision boundaries, thereby circumventing the direct optimization of lengthy execution trajectories. By integrating cross-environment pretraining with self-supervised agent signals, MIRA significantly enhances performance on tasks such as theorem proving. Ultimately, it achieves notable advances in efficient compute allocation, cross-environment value transfer, and autonomous research capabilities.
📝 Abstract
Long-horizon research agents must decide both how to investigate and what to investigate next as evidence accumulates. This is hard to learn because such decisions are sparse in long execution traces, and their consequences may emerge several investigations later. We introduce Meta-reasoning for Iterative Research Agents (MIRA), a hierarchical architecture separating research allocation from execution. An outer-loop meta-reasoner curates context from a persistent research record, then writes a work order for the next investigation or ends the episode. A fresh inner-loop executor carries out each work order, making execution part of the transition between meta-reasoning actions. Without policy training, MIRA improves long-horizon inference and allocates additional compute more effectively in theorem proving and open-ended neural-architecture research. Its decision boundaries also provide natural units for credit assignment. At each boundary, we train a generative critic to forecast expected remaining return from partial states, outperforming token-level alternatives. Cross-environment pretraining improves forecasting and adaptation, yielding a transferable prior for valuing partial progress. We use this prior to initialize MIRA-AC, a generative actor-critic jointly trained to forecast remaining return and choose the next investigation, without a separate critic model. MIRA-AC concentrates policy optimization on meta-reasoning decisions, enabling efficient long-horizon reinforcement learning without directly optimizing the longer execution traces they initiate. Training MIRA-AC on the model's own proxy hill-climbing signals improves gold performance across four autoresearch environments; the actor transfers with cross-environment value initialization. Together, these results show that meta-reasoning can be learned as an explicit policy for directing long-horizon autonomous research.
Problem

Research questions and friction points this paper is trying to address.

long-horizon research agents
meta-reasoning
sparse decision-making
credit assignment
reinforcement learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Meta-reasoning
Hierarchical architecture
Long-horizon reinforcement learning
Generative actor-critic
Credit assignment