ARISE: Adapting to Evolving Capability Gaps in Agentic Reinforcement Learning

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the misalignment between static training criteria and dynamically evolving agent capabilities, where sparse feedback fails to guide effective exploration toward bridging emerging skill gaps. To this end, we propose ARISE, a framework introducing the first Rubric-Skill co-evolution mechanism. ARISE leverages rollout evidence to dynamically adjust evaluation rubrics and skill libraries, guiding exploration through partial progress rewards and targeted skill activation. Furthermore, it incorporates capability-aware adaptive sampling to focus on weaknesses, thereby achieving dynamically aligned reinforcement learning. Experimental results demonstrate that ARISE significantly improves both long-horizon task performance and training efficiency on the SkillsBench and Terminal-Bench benchmarks.
📝 Abstract
As a long-horizon agent improves through experience, previously observed weaknesses may recede while new limitations emerge, continually changing what it still needs to learn. Yet the learning process often remains tied to a static view of these needs: fixed behavioral criteria and training priorities can become misaligned with evolving agent capabilities, while sparse task-level feedback makes such misalignment more difficult to detect. Even when capability gaps are identified, rollouts from the current policy may repeatedly reproduce the same failures rather than explore better alternatives. To address this, we introduce Adaptive Rubric-Skill Co-Evolution (ARISE), a reinforcement learning framework that uses rollout evidence to continually adapt evaluation criteria, exploration guidance, and training priorities. Rubrics evolve to reward partial behavioral progress, while their paired skills are refined and selectively activated to guide exploration toward unresolved weaknesses. Alongside this co-evolution, capability-based adaptive sampling prioritizes tasks that target behaviors needing further improvement. Experiments on two challenging long-horizon agent benchmarks, SkillsBench and Terminal-Bench, demonstrate that ARISE successfully enhances both overall task performance and training efficiency. The project page is at https://foundation-model-research.github.io/ARISE .
Problem

Research questions and friction points this paper is trying to address.

Agentic Reinforcement Learning
Long-horizon Agent
Capability Gaps
Static Evaluation Criteria
Sparse Feedback
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agentic Reinforcement Learning
Adaptive Rubric-Skill Co-Evolution
Capability-based Adaptive Sampling
Long-horizon Agent
Exploration Guidance
🔎 Similar Papers
No similar papers found.