UniSkill: Learning Actor-Aligned Skill Proposals for an Evolving Policy

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of reward conflation and prohibitive evaluation costs in skill extraction and policy evolution for LLM-based agents. To this end, it proposes a contrastive action feedback mechanism grounded in shared policy trajectories to guide skill proposal learning, enabling the co-evolution of skill libraries without requiring additional rollouts. Furthermore, an actor alignment signal based on log-likelihood differences and a skill editing regularization term are designed to balance policy optimization with the preservation of exploratory behavior. By integrating reinforcement learning with contrastive learning, the proposed approach achieves success rates of 98.4% and 84.7% on the ALFWorld and WebShop benchmarks, respectively. Notably, it maintains robust performance even with smaller-scale models, thereby establishing a novel paradigm for efficient agent evolution.
📝 Abstract
Large language model agents can improve across tasks by retaining reusable skills distilled from prior interactions. Recent work jointly optimizes task execution and skill extraction, enabling the policy and skillbank to co-evolve. However, as the actor continues learning, rewarding skill proposals through their reuse in subsequent training steps may conflate skill benefits with actor improvement, while directly testing each proposed skill requires costly additional actor rollouts. In this paper, we introduce UniSkill, which uses a shared policy to interact with the environment and propose skillbank edits (Add, Update, or No Edit) from the resulting trajectories. Specifically, the actor learns from environment rewards, while contrastive action feedback guides skill proposal learning. This feedback provides an actor-alignment signal by measuring how replacing the retrieved skill with a proposed skill changes the current actor's action log-likelihood gap between previously collected successful and failed trajectories from the same task, thereby avoiding new rollouts for each proposal. Since proposal-level feedback may suppress an otherwise appropriate edit operation when the proposed skill content scores poorly, we further apply skill-edit support regularization to preserve exploration. Empirically, UniSkill achieves strong performance, reaching 98.4% success on ALFWorld and 84.7% on WebShop while maintaining stable joint training. Further ALFWorld experiments show that UniSkill remains effective when the shared policy uses a smaller backbone. Our implementation is available at https://github.com/LimOkii/UniSKill.
Problem

Research questions and friction points this paper is trying to address.

LLM agents
skill extraction
policy co-evolution
skill proposal evaluation
joint optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Skill Proposals
Contrastive Action Feedback
Actor-Aligned
Co-evolving Policy
Skill-Edit Regularization
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Yifei Lu
Yifei Lu
Northeastern University, Shenyang, China
Large Language Model
C
Cheng Liu
The Chinese University of Hong Kong
D
Dianzhi Yu
The Chinese University of Hong Kong
H
Hui Xiang
Qiantang Credit, Hangzhou, China; Ant Group, Hangzhou, China
J
Ji Zhang
Qiantang Credit, Hangzhou, China; Ant Group, Hangzhou, China
Y
Yuanchu Xiao
Qiantang Credit, Hangzhou, China; Ant Group, Hangzhou, China
R
Rong Liang
Qiantang Credit, Hangzhou, China; Ant Group, Hangzhou, China