AdaptEvo: Adaptive Agent Learning with Evolving Supervision
This study addresses the challenges of imperfect supervision signals and dynamically evolving evaluation criteria in rule-driven decision-making by proposing a framework that integrates confidence-adaptive policy optimization with evolutionary decision knowledge. Methodologically, we design the CA-GRPO algorithm to dynamically balance outcome and process rewards, alongside an evolutionary module that distills reusable knowledge from failure cases. Experiments conducted on the Qwen3.6 backbone using a multimodal content moderation dataset demonstrate that the proposed framework significantly outperforms existing baselines on industrial-scale data. Furthermore, it exhibits exceptional robustness under shifting rule conditions, effectively overcoming the limitations inherent in fixed reward-mixing approaches.