Qwen-Planner-Agent: A Closed-Loop AI-for-AI Framework for Real-World Mobile Planner Agents

πŸ“… 2026-09-24
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the bottlenecks of low reliability in long-horizon tasks and high real-device interaction costs for mobile agents by proposing a closed-loop AI-for-AI framework. Methodologically, it introduces CARE reward engineering to reduce inference overhead, alongside a human-in-the-loop data flywheel and a model-environment co-evolution mechanism. Efficient iterative training is achieved through supervised cold-start initialization, online reinforcement learning in hybrid environments, action-feedback verification contracts, and memory-skill-tool orchestration. Experimental results demonstrate that the proposed approach achieves state-of-the-art performance on MobilePA-Bench, significantly enhancing tool invocation and multi-step coordination capabilities while preserving general-purpose generalization.
πŸ“ Abstract
The rapid progression of large language models is extending AI from passive content generation into the active workflows of engineering and scientific discovery. This shift raises a compelling question: can AI be both the object of development and an active participant in building next-generation AI systems? We explore this question by building Qwen-Planner-Agent within a closed-loop AI-for-AI framework for scalable development and iterative improvement. Mobile planning offers a demanding test of this approach: complex, long-horizon tasks challenge agent reliability, while costly real-device interaction limits development scalability. The framework connects data production, model training, and deployment through a shared action-feedback-verification contract. (i) AI for Data builds a human-gated agentic data flywheel in which specialized agents construct tasks, collect interaction trajectories, curate and balance training data, and use training feedback to guide subsequent data generation. (ii) AI for Training combines a supervised planning cold start with hybrid-environment online agentic reinforcement learning, where we introduce Competence-Aware Reward-and-Advantage Engineering (CARE) to reduce reasoning and tool-use costs while preserving task performance. (iii) AI drives model--harness co-evolution through an execution-evidence-driven loop that orchestrates memory, skills, and tools at runtime and feeds structured action feedback and preserved failure traces back into coordinated model and harness adaptation. Qwen-Planner-Agent achieves the best overall performance among all evaluated models and systems on MobilePA-Bench, improving over its base model across tool use, memory, skills, and sub-agent coordination. Further evaluations of our model show improvements across non-mobile agentic benchmarks while largely preserving general capabilities.
Problem

Research questions and friction points this paper is trying to address.

AI-for-AI
mobile planner agents
long-horizon tasks
scalable development
agent reliability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Closed-loop AI-for-AI framework
Agentic data flywheel
Competence-Aware Reward-and-Advantage Engineering (CARE)
Online agentic reinforcement learning
Model-harness co-evolution