Score
Designs and implements mechanisms for parametrically varying and controlling task difficulty to create graded challenges. Builds and calibrates task generators, curricula, progression rules, and difficulty-annotated evaluation suites that match task demands to agent skill levels and measure performance across difficulty levels.
To address the limited adaptability and long-horizon decision-making capabilities of large language model (LLM)-based agents in complex, realistic, interactive environments, this paper proposes the Generate-Execute-Feedback (GEF) loop framework. It is the first to systematically analyze the environment’s role—centered on the environment itself—across three phases: task generation, dynamic execution, and multi-granularity feedback. The work innovatively unifies fragmented environment extension approaches into a coherent analytical framework, integrating reinforcement learning paradigms, automated task generation, dynamic environment modeling, and rollout-based evaluation with feedback. We design multi-stage environment extension strategies and a corresponding benchmarking suite. Experiments demonstrate that the GEF framework significantly improves agents’ experience-efficient learning and long-term planning performance. It clarifies a scalable, environment-driven pathway for agent capability advancement, offering both theoretical foundations and practical guidance for embodied intelligence and autonomous agent research.
This study addresses the limited scalability of agent-executable environments and the skill-to-task translation gap by proposing a capability-oriented task synthesis framework. The approach introduces a novel capability-requirement-guided environment synthesis mechanism that leverages difficulty modes to translate low-level skills into task blueprints, jointly generating instructions, environments, and evaluators. Furthermore, it incorporates a solver-feedback-driven iterative hardening technique to dynamically escalate task challenge levels. Experimental results demonstrate that this framework achieves consistent performance improvements across diverse agent benchmarks, effectively validating its superiority in automatically constructing high-quality, progressive training environments.
This work addresses the limitation in existing mobile GUI agent training data, which lacks fine-grained control over task difficulty, often resulting in a mismatch between task complexity and agent capabilities that hinders effective learning. To overcome this, the authors propose MobileGen, a novel framework that decouples task difficulty into structural and semantic dimensions for the first time. MobileGen employs a multi-agent controllable generator to dynamically model the agent’s capability boundary and adaptively synthesizes high-quality interaction trajectories and task instructions aligned with the agent’s current proficiency through distribution-aware sampling. This enables difficulty-adaptive curriculum learning tailored to mobile GUI environments. Experimental results demonstrate that the proposed approach improves agent performance by an average of 1.57× across multiple challenging benchmarks, significantly outperforming existing data generation strategies.
This study addresses the challenges of uncontrollable difficulty and high refresh costs in browser agent benchmarks by proposing a method that reframes task difficulty as a programmable property. Through controlled environmental interventions, deterministic, detectable, and recoverable state perturbations are introduced across different layers of the web stack while preserving user instructions and success criteria unchanged. These perturbations are further combined with cognitive primitive annotations to construct challenging tasks. Experiments conducted on self-hosted websites using an automated evaluation framework demonstrate that this approach reduces the average pass rate of agents by 22.9% and reveals that 75% of failures stem from belief errors. The associated code and datasets have been made publicly available.
Existing GUI task difficulty metrics predominantly rely on motor-based indicators (e.g., step count), neglecting cognitive load. Method: We propose the “Cognitive Chain” framework, decomposing pre-execution cognition into three stages—discovery, decision-making, and computation—and quantify the difficulty of each stage using information-theoretic measures. Leveraging large language models, we automatically extract cognitive chains from user interaction traces to construct a computable Cognitive Difficulty Index (CDI). Contribution/Results: This work establishes the first cognitively grounded model for GUI task difficulty, moving beyond behaviorist paradigms. Empirical evaluation demonstrates that CDI significantly predicts per-step completion time (R² = 0.46) and reveals substantial performance degradation in state-of-the-art GUI agents under high cognitive load—validating both theoretical rigor and practical utility for HCI and AI-agent evaluation.
This study addresses the capability degradation and goal-contract invalidation of AI agents caused by shifting conditions during model transfer, cross-domain deployment, or scaling. To mitigate these issues, we propose a three-layer interactive calibration methodology that decouples trainable policies from frozen configurations, establishing a multidimensional constraint system encompassing information preservation, execution framework adaptation, and user acceptance. Technically, the approach integrates semantic checkpoint repair, tool substitution, local replanning, and output contract enforcement mechanisms. For evaluation, we introduce a factorial testing holdout system to prevent aggregated gains from masking localized failures. This work provides a methodological framework enabling the coexistence of shared standards and market-specific adapters in global e-commerce scenarios, with empirical validation reserved for future research.
研究通过分析CoderForge-Preview数据集中的任务特征,使用集成方法、SHAP归因和效应量分析,探究了软件问题解决任务的难度,并发现难度主要受补丁碎片化和仓库规模影响。
This work addresses the absence of benchmarks evaluating language model agents’ ability to consistently adhere to complex, constraint-laden instructions—such as corporate policy manuals—over long contexts and multi-turn tool interactions. The authors introduce the first benchmark for this challenge, comprising 65 tasks across five professional domains, which requires agents to operate within a simulated office environment (e.g., email, calendar, chat) guided by dynamic policy manuals ranging from 20 to 124 pages. Leveraging expert-authored, non-redundant manuals and 824 deterministic scoring rules, the benchmark enables fully automated, stringent evaluation where all criteria must be satisfied. Experiments reveal that even the best-performing configuration among 30 state-of-the-art models passes only 36.2% of tasks, with most scoring below 25%, exposing systemic deficiencies in policy compliance and behavioral consistency.
This study addresses the scarcity and synthesis difficulty of training data for skill invocation in large language model (LLM) agents by proposing an automated data synthesis and training framework. The framework introduces a novel method for automatically constructing offline skill environments with controllable difficulty levels and built-in execution verifiers. By integrating web crawling with a builder-reviewer pipeline, it enables large-scale synthesis of high-quality trajectories, followed by agent training via supervised fine-tuning (SFT) and automated verification mechanisms. Experimentally, the approach constructs 6.8k environments and 19k trajectories, improving the skill reading rate of the Qwen3.5-9B model from 28% to 96%. Notably, the resulting trained model surpasses the performance of an untrained 397B-parameter baseline.
This study addresses the challenge of filtering erroneous and non-transferable knowledge from skill libraries of self-evolving agents by proposing a Proposer-Builder-Verifier architecture to validate skill reusability in unseen tasks. The method dynamically synthesizes test scenarios through conditional constraint generation, overcoming the limitations of traditional static evaluation. Furthermore, it employs an automated execution comparison mechanism to quantitatively assess skill utility and efficiency, enabling reliable retention or rejection decisions. Experimental results demonstrate that the proposed framework significantly improves downstream task performance and execution efficiency on the ALFWorld and WebShop benchmarks, while also confirming the accuracy of its skill reusability assessment.
This study addresses the tendency of agents to ignore multi-stage procedural requirements by focusing solely on final outcomes, proposing a "grounded skill following" framework. This approach formalizes expert skills as runtime-aware contracts and employs progress-credit reward verification alongside state feedback mechanisms to guide agents in strictly adhering to each required execution stage. Built upon the Qwen3.5-4B model and integrated with reinforcement learning, the framework enables grounded decision-making from environmental observations and contract state tracking. Experimental results demonstrate that the proposed method achieves protocol completion rates of 99.27% and 99.96% on mathematical and search tasks, respectively, while significantly improving task success rates. Ultimately, this work realizes verifiable procedural execution for intelligent agents.