Score
Designs and implements benchmark suites and evaluation protocols that measure agent or model performance over extended temporal horizons, including ultra-long or continuous runs. Builds task suites, streaming/time-aware instrumentation, and synthetic or real-world scenario generators (e.g., worldlines, implicit‑reward conditions) that produce delayed, noisy, or nonstationary feedback and multilevel signals to enable reproducible comparison of sustained operation, long-term adaptation, and temporal robustness.
Current evaluations of agent tool use often conflate workload specifications, action generation, and evidentiary criteria, lacking a unified and auditable framework. This work proposes an evaluation paradigm centered on “evidence admissibility gating,” which explicitly decouples workloads, drivers, and verification evidence through a shared evidence admissibility contract. The framework integrates diverse environments—including WebArena Verified, a subset of SWE-Gym, and MiniWoB++—and employs a standardized reporting pipeline comprising a universal workload adapter, declarative drivers, task manifests, event schemas, and replay/freeze strategies. It uniformly logs multidimensional metrics such as latency, invalid actions, and patching costs, enabling consistent differentiation of controller performance under identical workloads while ensuring relevance, reproducibility, and auditability in agent evaluations.
Existing reinforcement learning agents often overfit to idiosyncratic patterns in closed environments and lack verifiable behavioral generalization. This work proposes the first cross-domain, long-horizon, multi-tool post-training framework, built upon the open-source MoE model Qwen3.5-122B-A10B and combining two-stage supervised fine-tuning (SFT) with reinforcement learning (RL). Training is conducted on 363 tasks across 27 categories within the MCP benchmark, strictly isolating external evaluation tasks and reward signals. Experimental results demonstrate that the proposed approach substantially enhances out-of-distribution transfer performance, achieving consistent gains across five external benchmarks—including Toolathlon (+9.6 percentage points) and τ²-Bench (+5.3 pp)—and even improves performance on SWE-Bench Pro and Terminal-Bench 2 despite the absence of software engineering tasks in training. The study further uncovers four consistent cross-scenario behavioral divergence patterns.
Existing evaluation of generative AI is hindered by the scarcity of high-quality benchmarks, whose manual construction is costly and time-consuming. Method: We propose the first automated benchmark construction framework powered by collaborative large language model (LLM) agents, decomposing benchmark creation into four sequential stages—planning, generation, verification, and evaluation—integrating task decomposition, agent coordination, human-in-the-loop feedback, and explicit constraint-satisfaction assessment. Contribution/Results: The framework significantly enhances data diversity and metric reliability. Leveraging it, we construct the first high-quality benchmark specifically targeting planning and constraint-satisfaction capabilities in text generation. We systematically evaluate seven state-of-the-art models, uncovering shared failure modes and fine-grained capability disparities. Our work establishes a scalable, reproducible paradigm for evaluating generative AI capabilities, advancing both benchmark methodology and empirical analysis.
Current agent evaluation methods predominantly focus on performance ceilings in static environments, failing to capture robustness in real-world dynamic scenarios—particularly in task scheduling, active exploration, and continual learning. To address this gap, this work proposes EvoEnv, the first dynamic, multi-dimensional evaluation framework tailored to realistic operational settings. EvoEnv simulates “intern” agents continuously exploring and learning within streaming, uncertain environments, assessing capabilities across three dimensions: context-aware scheduling, active information acquisition, and policy generalization. The framework integrates multimodal large language models, dynamic task generation, active exploration mechanisms, and experience distillation to construct a scalable evaluation environment. Experiments demonstrate that state-of-the-art agents exhibit significant performance degradation under dynamic conditions, underscoring EvoEnv’s effectiveness and necessity in evaluating real-world deployment reliability.
Existing long-horizon benchmarks merely show that agent performance degrades as task length increases, yet they cannot distinguish whether this decline stems from the intrinsic difficulty of extended tasks or from error accumulation across stages. This work introduces the "horizon residual" metric, which quantifies the additional difficulty beyond what is attributable to compounding errors by comparing the actual success rate on full-length tasks against a baseline predicted from short-segment performance. We formally define this concept for the first time and establish a comparable short-task baseline framework incorporating trajectory-induced degradation analysis, context decay modeling, and log-ratio metrics. Our approach emphasizes the necessity of predefined stage segmentation and resource allocation to control confounding variables, providing an attribution tool for long-horizon evaluation and demonstrating that declining aggregate success rates alone are insufficient evidence of length-specific challenges.
This work addresses the absence of benchmarks evaluating language model agents’ ability to consistently adhere to complex, constraint-laden instructions—such as corporate policy manuals—over long contexts and multi-turn tool interactions. The authors introduce the first benchmark for this challenge, comprising 65 tasks across five professional domains, which requires agents to operate within a simulated office environment (e.g., email, calendar, chat) guided by dynamic policy manuals ranging from 20 to 124 pages. Leveraging expert-authored, non-redundant manuals and 824 deterministic scoring rules, the benchmark enables fully automated, stringent evaluation where all criteria must be satisfied. Experiments reveal that even the best-performing configuration among 30 state-of-the-art models passes only 36.2% of tasks, with most scoring below 25%, exposing systemic deficiencies in policy compliance and behavioral consistency.
Current evaluations rely solely on final scores, failing to capture the process dynamics, knowledge reuse efficacy, and decision evolution of agents in long-horizon research and development tasks. This work proposes the first fine-grained evaluation framework, integrating rule-driven behavioral metrics, structured analysis, and cross-task controlled experiments to systematically assess seven state-of-the-art models across 36 tasks. The study reveals that existing agents behave more like engineering optimizers than autonomous researchers: while capable of producing viable solutions, they exhibit poor inter-run stability, and their innovations largely consist of recombinations of known methods rather than genuine methodological breakthroughs. Furthermore, this work identifies, for the first time, key factors underlying performance instability.
This work addresses the limitations of traditional time series methods, which are constrained by fixed forecasting horizons and struggle to support contextual reasoning, tool invocation, and structured decision-making in real-world scenarios. The authors propose AION, a framework that formalizes time series tasks as a triplet of task specification, workspace, and validation interface, integrating six core modules—agents, skills, rules, memory, evaluation, and protocols—to emphasize temporal grounding, knowledge-guided reasoning, and reliability assurance. By incorporating process traceability, multi-level auditing, and post-hoc experimental analysis, AION overcomes the constraints of static evaluation paradigms. In a Kaggle store sales forecasting case study, AION substantially outperforms direct modeling approaches, generating richer reasoning traces, intermediate artifacts, and audit steps, thereby demonstrating its effectiveness and superiority in handling complex, real-world time series tasks.
Existing AI agent benchmarks primarily focus on short-duration tasks, making them inadequate for evaluating agents’ long-horizon planning, extended-context comprehension, and memory capabilities in software engineering. This work proposes SWE-Marathon—the first benchmark specifically designed for ultra-long-horizon software engineering tasks spanning multiple hours and involving tens of millions of tokens. It comprises 20 tasks, each equipped with an executable environment, human-authored reference solutions, and multi-layered validation mechanisms. Through adversarial testing, shortcut-prevention designs, and trajectory analysis, the study systematically uncovers critical deficiencies in state-of-the-art coding agents, particularly in self-verification, task persistence, and resistance to reward hacking—evidenced by a completion rate below 30% and reward hacking observed in 13.8% of trials. The benchmark, evaluation code, and agent trajectories are publicly released to advance research on long-horizon autonomous agents.