Score
Designs, implements, and analyzes standardized evaluation suites and protocols for models that map visual inputs and language to actions; this includes defining representative manipulation or control tasks, specifying embodiment and uncertainty settings, creating metrics and test splits, and running comparative evaluations of multiple policies.
Current vision-language-action (VLA) frameworks lack unified learning paradigms to support robust, generalizable, and interpretable instruction-driven robotic manipulation. Method: We conduct a systematic literature review of 102 models, 26 datasets, and 12 simulation platforms. Based on semantic richness and cross-modal alignment, we propose a novel two-dimensional dataset characterization framework—quantifying critical distributional gaps in existing VLA data for the first time. We further synthesize prevailing architectural paradigms, identifying three key technical pathways: scalable pretraining, modular embodied control, and cross-modal alignment. Contribution/Results: This work establishes a theoretical foundation and empirical basis for universal policy learning in embodied AI. It enables principled dataset curation, architecture design, and evaluation, advancing VLA systems toward greater robustness, generalization, and interpretability.
本文提出两种技术,通过增加推理时间和模拟真实部署环境来提高对齐评估的真实性,解决模型在测试与实际部署中表现差异的问题。
Existing GUI automation testing tools suffer from high maintenance overhead, fragility under UI changes, and tight coupling with specific UI frameworks. To address these challenges, this paper proposes a decoupled testing methodology integrating the ViewModel architectural pattern with Behavior-Driven Development (BDD). It pioneers the synergistic combination of a projectional Domain-Specific Language (DSL)—implemented using JetBrains MPS—with ViewModel and BDD for GUI test modeling. The DSL enables declarative specification and validation of presentation logic independently of the underlying UI framework (e.g., JavaFX), thereby fully separating test specifications from implementation details. Evaluated on a task manager case study, the approach automatically generates executable test code, substantially reducing specification authoring effort and eliminating test breakage caused by UI refactoring. Empirical evaluation demonstrates significant improvements in test maintainability, robustness against interface evolution, and practical applicability.
This work addresses the absence of benchmarks evaluating language model agents’ ability to consistently adhere to complex, constraint-laden instructions—such as corporate policy manuals—over long contexts and multi-turn tool interactions. The authors introduce the first benchmark for this challenge, comprising 65 tasks across five professional domains, which requires agents to operate within a simulated office environment (e.g., email, calendar, chat) guided by dynamic policy manuals ranging from 20 to 124 pages. Leveraging expert-authored, non-redundant manuals and 824 deterministic scoring rules, the benchmark enables fully automated, stringent evaluation where all criteria must be satisfied. Experiments reveal that even the best-performing configuration among 30 state-of-the-art models passes only 36.2% of tasks, with most scoring below 25%, exposing systemic deficiencies in policy compliance and behavioral consistency.
本文探讨了大型语言模型在软件工程中基于测试的方法,通过分析87篇研究文献,区分并比较了不同测试驱动任务的特点和机制,提出了未来研究方向。
A systematic audit of whether foundation models actually adhere to their developers’ published behavioral guidelines remains absent. Method: We propose a tripartite consistency evaluation framework that (i) parses behavioral statements via natural language processing, (ii) generates targeted prompts, and (iii) leverages the model itself as an internal judge to assess consistency among *guideline*, *model output*, and *self-judgment*—extending beyond conventional generator-verifier paradigms. Contribution/Results: The framework enables the first large-scale, cross-organizational, automated compliance audit across 16 models from six developers, covering over 100 behavioral statements. Experiments reveal up to a 20% compliance gap, exposing pervasive and substantial systemic inconsistencies in guideline adherence. This work delivers a scalable, empirically grounded assessment tool for AI governance and responsible model deployment.
Existing benchmarks primarily assess language models on localized programming tasks, failing to capture their capability to construct complete software systems from scratch. This work introduces ProgramBench, the first end-to-end evaluation framework grounded in behavioral equivalence, which requires agents to autonomously design and implement full codebases based solely on program specifications and documentation, with correctness verified through behavioral test suites. The benchmark encompasses 200 real-world software tasks—including CLI tools, FFmpeg, and SQLite—supports unconstrained, open-ended code generation, and incorporates agent-driven fuzz testing to automatically produce behavioral test cases. Evaluation across nine prominent language models reveals that none can fully solve any task; the best-performing model passes 95% of tests on only 3% of tasks and tends to generate single-file implementations structurally divergent from human-written code.
This study evaluates the capacity of general-purpose multimodal large models to translate visual understanding into embodied manipulation. We propose a "code-as-policy" agent framework that requires no fine-tuning, dedicated perception modules, or predefined policies. By leveraging only a shared robot API and RGB visual feedback, the framework prompts large models to generate and execute manipulation code, enabling end-to-end decision-making with closed-loop iterative correction. Experimental results demonstrate that the optimal configuration successfully completes 22 out of 25 tasks, achieving an average success rate of 73.3%. Furthermore, this work systematically identifies critical failure modes, such as spatial misalignment. Overall, these findings provide a valuable benchmark and empirical evidence for advancing general-model-driven embodied intelligence.
This work addresses the lack of a general, auditable dynamic control mechanism in existing training systems, which typically rely on framework-specific code. The authors propose the first cross-framework, open-source control plane that exposes training interfaces through a unified protocol, integrating declarative configuration, request validation, and secure control-point scheduling within the Aim workspace to enable metric monitoring, real-time intervention, and operational traceability. The system supports safe human and automated controller interventions during training while fully logging all operational trajectories. Experiments across five NLP and reinforcement learning tasks demonstrate its effectiveness, and the open-source implementation provides a foundation for reproducible human-in-the-loop training.
This study addresses the challenges of adaptability under varying tasks and equipment states, as well as human-review traceability in automated systems. We propose a multi-line task adjustment system integrating local large language models, digital twins, and human-machine collaboration. The core innovation lies in a "propose-verify-decide" workflow that combines structured requirement parsing with multi-source record tracing, establishing an end-to-end closed loop from intent translation and strategy generation to simulation-based validation, thereby ensuring semantic correctness and operational compliance at each stage. Experimental results demonstrate that all 18 test cases met expected outcomes, achieving an 87.5% interception rate for invalid inputs and full approval in engineering reviews, with an average initial response time of only 12.94 seconds.
Existing embodied intelligence approaches struggle to reliably learn and modularize heterogeneous hierarchical capabilities. This work proposes a “capability externalization” framework that decouples perception, reasoning, planning, and control into independently optimizable tools, orchestrated through a novel Embodied Tool Protocol (ETP) enabling tool registration, discovery, and dynamic invocation. The authors construct a library of over 100 tools and introduce EmbodiedToolBench, a benchmark for cross-platform evaluation in both simulation and real-world settings. Experiments demonstrate average performance improvements of 31% on EB-ALFRED and 36% on EB-Navigation, with substantial gains in cognitive and perceptual tasks, while execution-oriented tasks show more limited progress—highlighting that the timing, selection, and composition of tool usage remain critical challenges.