benchmark vision-language-action models

Designs, implements, and analyzes standardized evaluation suites and protocols for models that map visual inputs and language to actions; this includes defining representative manipulation or control tasks, specifying embodiment and uncertainty settings, creating metrics and test splits, and running comparative evaluations of multiple policies.

benchmarkvision-language-actionmodels

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.02
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Existing GUI automation testing tools suffer from high maintenance overhead, fragility under UI changes, and tight coupling with specific UI frameworks. To address these challenges, this paper proposes a decoupled testing methodology integrating the ViewModel architectural pattern with Behavior-Driven Development (BDD). It pioneers the synergistic combination of a projectional Domain-Specific Language (DSL)—implemented using JetBrains MPS—with ViewModel and BDD for GUI test modeling. The DSL enables declarative specification and validation of presentation logic independently of the underlying UI framework (e.g., JavaFX), thereby fully separating test specifications from implementation details. Evaluated on a task manager case study, the approach automatically generates executable test code, substantially reducing specification authoring effort and eliminating test breakage caused by UI refactoring. Empirical evaluation demonstrates significant improvements in test maintainability, robustness against interface evolution, and practical applicability.

Automated GUI testing faces high specification effortsExisting tools struggle with maintenance and flaky testsTesting presentation logic is coupled to GUI frameworks

This work addresses the absence of benchmarks evaluating language model agents’ ability to consistently adhere to complex, constraint-laden instructions—such as corporate policy manuals—over long contexts and multi-turn tool interactions. The authors introduce the first benchmark for this challenge, comprising 65 tasks across five professional domains, which requires agents to operate within a simulated office environment (e.g., email, calendar, chat) guided by dynamic policy manuals ranging from 20 to 124 pages. Leveraging expert-authored, non-redundant manuals and 824 deterministic scoring rules, the benchmark enables fully automated, stringent evaluation where all criteria must be satisfied. Experiments reveal that even the best-performing configuration among 30 state-of-the-art models passes only 36.2% of tasks, with most scoring below 25%, exposing systemic deficiencies in policy compliance and behavioral consistency.

agentic instruction followingbenchmarklong-context

SpecEval: Evaluating Model Adherence to Behavior Specifications

Sep 02, 2025
AA
Ahmed Ahmed
🏛️ Stanford University | Virginia Tech

A systematic audit of whether foundation models actually adhere to their developers’ published behavioral guidelines remains absent. Method: We propose a tripartite consistency evaluation framework that (i) parses behavioral statements via natural language processing, (ii) generates targeted prompts, and (iii) leverages the model itself as an internal judge to assess consistency among *guideline*, *model output*, and *self-judgment*—extending beyond conventional generator-verifier paradigms. Contribution/Results: The framework enables the first large-scale, cross-organizational, automated compliance audit across 16 models from six developers, covering over 100 behavioral statements. Experiments reveal up to a 20% compliance gap, exposing pervasive and substantial systemic inconsistencies in guideline adherence. This work delivers a scalable, empirically grounded assessment tool for AI governance and responsible model deployment.

Auditing model adherence to provider behavior specificationsEstablishing three-way consistency between specifications, outputs, and judgmentsIdentifying systematic compliance gaps across multiple foundation models

Latest Papers

What's happening recently
View more

Existing benchmarks primarily assess language models on localized programming tasks, failing to capture their capability to construct complete software systems from scratch. This work introduces ProgramBench, the first end-to-end evaluation framework grounded in behavioral equivalence, which requires agents to autonomously design and implement full codebases based solely on program specifications and documentation, with correctness verified through behavioral test suites. The benchmark encompasses 200 real-world software tasks—including CLI tools, FFmpeg, and SQLite—supports unconstrained, open-ended code generation, and incorporates agent-driven fuzz testing to automatically produce behavioral test cases. Evaluation across nine prominent language models reveals that none can fully solve any task; the best-performing model passes 95% of tests on only 3% of tasks and tends to generate single-file implementations structurally divergent from human-written code.

benchmarkingcode generationlanguage models

This study evaluates the capacity of general-purpose multimodal large models to translate visual understanding into embodied manipulation. We propose a "code-as-policy" agent framework that requires no fine-tuning, dedicated perception modules, or predefined policies. By leveraging only a shared robot API and RGB visual feedback, the framework prompts large models to generate and execute manipulation code, enabling end-to-end decision-making with closed-loop iterative correction. Experimental results demonstrate that the optimal configuration successfully completes 22 out of 25 tasks, achieving an average success rate of 73.3%. Furthermore, this work systematically identifies critical failure modes, such as spatial misalignment. Overall, these findings provide a valuable benchmark and empirical evidence for advancing general-model-driven embodied intelligence.

Agentic ReasoningBenchmark EvaluationCode-as-Policy

This work addresses the lack of a general, auditable dynamic control mechanism in existing training systems, which typically rely on framework-specific code. The authors propose the first cross-framework, open-source control plane that exposes training interfaces through a unified protocol, integrating declarative configuration, request validation, and secure control-point scheduling within the Aim workspace to enable metric monitoring, real-time intervention, and operational traceability. The system supports safe human and automated controller interventions during training while fully logging all operational trajectories. Experiments across five NLP and reinforcement learning tasks demonstrate its effectiveness, and the open-source implementation provides a foundation for reproducible human-in-the-loop training.

auditable trainingcontrol planehuman-in-the-loop

This study addresses the challenges of adaptability under varying tasks and equipment states, as well as human-review traceability in automated systems. We propose a multi-line task adjustment system integrating local large language models, digital twins, and human-machine collaboration. The core innovation lies in a "propose-verify-decide" workflow that combines structured requirement parsing with multi-source record tracing, establishing an end-to-end closed loop from intent translation and strategy generation to simulation-based validation, thereby ensuring semantic correctness and operational compliance at each stage. Experimental results demonstrate that all 18 test cases met expected outcomes, achieving an 87.5% interception rate for invalid inputs and full approval in engineering reviews, with an average initial response time of only 12.94 seconds.

Digital TwinHuman-AI CollaborationLarge Language Model

Existing embodied intelligence approaches struggle to reliably learn and modularize heterogeneous hierarchical capabilities. This work proposes a “capability externalization” framework that decouples perception, reasoning, planning, and control into independently optimizable tools, orchestrated through a novel Embodied Tool Protocol (ETP) enabling tool registration, discovery, and dynamic invocation. The authors construct a library of over 100 tools and introduce EmbodiedToolBench, a benchmark for cross-platform evaluation in both simulation and real-world settings. Experiments demonstrate average performance improvements of 31% on EB-ALFRED and 36% on EB-Navigation, with substantial gains in cognitive and perceptual tasks, while execution-oriented tasks show more limited progress—highlighting that the timing, selection, and composition of tool usage remain critical challenges.

capability externalizationembodied intelligenceheterogeneous capabilities

Hot Scholars

MS

Muhammad Shafique

Professor, ECE, New York University (AD-UAE, Tandon-USA), Director eBRAIN Lab
Embedded Machine LearningBrain-Inspired ComputingRobust & Energy-Efficient System DesignSmart
JN

Jingwei Ni

Doctoral Researcher in NLP, ETH Zurich
NLP for social goodclaim verificationcausal NLPcomputational social science
SX

Shicheng Xu

Institute of Computing Technology, Chinese Academy of Sciences
Agentic LLMsTool UsingReasoningKnowledge
YY

Yifan Yang

Senior Research SDE, Microsoft Research Asia
Multi-modalityComputer VisionMachine LearningArtificial Intelligence
ZT

Zhiqiang Tao

Assistant Professor, Rochester Institute of Technology
Machine LearningData MiningDeep LearningComputer Vision