Score
Designs and builds evaluation artifacts—behavioral test suites, stratified benchmarks, judging protocols, and quantitative outcome metrics—that measure observable behaviors of AI systems (including LLMs) across targeted scenarios. Analyzes and reports behavior-stratified results, producing grounded, auditable judgments and measurement protocols to resolve ambiguous relevance, support alignment with user preferences, and guide system improvements.
Current LLM agent evaluation lacks a systematic framework, particularly neglecting enterprise-specific requirements such as role-based access control, regulatory compliance, and long-horizon interactive behavior. To address this gap, we conduct a comprehensive literature review and propose the first two-dimensional evaluation taxonomy: one axis captures *objective dimensions*—including behavior, capability, reliability, and security—while the other captures *process dimensions*—encompassing interaction paradigms, benchmark datasets, evaluation metrics, and toolchains. Crucially, our taxonomy explicitly incorporates enterprise challenges, exposing critical shortcomings in existing work regarding holistic coverage, scalability, and real-world applicability. The framework provides researchers and practitioners with a structured, actionable reference for designing, evaluating, and deploying LLM agents in complex, mission-critical environments. It advances the field toward trustworthy, production-ready LLM agent systems. (138 words)
This work addresses the fragmented and model-centric nature of existing evaluation methods for large language model (LLM) agents, which often overlook the influence of architectural components—such as planners, memory modules, and tool routers—on agent behavior, resulting in assessments that lack diagnostic precision and specificity. To bridge this gap, the paper proposes a lightweight, architecture-aware evaluation framework that systematically establishes the first explicit mapping between internal agent components, observable behaviors, and evaluation metrics. This approach shifts the paradigm from black-box assessment toward interpretable, component-level diagnosis. Through architecture-aware analysis, behavior-component modeling, and tailored metric design, the framework is validated on real-world LLM agents, demonstrating significant improvements in evaluation transparency, target specificity, and practical utility.
Current personality assessments of large language models (LLMs) predominantly rely on first-person self-report questionnaires, which are susceptible to prompt perturbations and lack behavioral grounding. This work proposes a behavior-data (B-data) framework grounded in contextualized scenarios, employing 3,200 contrastive behavioral situations to capture stable behavioral patterns of LLMs across diverse interaction contexts. It introduces the first behavior-mode axis (BMA) derived from chain-of-thought reasoning, enabling precise modulation of LLM behavioral styles within an activation space. Integrating established psychometric scales (BFI-2, DOSPERT, HEXACO) with behavioral trajectory analysis, the study demonstrates that LLMs exhibit model-specific, context-dependent yet stable behavioral profiles. Furthermore, it shows that the chain-of-thought–derived BMA significantly enhances both the stability and mechanistic fidelity of behavioral control compared to response-chain approaches.
To address unexpected behavioral shifts in language models (LMs) following fine-tuning or deployment, this paper introduces Behavioral Shift Auditing (BSA), a continuous monitoring framework. BSA operates without access to model parameters or gradients, and—uniquely—establishes the first unsupervised, statistical hypothesis testing framework for text generation comparison, leveraging the Kolmogorov–Smirnov test and bootstrap resampling to reliably detect distributional shifts in critical capabilities such as toxicity and translation. The method provides theoretically grounded false positive control and supports configurable tolerance thresholds to accommodate diverse application scenarios. Experiments demonstrate that BSA achieves stable detection of significant behavioral shifts using only hundreds of samples, attaining high sensitivity and low false positive rates on both toxicity and machine translation tasks. Overall, BSA establishes a lightweight, robust, and interpretable paradigm for continuous auditing of LM behavioral evolution.
This work addresses a critical gap in evaluating tool-augmented large language models (LLMs), as existing metrics predominantly emphasize linguistic alignment or task success while overlooking the structural relationship between linguistic signals and executable actions across varying autonomy architectures. To remedy this, the study proposes a behavior-centric evaluation framework grounded in the execution layer, introducing a two-dimensional action–refusal (A–R) space defined by action rate (A) and refusal signals (R), along with a divergence metric (D) to quantify their coordination. Systematic experiments across four canonical scenarios and three autonomy configurations—direct execution, planning, and reflection—reveal significant behavioral distributional differences: reflective scaffolding consistently increases refusal rates in high-risk contexts, yet models exhibit structurally heterogeneous redistribution patterns. By replacing scalar safety scores with separable behavioral dimensions, this approach enables fine-grained, comparable, and interpretable characterization of tool-augmented LLM behaviors.
Evaluating LLM-based agents is challenging due to their dynamic, probabilistic, and continuously evolving nature—traditional predefined benchmarks fail to capture open-ended behaviors, emergent outcomes, and lifecycle adaptation. Method: We propose an evaluation-driven agent development paradigm, integrating online runtime evaluation with offline reconstruction evaluation. Our hybrid framework enables real-time feedback injection, human-AI collaborative closed-loop refinement, and iterative optimization across the full stack (pipeline, architecture, and LLM), incorporating both human and AI evaluators. Contribution/Results: We introduce the first evaluation-centric process model and reference architecture for LLM agent development. It uniquely supports open-behavior capture, emergent-result governance, and dynamic alignment. Experiments demonstrate that the framework effectively enables safe, controllable, and continuous agent iteration under objective drift, requirement changes, and regulatory evolution—achieving robust adaptability without compromising reliability or compliance.
This work addresses the limitations of existing evaluation methods that focus solely on final outcomes, which fail to distinguish reliable reasoning from accidental success or diagnose process-level flaws in long-horizon tasks. To this end, we propose ClawTrack, a dual-dimensional evaluation framework that jointly assesses task completion (Task Score) and reasoning process quality (Process Score). Spanning 320 tasks across eight domains, ClawTrack introduces fine-grained, stepwise scoring along four dimensions, enabling the first interpretable, process-level evaluation of autonomous agent reasoning trajectories. Our Process Grader combines rule-based logic with large language models, incorporating 12,541 task-specific scoring criteria and integrating over 25 deterministic simulation environments. Validation across 21 models and more than 16,000 trials demonstrates that process scores effectively attribute success or failure, filter out spurious successes, and—when used to select high-quality reasoning trajectories—significantly boost performance across model scales, with consistent results across different evaluator models.
This study addresses the limited generalizability of behavior rules for software engineering agents derived from single-framework studies. Through a large-scale experimental ecosystem encompassing 126 agent configurations, 43 frameworks, and 64,380 SWE-bench runs, the authors systematically investigate the relationship between behavioral signals and problem-solving performance by controlling either the large language model (LLM) or the framework layer. Employing variance decomposition, behavioral feature statistics, and directional consistency analysis, they find that framework-level factors explain behavioral differences more significantly than LLM choice. Notably, in nearly half of the configurations, key behavioral signals—such as error rates—exhibit opposing effects across frameworks, revealing for the first time that identical behaviors can carry divergent or even contradictory semantic interpretations depending on the framework. These findings challenge the assumption of cross-framework universality in single-framework-derived rules and underscore the necessity of multi-framework validation.
This work proposes a “layered attribution” diagnostic framework to disentangle the origins of inscrutable behaviors exhibited by AI agents in complex social systems, which are often conflated between internal representations and external constraints. The framework systematically distinguishes a foundational computational layer—encompassing architecture, memory, and perception—from a behavioral modulation layer comprising identity, goals, social interactions, and institutional constraints, thereby integrating representation learning, multi-agent modeling, and institutional analysis into a unified two-tier diagnostic architecture. It yields three key insights: behavioral substitutability validity hinges on the coupling among model, task, and layer; human–AI behavioral discrepancies can serve as diagnostic signals; and effective governance presupposes precise source attribution. This approach establishes a theoretical foundation for interpreting and governing AI behavior.
Current AI evaluation predominantly relies on static responses, which inadequately capture the systematic capabilities of large language models in dynamic scenarios involving tool use, environmental interaction, and multi-agent collaboration. This work proposes “interactive evaluation” as a distinct paradigm, introducing a two-axis taxonomy to clarify its design principles and reporting standards. By leveraging trajectory modeling and a multi-dimensional scoring mechanism, the framework extracts evidence from interaction processes to holistically assess model performance across dimensions such as procedural fidelity, recoverability, coordination, robustness, and system-level efficacy. The study delineates core challenges in interactive evaluation and redefines the logical mapping from empirical evidence to performance judgments, thereby establishing a theoretical foundation for a unified, comparable, and interpretable next-generation AI evaluation framework.