Score
Designs and executes systematic evaluations that measure and document an entity’s existing processes, resources, capabilities, configurations, and performance to establish a baseline. Analyzes collected evidence to identify strengths, weaknesses, compliance gaps, and areas requiring change, producing reports or artifacts (for example inventories, process maps, metrics) that guide planning and decision-making.
Existing process mining benchmarks provide only macro-level performance metrics (e.g., throughput time, completion rate), hindering identification of concrete improvement opportunities. To address this, we propose an executable process execution benchmarking method: it aligns event logs from the target organization and benchmark processes based on behavioral similarity, automatically identifying semantically equivalent and substitutable activity units; then constructs a joint feasibility–performance-impact assessment framework to generate ranked, evidence-driven process modification recommendations. This work pioneers the shift from descriptive benchmark analysis to prescriptive, “actionable” improvement guidance. Evaluated across multiple real-world process scenarios, our approach achieves an average throughput time reduction of 12.7%, significantly enhancing both the precision and implementability of process optimization.
Existing process conformance checking techniques identify deviations between process executions and models but cannot assess their desirability—i.e., whether they are problematic, acceptable, or beneficial—leading to subjective, inefficient, and non-reproducible manual evaluation. To address this gap, we propose the first structured, reproducible framework for assessing deviation desirability. Grounded in a systematic literature review and semi-structured expert interviews, the framework defines three mutually exclusive desirability categories—problematic, acceptable, and beneficial—each accompanied by actionable recommendations that integrate theoretical conceptualization with frontline practical insights. We empirically validate the framework through task-oriented experiments, demonstrating significant improvements in analysts’ assessment efficiency and inter-rater consistency. Crucially, it maintains comprehensiveness while supporting concise, actionable decision-making. This work provides a methodological foundation for evidence-based process deviation governance.
Existing research lacks systematic methods to assess how requirements engineering (RE) impacts downstream development activities, hindering RE process optimization. Method: This paper proposes the first fitness-for-purpose RE impact assessment model, integrating a systematic literature review with multi-source empirical data to identify and structure 24 downstream development activities affected by requirements and 16 quantifiable attributes. Contribution/Results: The model bridges two critical gaps in requirements quality assessment—namely, the “activity dimension” and “measurability of impact”—by enabling empirical analysis of how specific requirements artifacts and processes concretely influence development practices. It provides a theoretically grounded framework and evidence-based decision support for precise, targeted optimization of the RE phase.
This study addresses the challenge that relying solely on final outputs fails to capture process-level behavioral drift during the skill evolution of enterprise AI agents. To this end, it proposes a continuous evaluation framework integrating both outcome and process assessments. The method independently computes ground-truth references and designs reusable test templates, combining programmatic checks with constrained LLM judges to enable fine-grained monitoring of tool selection, parameter configuration, and execution order. Furthermore, dependency attribution techniques are introduced to substantially reduce false-positive noise. Experimental results demonstrate that 92.6% of runs passing final numerical checks still exhibit process deviations, while dependency attribution reduces the average number of failed checks from 6.34 to 2.65, effectively revealing differences in specification sensitivity.
This study addresses the fragmentation of evaluation criteria for automated research systems and the difficulty of direct cross-task comparison. Employing a systematic literature review, it comprehensively examines evaluation designs across six task categories, including literature synthesis and ideation. By comparing benchmark construction and scoring protocols, this work proposes a complementary evaluation framework encompassing output-level, process-level, and human-subject assessments. It reveals the capability differences reflected by distinct designs and underscores the critical role of calibration specificity and resource budgets in performance interpretation. Furthermore, the project identifies gaps in diagnostic evaluation and provides recommendations for standardized reporting and auditing. Ultimately, these contributions offer practical guidance for benchmark selection and future research design in evaluating automated scientific discovery systems.
Tool-use agents frequently fail by acting on insufficient evidence or when preconditions in multi-step workflows remain unsatisfied. This study elucidates the mechanisms underlying evidence chain fragmentation from decision-making to execution, highlighting fundamental discrepancies between static evaluation and dynamic execution. To address these issues, this work proposes SafeActBench, a novel benchmark that introduces a provenance-bound evidence ledger and a deterministic trajectory evaluator. Through systematic investigation incorporating multi-model configuration testing, workflow dependency tracing, and evidence integrity verification, the results demonstrate that agent failures fundamentally stem from executing actions without establishing sufficient evidence and from inadequately resolving prerequisite dependencies within complex procedural workflows.
This study addresses the scoring distortion in existing tool-agent benchmarks, which assume correct interface behavior while overlooking defects in underlying implementations. We model tool interfaces as executable contracts and propose an auditing framework that integrates static code analysis with dynamic state tracking to systematically verify consistency between tool implementations and interface declarations, thereby tracing the origins of scoring discrepancies. Evaluations across four mainstream benchmarks reveal seven latent defects, demonstrating the predominance of static analysis in defect detection and the limitations of dynamic verification. Furthermore, our findings expose misjudgments in current evaluators, showing that most reported scores fail to reflect actual state changes. This work establishes a new paradigm for enhancing the reliability of agent evaluation.
本文提出了一种基于集合论的操作方法和AAS适用性模型,以解决不同AAS实例在结构、内容及完整性上的差异问题,从而提高其在特定应用场景中的可比性和适用性。
One aspired outcome of empirical research on quantitative data is a variance theory, i.e., a quantification of the effect of an independent on a dependent variables. The validity of variance theories stems from the synthesis of multiple pieces of evidence, which increases its validity beyond the findings of a single study. However, research synthesis in SE is rare and if done mostly limited to purely narrative syntheses. At best, researchers perform meta-analyses to synthesize variance theories from several quantitative results. But even meta-analyses only produce reliable results when synthesizing exact replications yet fail to generalize from variations. We aim to extend the frontier of research synthesis beyond the state-of-the-art to systematically manage empirical evidence and its evolution. We apply method engineering to construct a framework for research synthesis from proven, individual method fragments. The framework allows researchers to put new evidence in a clear relation to an existing body of evidence and systematically expand knowledge about a studied phenomenon. We demonstrate the application of this framework to two fields of research by explicitly modeling the relationship between existing pieces of evidence. The framework puts three types of evolution of evidence into relation: (1) replications investigate the same hypothesis in a new context to improve external validity, (2) revisions challenge an existing hypothesis to improve internal validity, and (3) reanalyses replace analysis methods to improve conclusion validity. Through a systematic evolution of evidence and clear assessment criteria for each dimension of validity, the proposed framework can determine the frontier of a field of research. The framework provides a perspective to systematically evolve empirical evidence in SE, supporting more constructive and productive advances in our field.
研究通过语言模型解释问题,确定性策略选择并运行预批准分析程序的方法解决企业分析问题,使用关系操作等确保结果可重现。