process-level evaluation

Designs and implements evaluation frameworks and benchmarks that assess systems at the process level, including stepwise outputs, template adherence, and overall system behavior. Develops metrics, scoring procedures, and diagnostic analyses to score intermediate steps and final answers and to reveal reasoning failures or correctness issues that final outputs can mask.

process-levelevaluation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.1
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$215K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

A Procedural Framework for Assessing the Desirability of Process Deviations

Jun 13, 2025
MG
Michael Grohs
🏛️ University of Mannheim | SAP Signavio

Existing process conformance checking techniques identify deviations between process executions and models but cannot assess their desirability—i.e., whether they are problematic, acceptable, or beneficial—leading to subjective, inefficient, and non-reproducible manual evaluation. To address this gap, we propose the first structured, reproducible framework for assessing deviation desirability. Grounded in a systematic literature review and semi-structured expert interviews, the framework defines three mutually exclusive desirability categories—problematic, acceptable, and beneficial—each accompanied by actionable recommendations that integrate theoretical conceptualization with frontline practical insights. We empirically validate the framework through task-oriented experiments, demonstrating significant improvements in analysts’ assessment efficiency and inter-rater consistency. Crucially, it maintains comprehensiveness while supporting concise, actionable decision-making. This work provides a methodological foundation for evidence-based process deviation governance.

Assessing desirability of process deviations systematicallyProviding step-by-step framework for deviation categorizationStreamlining manual, subjective desirability evaluations

Identifying Process Improvement Opportunities through Process Execution Benchmarking

Apr 22, 2025
LA
Luka Abb
🏛️ University of Mannheim | SAP Signavio

Existing process mining benchmarks provide only macro-level performance metrics (e.g., throughput time, completion rate), hindering identification of concrete improvement opportunities. To address this, we propose an executable process execution benchmarking method: it aligns event logs from the target organization and benchmark processes based on behavioral similarity, automatically identifying semantically equivalent and substitutable activity units; then constructs a joint feasibility–performance-impact assessment framework to generate ranked, evidence-driven process modification recommendations. This work pioneers the shift from descriptive benchmark analysis to prescriptive, “actionable” improvement guidance. Evaluated across multiple real-world process scenarios, our approach achieves an average throughput time reduction of 12.7%, significantly enhancing both the precision and implementability of process optimization.

Identifies gaps in current process mining benchmarking toolsProposes prescriptive technique for targeted process improvementsRecommends feasible changes based on behavioral similarity analysis

Eval Factsheets: A Structured Framework for Documenting AI Evaluations

Dec 03, 2025
FB
Florian Bordes
🏛️ FAIR | Meta

Current AI evaluation methodologies suffer from a lack of standardized, systematic documentation, severely undermining reproducibility, transparency, and trustworthy decision-making. Method: This paper introduces Eval Factsheets—a novel framework that pioneers the application of structured documentation to AI evaluation. It establishes a five-dimensional taxonomy—encompassing Context, Scope, Structure, Methodology, and Alignment—to uniformly characterize diverse evaluation paradigms, including traditional benchmarks and LLM-as-judge approaches. A taxonomy-guided questionnaire specifies mandatory and recommended fields covering the entire evaluation lifecycle. Contribution/Results: Empirical validation across multiple benchmark cases demonstrates that Eval Factsheets consistently represent heterogeneous evaluation practices, significantly enhancing cross-evaluation comparability, reproducibility, and transparency. The framework provides a foundational, extensible tool for standardizing AI evaluation documentation and practice.

Challenges in reproducibility and transparency due to benchmark proliferation.Lack of systematic documentation standards for AI evaluation methodologies.Need for structured framework to document diverse evaluation paradigms.

A Task Taxonomy for Conformance Checking

Jul 16, 2025
JR
Jana-Rebecca Rehse

Existing visualization tools for compliance checking lack systematic characterization of analytical tasks, hindering rigorous effectiveness evaluation. This paper introduces the first multidimensional task taxonomy specifically designed for compliance checking, modeling core trace-to-model alignment tasks in process mining along six dimensions: objective, method, constraint type, data characteristics, data target, and cardinality. Crucially, this taxonomy explicitly links the semantic requirements of compliance checking with established visual analytics design principles—thereby bridging the semantic gap between process mining and visual analytics. It provides a reusable theoretical framework to rigorously define visualization purposes, evaluate tool effectiveness, and support co-design of analysis systems. As a result, the interpretability and practical utility of complex compliance analysis outcomes are significantly enhanced.

Clarify purposes of diverse conformance checking visualizations.Classify tasks in conformance checking analyses.Enable systematic evaluation of visualization usefulness.

Latest Papers

What's happening recently
View more

Traditional AI evaluation methods are primarily designed for static model selection and often fail to diagnose root causes of performance degradation in production or guide targeted improvements. This work proposes EvalLoop, a novel methodology that embeds evaluation into a continuous optimization loop. By integrating dimensional metric grouping, failure mode categorization, single-variable controlled experiments, and a human-in-the-loop gating mechanism, EvalLoop enables precise mapping from failure attribution to actionable refinements and supports deployment-aware model selection. Evaluated on a sales intelligence briefing generation task, the approach increased overall accuracy of the best-performing model from 82.6% to 94.6%, improved performance on critical dimensions by over 16 percentage points, and reduced human review effort by 94%.

business AI systemsevaluationfailure diagnosis

Current computer-using agent (CUA) benchmarks rely on fragile scripted evaluators that frequently produce erroneous failure judgments, obscuring true performance bottlenecks. This work proposes the first reliability-focused evaluation framework encompassing the entire pipeline—from task construction and trajectory observation to scoring and reporting—and introduces a three-tier failure diagnosis taxonomy. Through manual auditing and attribution analysis of 150 publicly reported failure trajectories, we find that 15.3% of failure labels are incorrect, with 10.7% stemming from evaluator misjudgment and 4.7% arising from task design flaws. Building on these insights, we derive phased design principles for long-horizon CUA evaluation, substantially improving assessment accuracy and interpretability.

benchmarkingcomputer-use agentsevaluation reliability

Hot Scholars

WZ

Wentao Zhang

Institute of Physics, Chinese Academy of Sciences
photoemissionsuperconductivitycupratehtsc
MA

Mengting Ai

PhD Student, University of Illinois, Urabana-Champaign
efficient machine learninglarge language model
CH

Conghui He

Shanghai AI Laboratory
Data-centric AILLMDocument Intelligence
ZZ

Zibin Zheng

IEEE Fellow, Highly Cited Researcher, Sun Yat-sen University, China
BlockchainSmart ContractServices ComputingSoftware Reliability
JC

Jiajun Chai

Meituan Inc.
Reinforcement LearningLLMsAgentic Learning