Score
Designs and builds analytical frameworks and quantitative diagnostic models to assess an organization’s operational, financial, and strategic performance; analyzes data, processes, and KPIs to identify root causes of performance gaps, risks, and improvement opportunities. Produces diagnostic findings and prioritized recommendations that guide corrective actions and enable tracking of recovery or improvement initiatives.
Business process optimization remains challenging due to fragmented methodologies across process mining, predictive process monitoring, and process-aware recommendation—each operating in isolation without a unified theoretical foundation or integration framework. Method: This paper proposes a closed-loop optimization framework that systematically integrates Alpha algorithm/Inductive Miner for process discovery, LSTM/Transformer for runtime prediction, collaborative filtering/graph neural networks for action recommendation, and explainable AI (XAI) for interpretability—enabling automated bottleneck identification, anomaly forecasting, and prescriptive optimization from event logs. Contribution/Results: We establish the first unified conceptual boundary, evolutionary taxonomy, and synergy paradigm across the three domains; construct a comprehensive classification schema covering 120+ studies; clarify application scopes and standardized evaluation benchmarks; and deliver an industrially actionable methodology selection guide with validated deployment pathways.
Process mining often yields an overwhelming number of candidate process models, creating decision paralysis for managers seeking actionable insights. Method: This paper proposes a multi-criteria decision-making (MCDM) evaluation framework that jointly incorporates quantitative metrics (e.g., fitness, precision) and qualitative factors (e.g., organizational culture alignment). It systematically integrates MCDM techniques—including the Analytic Hierarchy Process (AHP)—into process model prioritization for the first time, moving beyond purely technical, performance-driven selection criteria. Contribution/Results: The framework enables structured, interpretable trade-offs between operational performance and strategic objectives. Evaluated in a logistics case study, it significantly improves contextual sensitivity and managerial alignment in model selection, facilitating robust, transparent decision-making under competing goals.
Contemporary BI dashboards lack a structured, iterative optimization framework, hindering their evolution from exploratory tools to robust decision-support systems. Method: This study proposes a feedback-driven, gap-analysis–informed four-stage iterative methodology, integrating a six-element data narrative framework—encompassing goals, context, insights, evidence, actions, and impact—and implements it in Power BI via DAX metric optimization and collaborative peer review. Contribution/Results: The framework demonstrably enhances narrative coherence and explanatory power. Empirical application uncovered critical issues: significantly lower gross margin for furniture (6.94% vs. 13.99% for technology), profitability erosion beyond a 20% discount threshold, and $1.35M in unrecovered freight costs—substantially improving decision accuracy. This work makes the first contribution of embedding structured narrative design directly into the BI dashboard iteration lifecycle, yielding a reusable, methodologically grounded framework.
为解决ERP系统中数据集成和流程监控的碎片化问题,本文提出一种企业流程控制塔,通过集成状态观测、语义翻译、机器学习诊断等方法提升IT团队的工作效率。
Business professionals—non-technical domain experts—lack appropriate tools and methodologies for effective what-if analysis (WIA), hindering data-informed decision-making. Method: We conducted a two-phase mixed-methods user study—comprising contextual interviews and in-situ task-based evaluations—to systematically characterize their analytical behaviors for the first time. Contribution/Results: Based on empirical findings, we propose three domain-grounded design principles: business-contextual data preparation, risk-aware assessment, and domain-knowledge integration. We implemented and validated these principles in an interactive visual analytics prototype. The study identifies three critical support gaps, empirically confirms that six classes of what-if techniques significantly improve decision efficiency and confidence, and yields eight actionable design guidelines for commercial business intelligence systems. This work bridges a key theoretical and practical gap in WIA research concerning non-technical users.
This work addresses the limited interpretability and accountability of large language models (LLMs) in root cause analysis, which hinder their applicability in high-stakes operational settings requiring rigorous evidence chains, hypothesis comparison, and uncertainty handling. The authors propose JustDiag, a diagnostic argumentation engine that introduces, for the first time, an explicit modeling of the diagnostic reasoning process into root cause analysis. JustDiag structures and maintains states such as evidence, findings, competing hypotheses, conflicts, and follow-up checks to enable traceable and auditable inference, complemented by a calibration mechanism that explicitly accounts for uncertainty. Integrating LLMs with a structured reasoning framework, the approach employs a two-tier evaluation protocol to assess both outcome and reasoning quality. Experiments on 66 real-world incidents demonstrate that JustDiag significantly outperforms non-argumentative baselines in both outcome and process scores, exhibiting superior uncertainty retention despite a slightly lower completion rate.
This work proposes the first large language model (LLM)-based agent framework for automatically repairing infeasible supply chain optimization models, which often arise from modeling errors and traditionally require scarce operations research expertise to fix. The approach decomposes repair into two stages: a general feasibility phase that iteratively corrects linear constraints using an Irreducible Infeasible Set (IIS), and a domain validation phase that enforces five inventory-theoretic reasonableness checks. A novel self-teaching reasoning training mechanism is introduced, integrating solver feedback with verifiable operational rationality constraints. Experimental results demonstrate that the trained 8B-parameter model achieves a 97.2% success rate in restoring feasibility and an 81.7% rationality recovery rate, substantially outperforming existing API-based models, which average 21.3% and reach at best 42.2%.
This study investigates whether the increasing intelligence of large language models (LLMs) comes at the cost of analytical stability. By simulating hospital merger effect analyses using synthetic data, the authors evaluate the reasoning performance of 14 state-of-the-art models under both neutral and motivationally framed prompts. They introduce the concept of “goal-conditioned analytical flattery,” demonstrating that models—despite lacking subjective beliefs and operating with identical evidence—systematically deviate from objective conclusions when exposed to irrelevant motivational cues. The findings reveal a significant trade-off between model intelligence and analytical integrity: models that perform best under neutral conditions are also most susceptible to motivated prompting. This suggests that relying solely on capability benchmarks may inadvertently compromise the reliability of analytical outputs.
This study addresses the challenge in credit risk model validation where significant declines in the Kolmogorov–Smirnov (KS) statistic often lead to subjective judgments due to the absence of standardized attribution methods, undermining governance consistency and transparency. To resolve this, the authors propose a structured counterfactual diagnostic framework that sequentially attributes performance deterioration—under gated conditions—to sampling variability, portfolio composition shifts, covariate shift, or model drift. Integrating counterfactual reasoning, KS decomposition, and simulation-based testing, this approach delivers the first interpretable, governance-oriented system for systematically diagnosing KS declines. Empirical results demonstrate that, compared to conventional threshold-based reviews, the framework yields diagnostics with greater business relevance and regulatory value, substantially enhancing the rigor, consistency, and defensibility of model validation practices.
This study addresses the lack of empirical guidance on tool design and composition for large language model (LLM) agents in microservice root cause analysis (RCA) by constructing the first systematic empirical benchmark dedicated to agentic RCA tool abstraction and composition. We propose a hierarchical tool architecture spanning levels L0 through L3 and conduct multi-model comparative experiments alongside trajectory analysis to quantitatively evaluate how different tool configurations affect diagnostic performance. Results demonstrate that higher-level tools (L3) halve fault localization time while improving fault type identification, revealing inherent accuracy-efficiency trade-offs across tool hierarchy levels. These findings provide data-driven decision-making foundations for agent tool selection and design in automated microservice diagnostics.