Score
Designs, implements, and evaluates the processes, models, controls, and systems that determine when and how an organization records, reports, and forecasts revenue and that attribute and quantify revenue contributions across contracts, channels, products, or campaigns. This includes defining accounting-compliant revenue recognition policies and automation, building revenue attribution and forecasting models, running revenue impact and performance analyses, and operating revenue planning, management, and quantification workflows (收益评估, 收益量化).
Traditional revenue forecasting approaches struggle to uncover the underlying customer behavioral drivers—such as customer acquisition, repeat purchase rates, and average transaction value—that influence revenue dynamics. To address this limitation, this work proposes the Customer-Based Multi-Task Transformer (CBMT), which uniquely integrates multi-task learning with a Transformer architecture to jointly model customer behavioral metrics and total revenue through shared representations. Furthermore, CBMT incorporates a downstream alignment mechanism to enhance both interpretability and predictive accuracy. Empirical evaluation on real-world customer transaction panel data demonstrates that CBMT outperforms existing methods across 23 out of 24 evaluation metrics, achieving a 30% reduction in total sales prediction error compared to the strongest baseline and significantly surpassing single-task models employed by 74.3% of firms.
Large-scale manufacturers face challenges in after-sales demand forecasting, including difficulty in fusing heterogeneous multi-source signals, weak modeling of COVID-19 disruptions, imbalanced prediction accuracy for long-tail versus high-revenue items, and insufficient business interpretability. Method: We propose an end-to-end interpretable ensemble forecasting framework integrating statistical models, deep learning, and large language models (LLMs). Key innovations include Pareto-aware segmented forecasting, horizon-aware weighted ensemble integration, LLM-driven automated attribution narrative generation, plus integrated change-point detection and WMAPE-optimized calibration. Contribution/Results: The system delivers city-item-level calibrated probabilistic forecasts across 90+ countries and 6,000 SKUs, simultaneously improving prediction accuracy, stability, and operational alignment. Through a performance scorecard and trend attribution module, it shifts evaluation from static accuracy metrics to a dynamic, intervention-oriented decision loop.
In data-driven economies, organizations lack systematic frameworks for evaluating and managing data value within internal business processes. To address this gap, this study develops a comprehensive data value assessment framework grounded in the Balanced Scorecard’s internal process perspective, integrating three interrelated dimensions: data quality, governance compliance, and operational efficiency. It introduces a novel, multi-layered taxonomy of data value—spanning technological, organizational, and regulatory dependencies—that resolves metric redundancy and establishes cross-dimensional conceptual linkages. Through systematic literature review, theoretical modeling, indicator clustering, and taxonomy design, the research produces a scalable, reusable data value metrics system. This system underpins standardized data valuation models and decision-support systems, offering both a methodological foundation and actionable implementation pathways for cross-sectoral data assetization. (149 words)
Existing business process modeling practices separate process flows from business rules, leading to fragmented mental models and impaired comprehension among expert process workers. Method: This study employs a mixed-methods design integrating eye-tracking and concurrent verbal protocol analysis, coupled with cognitive-behavioral coding, to investigate how domain experts perform sensemaking when interpreting integrated process-rule models. Contribution/Results: We identify fine-grained visual search patterns and cognitive bottlenecks that critically affect comprehension efficiency during information foraging and cognitive processing stages. Based on these findings, we propose empirically grounded design principles for personalized cognitive support targeting knowledge workers. The results provide actionable evidence to enhance integrated modeling languages, tool interfaces, and training strategies—advancing business process modeling from syntactic formalism toward cognitive alignment and human-centered design.
This paper addresses the challenge of accurately identifying demand under substantial temporal fluctuations and absent cost-variation information, focusing on the French railway industry. We systematically evaluate the economic performance of revenue management (RM) strategies using a novel identification framework that integrates time-series relative price changes, consumer rational expectations, and firms’ weak optimality conditions in pricing. Our methodology combines structural econometric modeling, counterfactual demand estimation, endogenous price treatment, and censoring-handling techniques to overcome identification issues arising from sales cutoffs and the lack of exogenous price variation. Results show that current RM practices significantly outperform uniform pricing but still incur a 16.7% revenue loss relative to theoretically optimal dynamic pricing. This study provides the first empirical quantification of RM’s net economic value in a real-world industrial setting and reveals its critical role in aggregating and processing information under demand uncertainty.
This study addresses the opacity and accountability challenges in AI-powered hiring systems, which stem from their complex supply chains that obscure the origins of algorithmic bias. Through regulatory analysis, system dependency modeling, and a multi-stakeholder perspective—complemented by case studies and an examination of implementation ambiguities—the work demonstrates for the first time that bias arises primarily from interactions among system components rather than from isolated modules. It further identifies a structural contradiction: deploying organizations bear legal responsibility yet lack technical visibility into upstream components. The research pinpoints two core barriers to effective bias assessment and accountability and proposes a holistic, supply-chain-wide governance framework featuring system-level audits, vendor guidelines, continuous monitoring, and cross-component documentation.
This work addresses the lack of traceable and tamper-resistant transparency mechanisms in large language models (LLMs) deployed in high-stakes decision-making contexts, which undermines accountability. To bridge this gap, the paper introduces the first LLM lifecycle auditing framework that integrates technical provenance with governance records. It proposes a reference architecture enabling cross-organizational traceability and implements a lightweight, open-source Python-based auditing layer. By leveraging append-only logs, event emitters, structured metadata, and an auditor interface, the system seamlessly integrates into existing LLM workflows with minimal intrusiveness. This design ensures complete, tamper-evident traceability across critical stages—including training, deployment, and monitoring—thereby facilitating robust accountability and responsibility attribution throughout the model’s lifecycle.
In non-contractual settings, customer churn is unobservable, rendering accurate counts of active customers challenging. This study identifies a category error in the conventional P(alive) metric, which conflates finite-horizon, verifiable repurchase probabilities with infinite-horizon extrapolations of customer survival. To address this, we propose counting customers based on auditable, finite-horizon repurchase probabilities and develop an interval estimation framework using the beta-geometric family of models, replacing prevailing point estimation approaches. Empirical analysis leveraging a seven-year panel dataset of 31,683 customers reveals that alternative model specifications can yield customer counts differing by up to 7.6-fold, while default parameter choices introduce biases as high as 42%. The proposed method substantially enhances both predictive accuracy and verifiability.
This study addresses the challenge of evaluating reasoning capabilities of large language models (LLMs) in the accounting domain by proposing the first systematic benchmarking framework tailored to this vertical field. The framework incorporates characteristics of model training data and introduces quantifiable, domain-specific tasks and evaluation criteria. Through carefully designed prompt engineering, the authors conduct a comprehensive assessment of GLM-6B, GLM-130B, GLM-4, and GPT-4. Results indicate that prompt formulation significantly influences model performance, with GPT-4 demonstrating overall superiority. Nevertheless, all evaluated models fall short of meeting the reliability standards required for enterprise-level accounting applications. This work establishes a methodological foundation and empirical benchmark for assessing LLM capabilities in specialized professional domains.