Score
Structured evaluation methods for measuring and comparing the regulatory, ethical, environmental, or equity consequences of models, designs, or policy choices to inform reporting and recommendations. This includes comparative analysis of alternative liability or regulatory regimes, pricing and value-allocation trade-offs, and guidance development based on observed empirical differences.
This study addresses ongoing academic and policy debates regarding the climate mitigation contributions of carbon pricing instruments (carbon taxes, emissions trading systems—ETS) and carbon offset mechanisms (voluntary carbon markets, REDD+), evaluating their empirical performance in effectiveness, equity, and environmental integrity. Employing an umbrella review methodology—systematically screening literature per PRISMA guidelines, assessing systematic review quality with AMSTAR-2, and quantifying risk of bias in primary studies using ROBINS-I—we provide the first methodologically rigorous cross-tool comparison. Results indicate that carbon taxes and ETS achieve only moderate emission reductions, constrained by narrow sectoral coverage and high price volatility. Voluntary carbon markets suffer from pervasive methodological flaws and monitoring failures, yielding high credibility risks and undermining their viability for credible net-zero pathways. This work establishes a methodological benchmark for evidence-based climate policy evaluation and delivers critical empirical evidence to inform tool selection and design.
This study addresses the challenge of jointly weighing loss aversion, adverse effects, and implementation costs in clinical or public health intervention decisions. Methodologically, it proposes a formal Bayesian utility-theoretic decision framework that integrates posterior distribution inference with a structured, non-arbitrary 1–9 value scale to map effect size, inter-individual response variability (standard deviation), and multidimensional costs onto a unified expected utility score. The framework further incorporates credible intervals, sensitivity analyses, and estimation of response-type proportions (benefit, no effect, harm). Its primary contributions are: (1) establishing the first comparable, interpretable, and operationally actionable quantitative utility paradigm for intervention evaluation; (2) generating both a single composite utility score and a full population-level response distribution; and (3) implementing the framework in an accessible spreadsheet tool, thereby providing a transparent, robust, and user-friendly aid for evidence-informed decision-making.
Organizations struggle to quantify the commercial value of data assets due to fragmented, siloed valuation approaches—divided across economic, governance, and strategic perspectives—and the absence of actionable mechanisms. This paper proposes an integrated data valuation framework that unifies these three perspectives using a Balanced Scorecard–inspired hybrid model. The framework combines qualitative scoring, cost-utility estimation, data quality indexing, and Analytic Network Process (ANP)-based multi-criteria weighting to enhance transparency and strategic alignment. Adopting a design science research methodology, it is iteratively refined through embedded industrial case studies. Empirical evaluation demonstrates that the framework significantly reduces subjectivity in valuation, improves the precision of mapping data assets to organizational strategic objectives, and supports diverse monetization pathways—including Data-as-a-Service (DaaS). It exhibits cross-industry applicability and robustness.
Current evaluations of AI governance proposals often fall into binary oppositions, overlooking implicit value trade-offs and lacking transparent analytical tools. This work proposes a multidimensional policy analysis framework that integrates expert interviews with computational text analysis to construct an interpretable scoring system across policy attributes, enabling cross-proposal comparison through visualization. Its novelty lies in three aspects: first, a multidimensional evaluation approach that avoids predetermined conclusions and explicitly reveals inherent trade-offs; second, a transparent hybrid methodology combining qualitative expert insights with quantitative computational validation; and third, the introduction of a domain-calibrated model as a benchmark against general-purpose large language models. The framework enables comparable, interpretable assessments of AI governance proposals across multiple attributes, allowing stakeholders to evaluate proposal relevance and coherence according to their own normative priorities.
Financial regulators face challenges in backtesting risk measures due to the lack of a unified, robust, and manipulation-resistant framework. This paper introduces the first e-value- and e-process-based standard for risk measure evaluation and comparative backtesting, enabling model-free, nonparametric, dynamic validation of widely used measures—including Value-at-Risk (VaR) and Expected Shortfall (ES). By pioneering the application of e-processes to financial backtesting, it achieves unified treatment of both identifiable and incentive-compatible risk measures, bridging statistical rigor with regulatory practicality. The proposed method is inherently resistant to p-hacking, requires no distributional assumptions, and substantially enhances backtest robustness. Extensive experiments on simulated and real-world financial datasets demonstrate its statistical validity, computational feasibility, and direct applicability to regulatory practice.
This study addresses the challenge of aligning ESG controversy events—extracted from unstructured news—with international normative frameworks (e.g., the UN Global Compact). To bridge this gap, we propose a semi-automated, lightweight ontology construction method that integrates large language models (LLMs) with formal RDF schema design, transforming abstract sustainability principles into reusable, interpretable semantic templates. These templates enable structured event knowledge extraction from news texts and the construction of an ESG controversy knowledge graph explicitly aligned with global standards. Our key contribution is the automated, accurate, and interpretable mapping of normative principles to machine-readable rules—ensuring cross-regional consistency. The resulting framework supports precise identification and semantic provenance tracing of regulatory violations, providing a scalable, transparent knowledge infrastructure for regulatory compliance assessment and sustainable investment decision-making.
Accurately estimating the direct, unmediated effect of environmental amenities on housing prices (DUET) is crucial for welfare analysis, yet existing methodologies face significant limitations. This study leverages over one million property transactions in New York State from 1990 to 2024 to construct an empirical “ground truth” that preserves the true data-generating process. Using Monte Carlo simulations, it systematically evaluates the estimation accuracy of generalized difference-in-differences (DID), two-way fixed effects models, and causal machine learning methods—including causal forest DID—across varying sample sizes. The results demonstrate that generalized DID consistently outperforms benchmark models across all scenarios, while causal machine learning approaches exhibit strong performance when sample sizes exceed 3,000 observations, with causal forest DID achieving accuracy nearly on par with generalized DID, thereby offering a reliable methodological alternative for DUET estimation.
This study addresses the absence of an operational framework for effectively translating long-term environmental scenarios into counterparty credit risk metrics suitable for pricing and regulatory capital calculations. It proposes the first integrated framework—Environmental Credit Valuation Adjustment (Environmental CVA)—that jointly incorporates climate and nature-related factors. The approach maps environmental scenario drivers to default intensity, introduces ecosystem-specific tail generators to quantify scenario model risk, and employs Kullback–Leibler divergence-based distributionally robust optimization to account for directional model misspecification risk. Empirical results demonstrate that different ecosystem generators yield significantly divergent nature-related CVA estimates, revealing a linkage mechanism through which climate and nature risks co-propagate. These findings underscore the necessity and efficacy of an integrated Environmental CVA assessment framework.
Current AI ethics assessments are fragmented, focusing predominantly on fairness, transparency, privacy, and trust at the model or output level while neglecting inter-component system interactions, real-world harm contexts, and causal harm propagation pathways—resulting in evaluations disconnected from actual risk scenarios and lacking actionable thresholds. Method: Through a scoping review synthesizing nearly 800 ethics metrics, this study constructs the first four-dimensional relational framework—“System Components–Attributes–Risks–Harms”—to systematically map ethical assessment dimensions. Contribution/Results: The framework uncovers three critical gaps: insufficient system integration, weak contextual embedding, and poor actionability. It advances AI ethics evaluation from isolated metric measurement toward a systemic, traceable, and intervention-oriented paradigm—enhancing regulatory alignment and practical deployment efficacy in industry settings.
In the absence of head-to-head trials, indirect treatment comparisons play a critical role in health technology assessment (HTA). This work proposes a methodological framework tailored to the French HTA context, developed by a working group of French experts in market access and methodology, informed by guidance from the Transparency Committee of the French National Authority for Health and systematic reviews. The framework emphasizes early planning, rigorous control of confounding factors, and explicit evaluation of similarity and transitivity assumptions. It also establishes clear criteria for the appropriate use of single-arm trials with external comparators and integrates advanced methods such as network meta-analysis, population-adjusted indirect comparisons, and non-randomized study designs. The resulting practical guidance aims to enhance the quality, robustness, and validity of indirect comparisons, thereby supporting reliable decision-making in complex healthcare settings where direct evidence is unavailable.
This work addresses the limitation of existing evaluation benchmarks that oversimplify societal risk as a single scalar mean, thereby neglecting its multidimensionality, distributional structure, and tail extremities. To overcome this, the authors propose the SHARP framework, which models social harm as a multidimensional random variable encompassing bias, fairness, ethical considerations, and epistemic reliability. Risk aggregation is performed via an additive cumulative log-risk formulation, and tail-sensitive statistics—particularly Conditional Value-at-Risk at the 95th percentile (CVaR95)—are introduced for robust assessment. Experiments across eleven state-of-the-art large language models reveal that models with similar average risk exhibit more than twofold differences in tail exposure and volatility, uncovering heterogeneous failure modes invisible to conventional scalar metrics and demonstrating SHARP’s unique capacity to identify high-risk behaviors.