Score
Systematic evaluation of policies, governance structures, and their impacts using qualitative and quantitative methods to inform incentive design and equitable interventions. Applied to cost-accounting, governance-design proposals (e.g., human-manager models), and mapping access/control around powerful technologies and exclusion mechanisms.
Current evaluations of AI governance proposals often fall into binary oppositions, overlooking implicit value trade-offs and lacking transparent analytical tools. This work proposes a multidimensional policy analysis framework that integrates expert interviews with computational text analysis to construct an interpretable scoring system across policy attributes, enabling cross-proposal comparison through visualization. Its novelty lies in three aspects: first, a multidimensional evaluation approach that avoids predetermined conclusions and explicitly reveals inherent trade-offs; second, a transparent hybrid methodology combining qualitative expert insights with quantitative computational validation; and third, the introduction of a domain-calibrated model as a benchmark against general-purpose large language models. The framework enables comparable, interpretable assessments of AI governance proposals across multiple attributes, allowing stakeholders to evaluate proposal relevance and coherence according to their own normative priorities.
This study examines how QALY-oriented health insurance policies balance efficiency, equity, and sustainability. Method: We propose a reverse behavioral optimization framework integrating QALY-based health outcomes, ROI-driven incentives, and adaptive learning; introduce the System Impact Index (SII) to quantify macro-level policy effects; and apply behavioral elasticity analysis within the FOSSIL paradigm—using sample-sensitive importance weighting and regret-minimizing estimation—validated via OECD-WHO panel data simulations and sensitivity analyses. Contribution/Results: We demonstrate that the healthcare system operates near an efficiency saturation frontier, where minor behavioral parameter adjustments trigger nonlinear shifts in resilience, equity, and ROI. Empirically, equity optimization exhibits diminishing returns under saturation but substantially enhances systemic stability.
Existing AI governance frameworks predominantly emphasize high-level principles, overlooking practitioners’ operational concerns within real-world organizational contexts. Method: This study pioneers the use of over 100,000 user reviews of AI products from G2.com as empirical data, applying BERTopic-based topic modeling, semantic similarity analysis, and cross-domain thematic alignment to derive governance themes inductively. Contribution/Results: The analysis confirms the practical relevance of established themes—e.g., privacy and transparency—while systematically uncovering six critical practice-oriented dimensions neglected by current frameworks: project management, strategic alignment, customer engagement, organizational adaptation, process integration, and accountability assignment. By shifting from normative abstraction to empirically grounded, context-sensitive insights, this work provides a novel, operationally actionable foundation for AI governance—establishing a new paradigm that bridges theory and practice.
Existing research on human-AI collaborative decision-making often lacks systematic guidance in incentive mechanism design, undermining participant behavior and the reliability of findings. This study addresses this gap through a thematic literature review combined with qualitative analysis and pattern recognition, offering the first systematic synthesis of key dimensions and causal pathways underlying incentive design. Building on these insights, the work proposes a standardized yet flexible “Incentive-Tuning” framework that provides researchers with structured guidelines for designing, reflecting upon, and documenting incentive mechanisms. By doing so, the framework substantially enhances experimental validity and strengthens the reproducibility and generalizability of empirical findings in human-AI collaborative decision-making research.
In quantitative pairwise comparisons, expert judgments are vulnerable to bribery-based manipulation, leading to distorted global rankings. Method: This paper formally defines the “targeted manipulation” problem for the first time and introduces a unified modeling framework integrating game theory and graph theory to characterize adversarial interventions. It proposes three polynomial-time solvable manipulation algorithms capable of precisely achieving desired rankings. Contribution/Results: Theoretical analysis demonstrates that even minimal bribery costs can significantly distort ranking outcomes. Furthermore, the study uncovers structural properties and inherent vulnerabilities of manipulation strategies, providing a theoretical foundation for detecting anomalous judgments and designing robust aggregation mechanisms. This work bridges a critical gap in the robustness literature on pairwise comparisons by establishing the first formal model of adversarial intervention.
Traditional health evaluation metrics, such as QALYs and PALYs, struggle to simultaneously account for equity and productivity contributions. This work proposes a unified framework that integrates welfare economics and health measurement theory through a normative axiomatic approach. Within this framework, a new class of evaluation functions is developed, satisfying scale invariance and the Pigou-Dalton transfer principle. The authors derive a tractable power-form representation of these functions, offering a coherent basis for assessing interventions that jointly affect health and productive capacity. By doing so, the framework overcomes key limitations of existing metrics, providing a principled balance between efficiency and fairness in health policy evaluation.
This study addresses the challenge of evaluating whether policy audit reports generated by large language models are substantiated by valid evidence. The authors develop a controlled evaluation framework that fixes policy provisions, scoring criteria, and the underlying model while systematically varying the evidence interface to produce structured reports. These reports are then assessed by human annotators across multiple dimensions: correctness, relevance to policy provisions, diagnostic value, and evidence misuse. Innovatively framing internal model evidence as an evidence design problem, the work proposes a hybrid evidence package format and reveals a critical risk: when causal grounding is insufficient, models may erroneously reuse superficially plausible but irrelevant internal labels as evidence. Experiments on 600 reports derived from 60 AGORA cases demonstrate that the hybrid evidence interface yields optimal performance, whereas control conditions with scrambled evidence correlations show that models can generate seemingly coherent reports citing internally fabricated yet substantively irrelevant evidence—highlighting significant governance risks.
In multi-stakeholder platforms, software architecture decisions often implicitly entrench conflicting requirements without systematic support for mapping governance principles to technical design. This work proposes the first governance-architecture alignment framework, explicitly linking five core governance principles to the space of architectural decisions, thereby rendering implicit governance stances identifiable and contestable. The framework also exposes how default technical choices can obscure underlying value commitments. Feasibility is preliminarily demonstrated through a constructive case study of a pig-farming knowledge platform in Rwanda. Future work will employ pre- and post-intervention user judgment studies to evaluate the framework’s impact on actual governance outcomes.
This study investigates the mechanisms underlying the emergence of social bias in artificial intelligence systems, practitioners’ understandings of these issues, and potential mitigation strategies. Employing a qualitative multiple-case design grounded in an interpretivist paradigm, the research integrates intersectionality theory and cognitive science, drawing on semi-structured interviews, document analysis, and triangulation to examine AI practitioners’ experiences across design, development, and governance. Findings reveal that algorithmic bias is deeply rooted in historical inequities, exclusionary assumptions, and organizational pressures for efficiency, underscoring the insufficiency of purely technical fixes. The study proposes an innovative approach that embeds ethical considerations early in the development lifecycle, strengthens structural accountability, fosters diverse stakeholder participation and cognitive awareness, and actively reshapes organizational culture to cultivate AI systems that are transparent, accountable, and aligned with community values.