Score
Designs and carries out analyses that evaluate, interpret, and compare policies by modeling incentives, costs, trade‑offs, and likely impacts, and by developing or revising measurement and reporting frameworks. Produces concrete policy interpretations and actionable recommendations grounded in quantitative and qualitative evidence.
Current evaluations of AI governance proposals often fall into binary oppositions, overlooking implicit value trade-offs and lacking transparent analytical tools. This work proposes a multidimensional policy analysis framework that integrates expert interviews with computational text analysis to construct an interpretable scoring system across policy attributes, enabling cross-proposal comparison through visualization. Its novelty lies in three aspects: first, a multidimensional evaluation approach that avoids predetermined conclusions and explicitly reveals inherent trade-offs; second, a transparent hybrid methodology combining qualitative expert insights with quantitative computational validation; and third, the introduction of a domain-calibrated model as a benchmark against general-purpose large language models. The framework enables comparable, interpretable assessments of AI governance proposals across multiple attributes, allowing stakeholders to evaluate proposal relevance and coherence according to their own normative priorities.
Current what-if analysis lacks a unified conceptual framework, leading to terminological inconsistency across domains, structural ambiguity, and divergent interpretations. To address this, we conduct a systematic review of 141 papers in visual analytics and human-computer interaction, proposing Praxa—the first integrative framework that unifies scenario modeling, sensitivity analysis, and counterfactual analysis under a coherent paradigm. Praxa formally defines the underlying motivations, core components (hypothesis generation, intervention modeling, outcome evaluation), and a taxonomy of analytical types. It establishes a standardized terminology and structured model, exposing critical challenges including interpretability, causal modeling fidelity, and alignment with user intent. By clarifying conceptual boundaries and operational relationships among methods, Praxa significantly enhances cross-domain conceptual consistency and application clarity. The framework provides a rigorous foundation for theoretical advancement and the design of next-generation interactive analytical tools.
Data uncertainties—such as measurement errors, missing values, and erroneous links—undermine the credibility of policy decisions. Method: This paper proposes a decision-stability-oriented sensitivity analysis framework that shifts the analytical focus from parameter deviation to decision robustness. It introduces an interpretable, decision-level sensitivity metric and integrates counterfactual modeling, hypothesis-driven perturbation sampling, decision boundary tracking, and interactive visualization. Contribution/Results: Evaluated on two real-world policy domains—U.S. presidential vote prediction and childhood lead exposure assessment—the framework significantly enhances policymakers’ awareness of analytical robustness, explicitly delineates credible decision intervals, and provides an actionable confidence assessment tool for data-informed policymaking under data imperfections.
Business professionals—non-technical domain experts—lack appropriate tools and methodologies for effective what-if analysis (WIA), hindering data-informed decision-making. Method: We conducted a two-phase mixed-methods user study—comprising contextual interviews and in-situ task-based evaluations—to systematically characterize their analytical behaviors for the first time. Contribution/Results: Based on empirical findings, we propose three domain-grounded design principles: business-contextual data preparation, risk-aware assessment, and domain-knowledge integration. We implemented and validated these principles in an interactive visual analytics prototype. The study identifies three critical support gaps, empirically confirms that six classes of what-if techniques significantly improve decision efficiency and confidence, and yields eight actionable design guidelines for commercial business intelligence systems. This work bridges a key theoretical and practical gap in WIA research concerning non-technical users.
This study addresses the absence of a systematic framework in empirical economics for translating analytical findings into normative policy recommendations. Integrating statistical decision theory with the literature on policy choice, the authors develop a unified analytical framework and introduce two types of navigational maps to guide research design. They also implement an R package that automatically generates standardized, publication-ready visualizations of policy impacts. Demonstrated through applications in development economics, this approach substantially enhances the transparency, cross-study comparability, and empirical grounding of policy advice, thereby offering the first end-to-end methodological pipeline that bridges theoretical analysis and practical policy evaluation.
This study addresses the challenge of optimizing data collection to enhance social welfare in policy learning when unobserved heterogeneity is present. Accounting for latent individual differences in policy responses, the authors propose a repeated-measurement design based on proxy variables for latent traits and derive minimax regret bounds for policy rules that either incorporate or omit these latent variables. The theoretical analysis uncovers a novel trade-off between policy class complexity and estimation accuracy, leading to an optimal data collection strategy that allocates resources efficiently between measurement precision and sample size. In a development economics application, incorporating a proxy for entrepreneurs’ managerial ability increases social welfare by 5% and reduces the probability of welfare loss by 50%.
This study addresses a critical limitation in traditional reproducible research, where sharing only code and results fails to expose the implicit assumptions, expectations, and premises underlying an analyst’s reasoning—thereby hindering thorough evaluation of analytical quality. To overcome this, the paper proposes a formal modeling framework that explicitly translates the analyst’s tacit reasoning process into structured logical representations, statically capturing the construction logic of the analysis. This approach enables systematic scrutiny of the analytical chain of reasoning, assumption sensitivity, and conclusion robustness—even in the absence of the original data. Empirical validation on representative data analysis tasks demonstrates the framework’s effectiveness, achieving both logical visualization and data-free static assessment of analytical integrity.
This study examines the practical validity of the official evaluation criteria used in Brazil’s CNPq Research Productivity (PQ) fellowship system, revealing significant discrepancies between stated standards and actual peer-review outcomes. By operationalizing policy dimensions into measurable variables and integrating CV data with OpenAlex metrics to construct predictive features, the research treats evaluation criteria as testable hypotheses for the first time. Employing an enhanced block-wise Boruta algorithm alongside interpretable machine learning models, the analysis achieves an average AUC of 0.96, indicating strong predictive power of PQ levels. However, only a few features—such as publication output, graduate student supervision, and administrative roles—demonstrate statistical significance, while several officially emphasized indicators contribute negligibly. These findings highlight a misalignment between formal evaluation criteria and real-world decision-making, underscoring the need for greater transparency in research assessment systems.