Score
Translating model coefficients, evidence, and uncertainty into clear, actionable recommendations for stakeholders, accounting for trust, comprehension, and decision-making needs. It covers formats and processes that make findings and uncertainty legible and usable for public-health guidance and senior decision-makers.
A systematic methodology for assessing the potential harms of large language models (LLMs) in high-stakes public health contexts—such as infectious disease control, opioid misuse mitigation, and intimate partner violence response—remains critically absent. Method: Drawing on qualitative insights from focus groups with public health practitioners and lived-experience stakeholders, this study employs thematic analysis and participatory knowledge co-creation to develop, for the first time, a four-dimensional LLM risk taxonomy tailored to public health: individual behavioral impact, human-centered care integrity, information ecosystem stability, and technical accountability—all grounded in real-world scenarios. Contribution/Results: The study introduces a novel reflective questioning toolkit to foster interdisciplinary risk governance. It delivers a consensus-based terminology framework, an actionable assessment pathway, and decision-support instruments—enabling responsible, context-sensitive, and non-replacement-oriented deployment of LLMs in public health practice.
Health simulation models—such as agent-based models (ABMs)—are often too complex for key stakeholders (e.g., clinicians, policymakers, patients) to understand and adopt. Existing large language model (LLM)-generated explanations lack stakeholder specificity and fail to accommodate diverse information needs and linguistic preferences. Method: We propose the first stakeholder-centered explanation framework for health simulation, integrating qualitative interviews and quantitative surveys to systematically elicit heterogeneous user requirements, and designing a controllable text generation mechanism that guides LLMs to produce summaries jointly customized in both content and stylistic attributes. Contribution/Results: Rigorous multi-dimensional evaluation demonstrates significant improvements in explanation relevance, readability, and stakeholder acceptance. The framework establishes a reusable methodological foundation for human-AI collaborative explanation in health decision support systems.
This study systematically compares communication styles between large language models (LLMs) and human experts in health misinformation refutation, and examines reader preferences. Addressing the gap in understanding how LLM-generated explanations differ from human-authored ones in health contexts, we adopt a tri-dimensional framework grounded in health communication theory: linguistic features of messages, sender-level persuasive strategies, and alignment with recipient values. Methodologically, we integrate computational text analysis, expert annotation, a 99-participant double-blind user study, and a novel authoritative fact-checking dataset. Our multi-layered evaluation—first of its kind—reveals that LLM explanations significantly underperform humans in persuasiveness, certainty expression, and moral value alignment; yet over 60% of participants rated them as clearer, more comprehensive, and more persuasive, underscoring the critical role of structural coherence in reader engagement. These findings provide both a theoretical framework and empirical foundation for the trustworthy adaptation of LLMs in health communication.
This paper addresses the lack of a unified framework for assessing systematic error (bias) across causal and descriptive inference. We propose the first cross-paradigm, generalizable bias risk assessment method, integrating modeling assumptions, data-generating mechanisms, and inferential objectives to cover high-risk settings—including randomized controlled trials (RCTs), nonprobability sampling, and statistical extrapolation—beyond traditional medical RCT constraints. Our approach combines qualitative bias mapping, assumption sensitivity analysis, and standardized reporting criteria, mandating explicit documentation of untestable assumptions and model uncertainty. The framework has been adopted as a mandatory reporting requirement by leading journals and funding agencies, thereby enhancing the reliability, interpretability, and external validity of research findings. (132 words)
Clinical prediction models are frequently developed from observational data that include early treatments, rendering them vulnerable to confounding, selection bias, mediation effects, and dynamic treatment regimes—collectively termed “causal blind spots”—which lead to miscalibrated risk estimates and suboptimal clinical decisions. This paper formally defines “causal blind spots” for the first time and demonstrates that conventional modeling strategies—treating treatment as a covariate, stratifying by treatment, or omitting treatment—are all unreliable. We propose an intervention-oriented framework centered on the *interventional prediction estimand*, integrating causal diagrams, do-calculus, and potential outcomes theory. This framework mandates embedding causal inference into both model development and validation. By shifting predictive modeling from associative pattern recognition to causal intervention modeling, our approach provides a principled foundation for revising clinical prediction guidelines to ensure causal validity, thereby enhancing the scientific rigor and safety of treatment decisions.
Traditional meta-analyses struggle to quantify the strength of evidence for the presence or absence of an effect and cannot adequately assess the sensitivity of conclusions to publication bias or small-study effects. This work proposes a Bayesian evidence auditing framework tailored to meta-analytic corpora, integrating bias-aware models with unbiased baseline specifications through Bayesian model averaging and Bayes factors. It introduces “rigor” as a composite metric that jointly evaluates the strength of evidence for an effect and robustness to bias—allowing null effects to achieve high rigor scores when supported by strong evidence. Built upon Bayesian random-effects models, the approach employs simulation and resampling strategies within the ADEMP framework, including synthetic data generation, registered-report resampling, and contour-enhanced funnel weighting. Applied to nutritional intervention studies, the method frequently attenuates conventional effect estimates, revealing that many nominally significant findings lack robust evidential support. Full reproducible resources are publicly released.
This study examines how the U.S. Food and Drug Administration’s new food traceability rule transforms small-scale agricultural producers into uncompensated data laborers, exacerbating existing burdens related to labor, financial resources, and technological capacity. Drawing on data feminist theory—an approach newly applied to food regulatory policy analysis—the research employs qualitative coding of 1,198 public comments to systematically uncover structural inequities embedded in the rule’s implementation. The analysis identifies three core tensions: the invisible burden of data labor imposed on marginalized actors, the technical infeasibility of mandated tracking systems for small operations, and regulatory ambiguity that fuels inconsistent enforcement. These findings offer both empirical evidence and theoretical innovation to inform the development of more inclusive and equitable data governance frameworks in food safety regulation.
This study addresses the limitations of individualized clinical decision-making, which is often constrained by the high internal validity but limited external applicability of randomized controlled trials (RCTs) and the strong representativeness yet susceptibility to confounding bias in real-world data (RWD). To overcome these challenges, the authors propose a multi-source data integration paradigm grounded in an explicit causal inference framework. This approach systematically combines RCT and RWD by rigorously defining estimands, ensuring comparability across data sources, and conducting sensitivity analyses. The resulting methodology enhances the reliability and evidentiary strength of treatment effect estimates while providing a practical, regulatory-compliant pathway for generating individualized treatment recommendations.
Despite growing adoption of non-targeted analysis (NTA) in food safety, systematic evaluation of liquid/gas chromatography–high-resolution mass spectrometry (LC/GC-HRMS) NTA tools against FAIR principles and the BP4NTA operational pillars—laboratory validation, data/code availability, standardized formats, knowledge integration, and portable implementation—remains lacking. Method: We conducted a longitudinal, systematic audit of 103 NTA tools published between 2004 and 2025, assessing compliance across all six BP4NTA pillars. Contribution/Results: We found a marked increase in openness (56% → 86%) but a paradoxical decline in practical reproducibility (55% → 43%), quantifying for the first time the persistent “discoverable but not runnable” gap. Critical synergistic deficits between Pillar C1 (validation) and C6 (portable implementation) emerged as the primary bottleneck for regulatory-grade reproducibility. This study fills a key gap in food safety NTA tool assessment and proposes a multidimensional audit framework; only 17% of tools satisfy both validation and portability criteria—providing empirical grounding and actionable pathways toward fully reproducible NTA workflows.
This study addresses the challenge of effectively utilizing large-scale, heterogeneous patient-generated health data in clinical practice, where time constraints and limited data literacy among healthcare professionals hinder meaningful engagement. To bridge this gap, the authors propose an interactive system that integrates large language model (LLM)-generated summaries with a natural language conversational interface, enabling clinicians to rapidly comprehend and flexibly explore multimodal health data within cardiovascular disease risk reduction scenarios. Embedded into clinical workflows, the system combines automated summarization, dialog-driven interaction, and multimodal visualization. A mixed-methods evaluation involving 16 healthcare professionals demonstrated that AI-generated summaries significantly enhanced data interpretability, while conversational capabilities facilitated adaptive exploration, effectively mitigating disparities in data literacy. The study also identifies critical challenges concerning transparency, privacy preservation, and risks of overreliance on AI assistance.