Score
Techniques for combining socio-demographic, digital-access, and material-resource data across sources to diagnose coverage bias, attribute changes (e.g., seasonal or shock-driven) to socioeconomic or climatic causes, and adjust inference accordingly.
This study investigates how mode effects (e.g., face-to-face vs. online surveys) and mode selection bias jointly distort epidemiological inference. Using directed acyclic graphs (DAGs), we systematically model their interplay and demonstrate that conventional conditional adjustment may induce collider bias. We integrate DAG-based identifiability analysis, quantitative bias modeling, multiple imputation, and causal sensitivity analysis to characterize the direction and magnitude of resulting biases. Our contributions are threefold: (1) first formal distinction and joint modeling of mode effects and selection bias within mixed-mode survey designs; (2) identification of inherent limitations in standard statistical approaches—particularly naive covariate adjustment—for addressing such biases; and (3) proposal of a DAG-driven bias mitigation framework that enhances causal validity and estimation reliability in multi-mode data integration. (149 words)
A systematic review and practical framework for integrating Earth observation (EO) data with machine learning (ML) to enable causal inference in poverty geography remains absent. Method: This paper introduces the first taxonomy of five EO-based paradigms for causal inference—including outcome imputation, image-based deconfounding, and heterogeneous treatment effect modeling—and establishes a standardized EO-ML causal analysis workflow aligned with the Sustainable Development Goals. The workflow integrates spatial statistics, computer vision, causal discovery, and counterfactual reasoning, emphasizing structured remote sensing representation and causally interpretable modeling. Contribution/Results: We deliver an actionable guideline covering data selection, model adaptation, and evaluation metrics, enhancing credibility and reproducibility of causal analyses across multidimensional development indicators—particularly health outcomes and housing conditions.
Traditional demographic data (e.g., censuses) suffer from high cost and low timeliness, while mobile application data—despite offering high spatiotemporal resolution—exhibit systematic coverage bias due to digital inequality and lack standardized evaluation frameworks. Method: We develop the first reproducible, attribute-agnostic framework for quantifying coverage bias, integrating aggregated mobile data with census benchmarks to derive transparent coverage metrics; we further employ interpretable machine learning to uncover nonlinear geographic, socioeconomic, and demographic drivers of bias. Contribution/Results: Validated on four UK datasets, our approach challenges the assumption that multi-source data inherently improve representativeness. Although mobile data achieve higher overall coverage than traditional surveys, they exhibit pronounced and complex spatial biases. The framework establishes a methodological foundation and practical toolkit for rigorously assessing the reliability of digital trace data in demographic inference.
This study addresses disciplinary fragmentation and concealed biases in socio-technical research by proposing the first embedded bias-aware interdisciplinary framework. It integrates sociological qualitative methods with computer science–based quantitative techniques to systematically examine online harms experienced by ethnic minorities accessing digital social housing services in the UK. Methodologically, it combines grounded theory coding, LDA topic modeling, sentiment analysis, qualitative comparative analysis (QCA), and a mixed-methods design to ensure bias source visualization, methodological transparency, and result interpretability. The study identifies four structural vulnerability dimensions—discrimination, digital poverty, low digital literacy, and limited English proficiency—and provides the first empirical evidence that Black African communities disproportionately experience compounded vulnerabilities. The framework enhances the robustness, ethical rigor, and cross-disciplinary reproducibility of socio-technical research.
Existing research on assessing the socioeconomic impacts of climate disasters using textual data often suffers from ambiguous impact definitions, inadequate handling of spatiotemporal biases, and inconsistent modeling strategies, which undermine result transparency and comparability. This study addresses these limitations by systematically integrating large-scale textual sources—including news articles, social media posts, and official reports—and proposes the first standardized methodological framework tailored to this task. The framework explicitly defines impact criteria, corrects for spatiotemporal biases, and standardizes model selection. Leveraging natural language processing and large language models within a “text-as-data” paradigm, the work introduces a reproducible, transparent, and comparable set of best practices for extracting and quantifying disaster-related information, thereby significantly enhancing the accuracy of climate disaster impact assessment and attribution studies.
This study addresses the challenges of integrating multi-source, heterogeneous indicators and constructing transparent, comparable composite indices for environmental risk assessment under extreme climate events. It systematically compares four methodological approaches—weighted linear aggregation, principal component analysis, fuzzy logic, and data envelopment analysis—evaluating their applicability, underlying assumptions, and limitations. To enhance interpretability and policy relevance, the framework integrates expert knowledge with scenario-based simulation. Empirical validation is conducted using a streamlined dataset and real-world drought cases. Results reveal significant divergence across methods in dimensionality reduction capability, weight sensitivity, and robustness; notably, fuzzy logic and weighted aggregation demonstrate superior suitability for policy-oriented resilience assessment. The study contributes a methodological framework and a reusable, step-by-step guideline for constructing composite indices, supporting evidence-based decision-making in sustainable development and climate adaptation.
To address latent bias arising from unmeasured confounding in observational studies, this paper proposes a novel causal inference paradigm based on sample splitting: data are partitioned into planning and analysis samples, where the former adaptively selects robust design parameters (e.g., matching strategies, covariate sets), and the latter yields unbiased causal estimates. The method innovatively integrates multiple testing correction, heteroskedasticity-robust covariance estimation, and formal sensitivity analysis—extending support to multiple outcome variables, thereby relaxing the conventional single-outcome assumption. We establish theoretical guarantees of statistical validity under latent bias. Simulation studies demonstrate substantially higher statistical power than benchmark methods under strong unmeasured confounding. Empirical application to assessing the multidimensional impacts of floods in Bangladesh confirms practical feasibility and robustness.
This study addresses the pervasive issue of non-representative GPS mobility data in low- and middle-income countries, where coverage biases across data sources remain poorly understood. Integrating Facebook and Veraset GPS datasets with census data from 2,478 Mexican municipalities, the authors employ interpretable machine learning and spatial statistical models to systematically dissect the origins, spatial structure, and drivers of coverage bias. Findings reveal that coverage bias exhibits strong data-source specificity and spatial dependence: Facebook data demonstrate more uniform coverage, whereas multi-app aggregated data disproportionately represent wealthier, more digitally connected areas. Explicitly modeling spatial autocorrelation substantially improves explanatory power for these biases. The results underscore the necessity of tailoring bias-correction strategies to specific data sources and highlight that a portion of spatial variation in coverage cannot be fully accounted for by observable covariates.
This study addresses how AI-driven climate information systems exacerbate the global digital divide between the Global North and South, resulting in systemic biases in data representation, model validation, and knowledge articulation for vulnerable regions. By integrating high-resolution climate modeling, large language model analysis, and data equity assessments, the research uncovers structural inequities embedded throughout the AI system lifecycle. It proposes a paradigm shift from a “model-centric” to a “data-centric” approach, advocating for the co-development of climate digital public infrastructure and knowledge co-production mechanisms. The work advances a vision of democratized computational sovereignty and offers both technical pathways and policy frameworks to foster an inclusive, resilience-oriented global climate information ecosystem.
This study addresses the issue of model unfairness toward specific subpopulations in high-stakes clinical settings, stemming from biases in training data. Focusing on intensive care unit (ICU) environments, the authors systematically evaluate how incorporating external electronic health records—such as those from eICU and MIMIC-IV—affects fairness across subgroups. Through multi-source data integration, subgroup performance analysis, post-hoc calibration, and comparative data selection strategies, they demonstrate that simply increasing training data volume does not necessarily improve fairness and may even degrade subgroup performance, thereby challenging the prevailing “more data is better” paradigm. The work proposes a novel framework of joint data-model intervention, showing that combining targeted data augmentation with post-processing calibration effectively enhances both fairness and overall predictive performance.
This study quantifies the impact of extreme climatic factors on cause-specific mortality across diverse U.S. populations to address uncertainties in physical climate risk. Innovatively integrating compositional data analysis (CODA) into climate–mortality research, the authors employ principal component analysis (PCA) to reduce dimensionality of climate variables and construct generalized additive models (GAMs) to capture nonlinear relationships between factors such as temperature and sea-level rise and the proportional distribution of causes of death. The findings reveal that elevated temperatures and rising sea levels significantly increase the proportion of deaths attributable to hypertensive heart disease, with individuals aged 55–95 exhibiting heightened sensitivity. Moreover, the study identifies climate-driven natural hedging effects among different causes of death, offering novel insights for climate risk modeling and the design of insurance products.
This work addresses the persistent underperformance of machine learning models on intersecting sensitive subgroups—such as those defined by race and gender—attributed to inadequate bias metrics and insufficient representation in training data. The authors introduce coverage constraints into a bias mitigation framework for the first time, formulating the problem via integer linear programming to optimize data modification costs while ensuring adequate representation across all (intersecting) groups and bounding approximation error in bias reduction. This approach enables quantification of the “price of fairness,” facilitating principled trade-offs between equity and data efficiency in legal compliance and data governance contexts. Empirical evaluations across multiple benchmark datasets and classifiers demonstrate that the proposed framework effectively preserves predictive accuracy while substantiating the critical role of coverage constraints in safeguarding both fairness and performance in downstream models.