Score
Empirical methods for constructing and analyzing metrics of institutional quality and subgroup heterogeneity, used to assess moderation effects, composition changes, and dispersion across cohorts or events.
This study addresses the inconsistency in causal effect estimates between observational studies and randomized controlled trials (RCTs) by proposing the first unified framework for decomposing causal effect heterogeneity. The framework systematically identifies and quantifies three sources of heterogeneity: differences in covariate distributions, variation in mediating pathways, and shifts in outcome-generating mechanisms. Methodologically, it formally defines effect decomposition across data types (observational vs. experimental), integrating causal inference, sensitivity analysis, and decomposition modeling, while enabling robust parameter estimation under multiple hypotheses. Evaluated through simulation studies and an empirical analysis of the “Moving to Opportunity” experiment, the framework demonstrates improved interpretability, robustness, and policy generalizability in synthesizing evidence from heterogeneous data sources.
Existing causal mediation analysis relies heavily on untestable multiple ignorability assumptions, undermining the robustness of mechanistic inference. This paper proposes a novel identification framework grounded in heterogeneous treatment effects, which integrates explicit and implicit mediation pathways via causal decomposition—enabling simultaneous identification of total, direct, and indirect effects without requiring multiple ignorability. The method combines flexible heterogeneity modeling, Monte Carlo simulation–based validation, and a dedicated software implementation. It is empirically validated in two real-world applications: public resource governance and voter information dissemination. Simulation studies demonstrate that, compared to prevailing approaches, the proposed method achieves substantially improved estimation accuracy and bias control. By relaxing strong untestable assumptions and offering practical implementation tools, this work advances causal mediation analysis with a more robust, interpretable, and accessible framework for uncovering underlying causal mechanisms.
This study addresses the unclear mechanisms through which time-varying covariates (TVCs) influence latent class trajectory heterogeneity in nonlinear growth mixture models (GMMs). We propose a novel TVC decoupling framework that decomposes each TVC into two distinct components: a baseline trait—capturing its effect on initial growth factors—and a time-specific state—directly influencing observed outcomes—while allowing heterogeneous effects across latent classes. Methodologically, we integrate GMM with an extended mixture-of-experts (MoE) architecture and the proposed TVC decomposition, validated via Monte Carlo simulations and empirical longitudinal analysis. Simulation results confirm unbiased parameter estimation and nominal coverage of confidence intervals. Empirically, we identify significant between-class heterogeneity in both baseline and dynamic effects of reading ability on mathematics achievement trajectories. To our knowledge, this is the first work to systematically disentangle baseline versus dynamic TVC effects and model their cross-class heterogeneity within nonlinear GMMs, thereby enhancing theoretical precision and empirical interpretability in attributing trajectory heterogeneity.
This paper addresses a fundamental question in causal mechanism identification: under what conditions can heterogeneous treatment effects (HTEs) be used to infer the activation of a specific causal mechanism? It highlights that prevailing HTE detection methods—relying on pre-treatment covariates—implicitly assume linearity or additivity, rendering them invalid for mechanism inference under nonlinear outcome generation. Method: We formally characterize the necessary and sufficient conditions under which HTEs carry identifying information about mechanism activation. Using potential outcomes frameworks, mechanism identification theory, and HTE modeling, we prove that nonlinear transformations of outcomes generally eliminate the inferential value of HTEs for mechanisms. Contribution: We derive testable experimental design principles and an interpretive framework that explicitly delineate the boundaries of mechanism identification. Our results provide empiricists with robust identification guidelines and a systematic pathway for sensitivity analysis, bridging theoretical causality and applied HTE estimation.
Existing academic impact metrics lack organizational-level adaptability, particularly for research teams such as conference program committees (PCs) or journal editorial boards. Method: This paper proposes the alpha-index—a novel composite metric that jointly models statistical homogeneity (extending group h-index consistency) and collective h-group influence, moving beyond simple aggregation of individual h-indices. It is the first to unify homogeneity measurement with an extended group h-index framework. Results: Empirical evaluation in computer science demonstrates that alpha-index rankings of conference PCs align strongly with authoritative manual classifications (Pearson *r* > 0.85), significantly outperforming baselines including h-median and h-sum. The metric offers an interpretable, reproducible, and organization-aware assessment paradigm, directly supporting high-stakes decision-making in research funding allocation, project review, and editorial board selection.
This study addresses the excessive polarization observed in Italy’s Institutional Scientific Productivity Index (ISPD) rankings, which stems primarily from the unmodeled homogeneity of normalized scores within academic departments. The work formally characterizes, for the first time, the relationship between such intra-departmental score homogeneity and department size. To mitigate this bias, the authors propose an adjusted ISPD index based on maximum likelihood estimation and introduce a novel Betoidal probability distribution tailored for truncated and rounded publicly available data. Empirical evaluations using real Italian data from 2017 and 2022, complemented by simulation studies, demonstrate that the proposed method substantially alleviates ranking polarization and yields fairer departmental assessments compared to the original ISPD and other existing approaches.
This study addresses the absence of standardized metrics for quantifying the distributional alignment between survey samples and target populations across multidimensional demographic characteristics. To this end, it introduces the Global Representativeness Index (GRI), which—by incorporating total variation distance into survey methodology for the first time—establishes a symmetric [0,1] scoring framework to assess the fidelity of samples with respect to complex demographic structures. The GRI leverages benchmark demographic data from the United Nations and Pew Research Center and complements design effect to form a novel paradigm for sample quality evaluation, implemented via an open-source Python library. Validation across multiple international survey datasets reveals that even large-scale probability samples typically achieve fine-grained GRI scores below 0.36, underscoring substantial deficiencies in current surveys’ demographic representativeness.
This study addresses the challenge of estimating heterogeneous treatment effects when post-treatment variables—such as non-compliance—induce endogenous selection bias if naively conditioned upon. The authors propose a lightweight, assumption-lean empirical stratification framework that predicts latent post-treatment responses using baseline covariates to construct an empirical score, which in turn defines observable subgroups for effect estimation. This approach innovatively bridges empirical stratification with principal strata analysis: it recovers principal causal effects under principal ignorability, yet remains informative even when this assumption fails. To flexibly capture effect heterogeneity, the method introduces a projected ETE (Expected Treatment Effect) curve. Supported by theoretical guarantees and implemented via a semiparametric influence function estimator, the framework demonstrates strong empirical validity and robustness in real-data applications.
This study addresses the limitations of existing exhaustive subgroup treatment effect plots, which struggle to reliably assess heterogeneity under small sample sizes and multiple testing, and lack a formal quantification of the significance of observed heterogeneity under the null hypothesis of homogeneous treatment effects. The authors propose a computationally efficient strategy to construct homogeneity regions by leveraging a Doubly Robust learner to generate pseudo-outcomes for subgroup effect estimation. By constructing a reference distribution under homogeneity, the method provides the first framework to quantify evidence of heterogeneity directly within exhaustive subgroup plots. An explicit formula for homogeneity regions is derived, accompanied by several approaches for computing critical thresholds. Empirical evaluations in cardiovascular clinical trials and simulation studies demonstrate well-calibrated performance and substantial improvements over conventional methods based on subgroup mean differences.
Existing methods struggle to effectively quantify residual heterogeneity in causal effects after adjusting for covariates, particularly lacking suitable metrics in settings with continuous treatments or continuous outcomes. This work proposes, for the first time, P/N-CACE for binary treatments with continuous outcomes and P/N-CPICE for continuous treatments with continuous outcomes, both designed to characterize causal effect heterogeneity unexplained by observed covariates. Building upon the conditional average causal effect (CACE) framework and stochastic intervention strategies, we establish formal identification theorems and develop corresponding bounding analysis theory. Empirical applications on real-world data demonstrate the effectiveness and practical utility of the proposed measures in uncovering residual heterogeneity, substantially expanding the scope of answerable questions in causal inference.