estimate marginal estimands

Design and implement identification strategies and statistical estimators for marginal estimands—population- or policy-level averages or contrasts—including derivations of identification conditions and specialized approaches for settings with truncation by death where outcomes are undefined for some units; build estimators that adjust for confounders and covariates. Quantify estimator uncertainty (e.g., confidence intervals and variance estimation), derive properties such as bias and consistency, and evaluate robustness to alternative identification or truncation/missingness mechanisms.

estimatemarginalestimands

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.01
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$206K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Finite Population Identification and Design-Based Sensitivity Analysis

Apr 19, 2025
BK
Brendan Kline
🏛️ University of Texas at Austin | Duke University

This paper addresses the lack of robust design foundations for sensitivity analysis in finite-population causal inference. Methodologically, it introduces a novel sensitivity analysis framework grounded in the experimental design distribution—first integrating design-based distributions with partial identification theory to construct model-free, non-asymptotic confidence intervals for the average treatment effect (ATE). It further reinterprets the role of randomization in sensitivity analysis and provides a new design-driven rationale for covariate balance checks. Key contributions include: (1) model-free, finite-population inference under heterogeneous treatment effects; (2) robust ATE confidence intervals with clear identification-theoretic interpretation; and (3) empirical validation across three real-world applications, demonstrating reliability and practicality in small-sample and highly heterogeneous settings.

Analyzes randomization role and motivates covariate balance examinationConstructs design-based confidence intervals for heterogeneous treatment effectsDevelops sensitivity analysis using design distributions for finite populations

Inference for Interval-Identified Parameters Selected from an Estimated Set

Mar 01, 2024
SH
Sukjin Han
🏛️ University of Bristol | University of Colorado

This paper addresses the dual challenge of (i) interval identification of parameters—such as the average treatment effect—under observational data and noncompliance, and (ii) subsequent data-dependent policy selection (e.g., optimal complier selection) from the estimated interval. We propose the first statistical inference framework tailored to the joint setting of “interval identification + data-dependent selection.” Methodologically, we construct three novel classes of confidence intervals leveraging extreme-value theory, asymptotic analysis of set-valued mappings, and coverage probability control, ensuring uniform asymptotic validity under weak regularity conditions. Compared with conventional approaches that ignore selection bias, our method substantially improves coverage robustness across diverse policy evaluation settings. The framework provides a theoretically grounded and empirically reliable tool for causal inference and applied policy analysis.

Addressing data-dependent selection from estimated setsDeveloping confidence intervals for selected treatment effectsInference for interval-identified parameters under selection

In longitudinal studies, death often truncates non-fatal outcomes, rendering existing causal estimands either restricted to the survivor subpopulation or lacking a clear causal interpretation for the entire target population. This work proposes a novel class of marginal separable effects that, for the first time, defines a causally interpretable overall effect applicable to the full population, extending conditional separable effects into a population-level causal summary measure. Drawing on causal inference theory, we establish identification assumptions and develop an estimation approach that integrates longitudinal observational modeling with reweighting strategies. Reanalysis of a prostate cancer clinical trial demonstrates that different estimands can yield divergent conclusions about treatment efficacy, thereby highlighting the practical relevance and sensitivity of the proposed framework.

causal inferencelongitudinal studiesmarginal estimands

Identification and Semiparametric Estimation of Conditional Means from Aggregate Data

Sep 24, 2025
CM
Cory McCartan
🏛️ Pennsylvania State University | Yale University

This paper addresses ecological inference: estimating group-level means of individual outcomes using only aggregate-level data (e.g., regional averages). Existing methods rely on strong aggregation assumptions, yielding fragile identification conditions. We formally characterize weaker, more plausible identification conditions and propose a debiased machine learning estimator grounded in a partially linear structure. Our method accommodates multiple covariates, enables semiparametric sensitivity analysis, and supports asymptotically efficient inference for local individual effects. Integrating ecological inference, debiased ML, semiparametric modeling, and high-dimensional statistical inference, it delivers robust estimation of average treatment effects. Simulation studies and empirical applications demonstrate superior performance over leading alternatives. An open-source software implementation is provided.

Addressing ecological inference with weaker identification assumptionsDeveloping semiparametric estimators for conditional meansEstimating group-level outcome means from aggregated data

Weak Identification with Bounds in a Class of Minimum Distance Models

Dec 21, 2020
GF
Gregory Fletcher Cox
🏛️ National University of Singapore

To address inaccurate parameter estimation and poor confidence interval coverage under weak identification, this paper proposes a robust inference method within the minimum distance framework that incorporates parameter boundary information. Methodologically, it unifies asymptotic theory for weak identification with limit distribution theory for parameters lying on boundaries, thereby constructing a boundary-constrained identification-robust estimation and inference system. The approach builds upon minimum distance estimation, integrated with factor model identification analysis and boundary-constrained inference techniques. Simulation studies and an empirical application to parental educational investment demonstrate that the method substantially improves confidence interval coverage—bringing it close to the nominal level—and enhances estimation precision under weak identification. It overcomes the failure of conventional methods near parameter boundaries and establishes a novel paradigm for weak-identification inference in structural models featuring inequality constraints.

Addresses weak parameter identification in minimum distance modelsDemonstrates application in latent factor and GARCH model contextsIncorporates parameter bounds into identification-robust inference methods

Latest Papers

What's happening recently
View more

This study addresses the challenge of partial identification of causal effects in stratified randomized experiments, where attrition and heterogeneous treatment assignment proportions complicate inference. The authors propose a unified analytical framework that integrates Lee bounds, inverse probability weighting, and a global trimming strategy, accommodating both equal and unequal treatment allocation ratios and extending naturally to settings where strata are defined solely by observed covariates. Innovatively, they construct novel bounds tailored for small or imbalanced strata, and combine them with method-of-moments estimation and a design-consistent closed-form variance estimator to produce tight confidence intervals with accurate coverage. Simulation results demonstrate that the proposed approach substantially narrows interval width and improves inferential precision compared to conventional methods.

attritionheterogeneous treatment sharesLee bounds

When longitudinal outcomes are truncated by death, causal effects become challenging to define and estimate, and existing methods often lack clear causal assumptions and appropriate estimands. This study develops a unified framework that clarifies the definitional challenges and identification assumptions underlying various causal estimands in the presence of truncation by death. It proposes an integrated characterization combining stratum-specific average causal effects with restricted mean survival time, thereby revealing the intrinsically multifactorial nature of the problem. Building on Bayesian inference, the authors derive corresponding estimation procedures and evaluate their performance through simulations and real data from a randomized controlled trial on amyotrophic lateral sclerosis. The results demonstrate that the proposed framework yields a more comprehensive and accurate assessment of treatment effects.

causal estimanddeath censoringlongitudinal study

This study addresses the limited practicality of partial identification arising from overly wide identification regions, which often compels researchers to impose strong assumptions to achieve point identification at the cost of credibility. The authors propose a novel approach that tightens the partially identified set in linear prediction models by incorporating auxiliary moment information—such as means—provided by data publishers, without altering the original interval-censored data. This method systematically integrates moment constraints to substantially shrink the identification region without additional assumptions and elucidates how different moment conditions shape the geometry of the identified set. Leveraging tools from convex geometry and moment-constrained optimization, the paper derives directional measures of identification value under unconditional and conditional mean restrictions. Empirical analysis on survey wage data demonstrates that even minimal moment information can dramatically recover identification power lost due to data coarsening, markedly improving inference accuracy.

auxiliary moment restrictionsdata coarseningidentification region

This study addresses the challenge that existing clinical trial analysis methods struggle to jointly model completers, retrieved dropouts, and missing-at-random participants, often with unclear links between modeling assumptions and target estimands. The authors propose a likelihood-based unified framework that explicitly integrates data from these three participant types for the first time. By combining analysis of covariance with a probit model for treatment discontinuation, the approach clearly aligns its modeling assumptions with the hypothetical and treatment-policy strategies defined in ICH E9(R1). The method employs a maximum likelihood–based efficient estimation algorithm tailored for continuous endpoints and dropout mechanisms. Numerical studies demonstrate that the proposed approach substantially outperforms conventional imputation methods in terms of both bias and variability.

clinical trialscontinuous endpointsestimand framework

This study addresses the fundamental trade-off between bias in a narrow model and variance in a wider model under moderate misspecification—where the true data-generating process includes one additional parameter beyond the fitted narrow model. The authors introduce the concept of a “tolerance radius” to quantify the range of misspecification within which the narrow model yields superior performance. Building on large-sample theory, likelihood-based estimation, and bias–variance decomposition, they develop a novel estimator that achieves robustness and efficiency across both model classes. Theoretical analysis and extensive numerical experiments across multiple model settings demonstrate that the proposed estimator significantly improves estimation accuracy within the tolerance radius, offering a principled balance between robustness to misspecification and statistical efficiency.

estimator robustnessmodel misspecificationnarrow model

Hot Scholars

NL

Nils Lid Hjort

Professor of Mathematical Statistics, University of Oslo
Theoretical and applied statistics and probability theory
FL

Fan Li

Department of Statistical Science, Duke University
statisticscausal inferencecomparative effectiveness researchmissing data
GH

Guanglei Hong

Associate Professor of Comparative Human Development, University of Chicago
causal inferencemultilevel modelinglongitudinal data analysiseducation