marginal contribution estimation

Methods for estimating the incremental (marginal) effect of a dataset or action on model performance or reward without full retraining or sharing raw data. It encompasses credit-assignment techniques that quantify individual or synergistic contributions to multi-task metrics and answer quality.

marginalcontributionestimation

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Treatment Effects of Multi-Valued Treatments in Hyper-Rectangle Model

Sep 05, 2025
XT
Xunkang Tian
🏛️ European Research University

This paper addresses the identification of marginal treatment responses (MTRs) in multivalued treatment models, relaxing the restrictive assumptions of conventional hyper-rectangular frameworks—namely, known treatment thresholds and full determination of treatment choice by unobserved heterogeneity. We introduce an *ordered treatment assumption*, which permits unknown thresholds and partial independence between unobserved heterogeneity and treatment selection, enabling point or set identification of MTRs under more realistic conditions. Within this framework, we systematically derive policy-relevant treatment effects—including the marginal average treatment effect and rank-order effects—and develop nonparametric specification tests to assess policy effectiveness. Our approach enhances both the empirical applicability and credibility of multivalued treatment models in causal policy analysis.

Enables policy evaluation through treatment effect testingIdentifies marginal treatment responses in multi-valued modelsRelaxes restrictive assumptions of hyper-rectangle framework

Transfer Estimates for Causal Effects across Heterogeneous Sites

May 02, 2023
KM
Konrad Menzel
🏛️ New York University

This study addresses the problem of extrapolating causal effects from multi-site randomized controlled trials (RCTs) to a new target site with baseline survey data only. To handle site-level population heterogeneity and unobserved confounding, we propose modeling baseline covariates as functional data—thereby capturing site-specific confounding structures—for the first time. We then develop a design-oriented, nonparametric method to construct an optimal finite-dimensional feature space, ensuring optimal convergence rates for conditional average treatment effect (CATE) estimation. Our approach integrates functional data analysis, nonparametric regression, and causal transfer learning theory. Evaluated across five integrated multi-site RCTs on cash transfer programs, the method significantly improves prediction accuracy of treatment effects at target sites and quantifies the estimation gain attributable to adaptive transfer.

Adapting experimental estimates to target site characteristicsDetermining optimal feature space for causal predictionExtrapolating treatment effects across heterogeneous populations

Estimation and Inference for the Average Treatment Effect in a Score-Explained Heterogeneous Treatment Effect Model

Apr 23, 2025
KC
Kevin Christian Wibisono
🏛️ University of Michigan | Boston University

Real-world policy interventions—such as cutoff-based eligibility rules (e.g., admission thresholds or income criteria)—often induce non-random treatment assignment, rendering conventional regression discontinuity designs (RDD) inefficient: they discard observations away from the cutoff, leading to information loss and slow convergence. Crucially, most existing methods assume homogeneous treatment effects, contradicting empirical evidence of effect heterogeneity. This paper proposes a novel ATT estimator integrating difference-in-means with covariate matching, enabling full-sample utilization under heterogeneous treatment effects for the first time. The method simultaneously delivers nonparametric estimates of both conditional average treatment effects (CATE) and individual treatment effects (ITE). We establish asymptotic normality of the estimator, permitting valid robust inference. Extensive simulations and empirical applications demonstrate substantial gains in estimation accuracy and statistical power, effectively overcoming the dual limitations of RDD—information inefficiency and restrictive homogeneity assumptions.

Estimating average treatment effect with heterogeneous effectsOvercoming limitations of traditional cutoff-focused methodsProviding non-parametric CATE and ITE estimates

Estimating Model Performance Under Covariate Shift Without Labels

Jan 16, 2024
JB
Jakub Bialek
🏛️ NannyML NV | AI Institute | University of Waikato | LTCI | Telecom Paris | IP Paris

To address the challenge of unsupervised model performance estimation under covariate shift—where ground-truth labels are unavailable or delayed post-deployment—this paper proposes the Probability-Adaptive Performance Estimation (PAPE) framework. PAPE requires neither access to true labels nor knowledge of the original model’s architecture or feature representations; it operates solely on the model’s probabilistic outputs and confidence scores. By jointly leveraging density ratio estimation and performance generalization bound theory, PAPE models prediction distributions and applies adaptive reweighting to yield unbiased estimates of arbitrary classification metrics—without assuming a specific shift form or resorting to feature learning or generative modeling. Extensive evaluation across 900+ real-world census dataset–model combinations demonstrates that PAPE reduces mean absolute error by 37% compared to state-of-the-art proxy metrics and drift detection methods, significantly enhancing the reliability and generality of model monitoring in production environments.

Addressing performance degradation from data distribution shiftsEstimating model performance under covariate shift without labelsEvaluating binary classification models on unlabeled tabular data

Precise High-Dimensional Asymptotics for Quantifying Heterogeneous Transfers

Oct 22, 2020
FY
F. Yang
🏛️ Tsinghua University | Beijing Institute of Mathematical Sciences and Applications | Northeastern University | Stanford University | University of Pennsylvania

This work addresses the fundamental question in transfer learning: *when does jointly training on source and target task data outperform using target data alone?* Focusing on high-dimensional linear regression, we quantify positive and negative transfer effects. Within a hard parameter-sharing framework, we derive exact asymptotic expressions for the bias and variance of the estimator under the proportional limit—constituting the first such characterization. We uncover a phase transition in transfer risk driven by model shift and rigorously identify the critical condition separating positive from negative transfer. Our theoretical results extend to multi-task settings and are validated under random-effects models. This provides the first analytical feasibility criterion for transfer learning, filling a key theoretical gap in the precise, quantitative analysis of transfer performance across heterogeneous, high-dimensional tasks.

Analyzing phase transitions in transfer learning under model shiftsIdentifying conditions for positive or negative transfer in high-dimensional regressionQuantifying transfer learning performance between source and target tasks

Latest Papers

What's happening recently
View more

This work addresses the challenge of fairly attributing credit among multiple creators of AI-generated content—such as code, news articles, or short videos—within a contextual window. To this end, we propose an incentive-compatible credit allocation mechanism grounded in cooperative game theory, specifically leveraging the least core to guarantee that no subset of contributors is significantly undervalued. As the first study to apply the least core to contextual credit assignment, we introduce an efficient approximation algorithm that integrates constraint seeding and constraint separation techniques, substantially reducing the number of large language model (LLM) queries required. Empirical evaluation on a web retrieval credit allocation task demonstrates that our method achieves a high-quality approximation of the least core with orders of magnitude fewer LLM calls compared to baseline approaches.

AI-generated contentcooperative game theoryin-context credit assignment

This work addresses the challenge of predicting performance changes when a source-domain model is replaced by a new one. To this end, the authors propose TRACE, a novel framework that, for the first time, decomposes the risk difference between two models under covariate shift into four interpretable components: two generalization gaps, a model change penalty, and a covariate shift penalty. The framework establishes a computable upper bound to diagnose the causes of performance degradation. TRACE estimates model sensitivity via high-quantile input gradients, quantifies data distribution shift using either optimal transport (OT) or maximum mean discrepancy (MMD), and measures model change through output distances on target samples. Experiments demonstrate that TRACE’s diagnostic scores exhibit strong monotonic correlation with actual performance degradation and achieve superior performance in deployment gating, as measured by AUROC and AUPRC, thereby enabling label-efficient and safe model replacement.

covariate shiftdistribution shiftmodel replacement

This study addresses the disconnect between academic research and industrial practice in treatment effect estimation, where prevailing evaluation paradigms hinder real-world applicability. Through a large-scale empirical analysis, we systematically compare diverse meta-learners, base learners, and specialized causal models across semi-synthetic benchmarks and real-world datasets. Our findings reveal a pronounced inconsistency between counterfactual and observable performance metrics, and demonstrate that model rankings derived from semi-synthetic data fail to generalize to real settings. Notably, simple meta-learners paired with strong base models consistently outperform purpose-built causal models on real data, underscoring the critical importance of validation on real-world outcomes and observable metrics. These results challenge the dominant reliance on semi-synthetic evaluations and call for a paradigm shift toward more empirically grounded assessment protocols.

counterfactual metricsevaluation gapreal-world datasets

Paid media attribution often overestimates true incremental lift in the presence of channel overlap, leading to inaccurate ROI assessment and suboptimal budget allocation. This work proposes the first calibration framework that integrates incremental experiments with large-scale attribution systems: leveraging causal insights from incrementality trials as anchors, it translates sparse lift observations into daily corrected estimates and allocates cross-channel cannibalization effects under business-level hierarchical constraints. Combining causal inference, hierarchical optimization, and machine learning–based attribution models, the approach substantially reduces calibration error in offline evaluations. Following global deployment across multiple markets, it has driven strategic budget reallocations and achieved an empirical reduction in measured cannibalization rates by approximately 15 percentage points, effectively bridging the gap between attribution and true incrementality.

advertisingattributionbudget allocation

To address the limited external validity of individual treatment effect (ITE) estimation in few-shot and cross-domain settings, this paper proposes TL-TARNet—the first systematic investigation into the applicability boundaries of transfer learning within the causal model TARNet. Methodologically, it achieves robust transfer of large-sample causal representations from source to target (few-shot) domains via representation alignment and weight fine-tuning, under both random and non-random treatment assignment. Theoretically, it characterizes how transfer efficacy depends on source data quality and scale, and substantially mitigates small-sample bias under unbiasedness constraints. Empirical results show a 32% reduction in ITE estimation error and over 40% bias attenuation in simulations. In the IHDS-II application, TL-TARNet yields more robust ITE estimates of maternal firewood-collection time on children’s study duration, markedly improving reliability for causal extrapolation with limited target-domain data.

Addresses limited applicability in small datasets via knowledge transferImproves individual treatment effect estimation with transfer learningReduces bias and error in causal inference across diverse settings

Hot Scholars

ME

Maria Elena Filippin

Department of Economics, Uppsala University
Monetary PolicyFinancial Stability
AA

Alexis Akira Toda

Emory University
Macro-financeAsset price bubblesPower lawMathematical economics
WY

Wayne Yuan Gao

Department of Economics, University of Pennsylvania
EconometricsMicroeconomic TheoryNetworks
YK

Yiannis Karavias

Brunel University of London
EconometricsPanel DataStructural BreaksThreshold Regression