conditional expectation estimation

Design, build, or analyze estimators that approximate the conditional expectation E[Y | X] (including procedures induced from pretrained models such as LLMs) from observed inputs and outputs; evaluate and tune these estimators for squared‑loss performance relative to the Bayes optimal predictor, ensure downstream inferences depend only on the conditional mean, and study their large‑sample behavior (convergence to irreducible variance plus any estimator bias).

conditionalexpectationestimation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.21
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Traditional social science experiments rely on human subjects, which are costly, inefficient, and susceptible to sampling bias, thereby hindering accurate estimation of statistical quantities such as conditional expectations. This work formalizes pre-trained large language models (LLMs) as misspecified function estimators and establishes, for the first time, their risk equivalence to the Bayes optimal estimator within a restricted function class. By introducing a decomposition framework that separates representation bias from optimization error—and leveraging Le Cam’s two-point method, Pinsker’s inequality, and finite-sample concentration bounds—the study rigorously demonstrates that, under appropriate range conditions and proper model calibration, LLMs can asymptotically approach the Bayes optimal risk for estimating the conditional mean of human responses, substantially reducing reliance on expensive human experimentation.

Bayes optimal riskconditional expectationhuman-response data

To address low estimation accuracy of causal effects under high-dimensional confounders, this paper proposes a novel method integrating large language models (LLMs) with double machine learning (DML). Specifically, the LLM leverages its semantic understanding and generative capabilities on historical auction texts to directly model and predict conditional expectation functions; the resulting outputs serve as enhanced covariate embeddings within the DML framework. This approach circumvents information loss inherent in conventional high-dimensional feature engineering and embedding compression, thereby mitigating the curse of dimensionality. Empirical evaluation on a small-sample online jewelry auction dataset demonstrates that, compared to methods relying solely on pretrained embeddings, the proposed approach significantly improves both accuracy and stability of average treatment effect (ATE) estimation. The results validate the feasibility and advantages of systematically integrating LLM-derived prior knowledge with structured causal inference frameworks.

Enhancing double machine learning through LLM predictions as additional predictorsImproving causal estimation efficiency using LLM-generated conditional expectationsOvercoming dimensionality curse in causal inference with generative models

This work addresses the high cost of large model fine-tuning by tackling the challenge of accurately predicting post-fine-tuning performance beforehand—a task whose theoretical limits remain unclear. We formulate pre-fine-tuning performance prediction as a stochastic estimation problem under information constraints and introduce a predictive risk decomposition framework that separates it into an irreducible intrinsic limit and an optimizable variance term, thereby revealing fundamental bounds on predictability. Leveraging information theory and statistical learning theory, we establish a theoretical lower bound on variance decay through optimization and construct a predictability phase diagram that delineates three distinct task regimes. Experiments on both synthetic and real-world benchmarks validate the efficacy of this phase diagram, and our proposed budget-optimal probing strategy significantly enhances prediction efficiency, offering both theoretical grounding and practical tools for pre-fine-tuning decision-making.

fine-tuning costLLMspre-hoc prediction

Adapting to Misspecification

May 23, 2023
TB
Timothy B. Armstrong
🏛️ USC | UC Berkeley | NBER | UCL | CEMFI

This paper addresses the trade-off between robustness and efficiency under model misspecification, proposing an adaptive estimation framework that does not require a pre-specified upper bound on bias. The core challenge is to construct an estimator whose worst-case risk—relative to an oracle knowing the true bias bound—is minimized. Methodologically, we formulate an adaptive shrinkage estimator via weighted convex minimax optimization, calibrated against the oracle risk, and develop a lookup-table-based fast algorithm. Theoretically, our approach departs from conventional hypothesis-testing paradigms and achieves, for the first time, direct adaptation to the degree of misspecification. Empirically, the method substantially improves estimation accuracy and robustness across multiple canonical studies, offering both strong theoretical guarantees and practical computational efficiency.

Adapting to unknown bias bounds in restricted estimatorsBalancing robustness and efficiency in parameter estimationSolving weighted convex minimax for adaptive estimation

Latest Papers

What's happening recently
View more

This work proposes a decision-theoretic neural pretraining framework to address key challenges in time series analysis, including finite-sample bias, poor calibration, and forecast combination. By jointly modeling the data-generating process and decision objectives within a simulated environment, the method trains neural networks via hierarchical simulation to approximate optimal decision rules, enabling high-quality zero-shot inference without real-world data. The approach innovatively integrates decision theory with deep learning, allowing explicit control over risk, bias, minimax performance, and calibration consistency, thereby effectively solving problems that are analytically intractable or computationally prohibitive. Empirical results demonstrate substantial improvements over conventional methods such as maximum likelihood estimation and AICc in AR(p) modeling and forecast combination tasks, while achieving competitive or superior performance against state-of-the-art statistical and deep learning models on real-world benchmarks.

decision-theoretic inferencefinite-sample biasforecast combination puzzle

Traditional sensitivity analyses in matched observational studies often assume that unobserved confounding is nearly perfectly correlated with potential outcomes, rendering them overly conservative and lacking realistic flexibility. This work proposes a stochastic sensitivity analysis framework that models unobserved confounding as a random variable with an unknown conditional distribution given the potential outcomes and observed covariates. Rather than optimizing over worst-case realizations, the approach evaluates the robustness of causal conclusions by optimizing over the least favorable conditional distributions. By introducing controlled randomness, the method permits imperfect alignment between unobserved confounders and potential outcomes and incorporates both nonparametric interpretable distribution classes and Bernoulli conditional models in the optimization. Empirical results demonstrate that even minimal stochasticity substantially enhances the ability to report robustness against hidden bias.

hidden biasmatched observational studiessensitivity analysis

This work addresses the limitations of deterministic regression in settings involving coarse-graining, partial observability, or inverse problems, where input–output relationships are inherently one-to-many and conditional distributions exhibit irreducible stochasticity. To diagnose these challenges under finite data, the authors introduce a framework centered on the “conditional mean barrier,” proposing two diagnostic tools: a residual–feature orthogonality test and an upper-bound analysis of the coefficient of determination. These tools effectively disentangle model underfitting from irreducible conditional variance. Leveraging this diagnostic framework, the study systematically evaluates distribution-learning approaches—including negative log-likelihood, moment matching, variational objectives, adversarial divergences, and score matching—on benchmark problems such as bimodal distributions and multiscale Lorenz-96 closure tasks. Empirical results demonstrate that the framework clearly identifies the inadequacies of deterministic models and reveals the true variability of underlying conditional distributions.

aleatoric uncertaintyconditional-mean barrierdistribution learning

This study addresses the estimation of history-dependent conditional prediction revision scales in sequential models—the magnitude by which predictions are updated upon observing new data—a quantity inherently unobservable due to its dependence on unknown conditional means. We systematically evaluate block bootstrap, conditional heteroskedasticity models, state-space filters, O(1) streaming smoothers, and pretrained RNN forget gates across varying structural assumptions and computational budgets. Theoretical and empirical analyses reveal that lagged smoothers are inconsistent under rapid dynamics, while structurally aligned state-space filters can surpass conventional convergence rate limits. Although neural forget gates do not explicitly encode this scale, it can be effectively decoded via linear probing. These findings motivate a practical guideline: “identify structure, match method, choose minimal cost.” In volatility-driven settings, conditional variance models outperform block bootstrap by orders of magnitude in both speed and accuracy; in state-driven scenarios, only structurally matched filters reliably track abrupt changes.

conditional forecast-revision scaleconditional second momentestimation

Hot Scholars

RN

Razieh Nabi

Rollins Assistant Professor of Biostatistics, Emory University
Causal InferenceMissing DataAlgorithmic FairnessGraphical Models
YL

Yiping Lu

Assistant Professor, Northwestern University
Scientific Machine LearningMachine LearningStatisticsStochastic Control
YR

Yinuo Ren

ICME, Stanford University
Applied and Computational Mathematics
DS

Dino Sejdinovic

Professor of Statistical Machine Learning, Adelaide University
Machine LearningStatistical InferenceInformation Theory
KG

Kun Gai

Senior Director & Researcher, Alibaba Group
Machine LearningComputational Advertising