Adaptive Mean Estimation by In-Context Learning: A Gradient-Flow Analysis

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses how prior-fitting networks achieve statistically adaptive mean estimation under unknown data distributions. The proposed approach leverages a Transformer architecture with Softmax mixture-of-experts, employing gradient flow theory to analyze how the attention mechanism distinguishes between Gaussian and uniform distributions. By proving that the Softmax operation computes derivatives of the cumulant generating function, this work rigorously disentangles expert, routing, and normalization errors to elucidate the underlying interpolation estimator principle. The primary contribution lies in theoretically characterizing the end-to-end specialization mechanism. Empirically, experiments demonstrate that the model attains asymptotic efficiency on Gaussian tasks while approximating minimax rates on uniform tasks, thereby validating its distribution-adaptive capabilities.
📝 Abstract
Prior Fitted Networks (PFNs) such as TabPFN now rival established statistical procedures across prediction and estimation tasks. A natural explanation is that PFNs have the property of statistical adaptivity, that is, they perform nearly as well as a method tailored to the true data-generating model for a heterogeneous set of models, while not being told which model the data comes from. We study how such adaptivity is learned in a controlled location-estimation problem. Each task is an unlabeled sample whose family is hidden: Gaussian data call for averaging, with error of order $n^{-1}$, whereas uniform data are best estimated from their extremes, at the faster rate $n^{-2}$. We also provide the example of a symmetric Gaussian mixture, for which a rate of $σ^2_n/n$ can be attained. On scalar inputs, softmax attention computes the derivative of the empirical cumulant-generating function. A single primitive therefore both supplies features that distinguish the families and forms estimators interpolating between the sample mean and the mid-range. We combine attention experts through either a softmax mixture of experts or a gated linear unit (GLU), and analyze stagewise gradient flow. With $\widetildeΩ(n^{1+ε})$ pretraining tasks, the learned estimator is asymptotically efficient on Gaussian tasks, within a factor $n^ε$ of the minimax rate on uniform tasks, and order-optimal on mixtures in a shrinking-variance regime. These guarantees extend to new locations and longer contexts. A risk decomposition separates expert error, routing error and normalization error, which clarifies the architectural contrast. Softmax gating enforces normalization and exact translation equivariance, whereas the GLU must learn it: its dynamics separate into fast bias removal followed by slow expert selection. End-to-end experiments recover the predicted specialization.
Problem

Research questions and friction points this paper is trying to address.

In-Context Learning
Statistical Adaptivity
Prior Fitted Networks
Mean Estimation
Gradient Flow
Innovation

Methods, ideas, or system contributions that make the work stand out.

In-Context Learning
Prior Fitted Networks
Gradient Flow
Attention Mechanism
Mixture of Experts
🔎 Similar Papers
No similar papers found.