Detecting a Shift Is Not Enough: Exact Minimax Limits of Linear Representation Repair

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that mean shifts are readily detectable yet difficult to remove, leaving the statistical limits of linear representation repair unclear. To investigate this, the authors formulate shift removal as a statistical decision problem, employing linear mappings and Poisson modeling integrated with paired calibration measurement techniques, and validate their framework using clinical EEG data. The primary contribution is the first derivation of a closed-form solution for the finite-sample minimax risk, which reveals a fundamental theoretical gap between detection and repair. Furthermore, this work demonstrates that projection-based methods achieve optimality without requiring prior knowledge. By accurately predicting residual shifts in unseen data, these findings provide rigorous theoretical guidance for evaluating and improving concept erasure mechanisms.
📝 Abstract
A mean shift between two data sources can be easy to detect but hard to remove without substantially changing their representations. We cast its removal as a statistical decision problem: from noisy differences between paired calibration measurements in $\mathbb{R}^d$, learn one linear map, applied to both sources under a hard distortion budget, that leaves as little of the shift as possible on fresh data. We derive the exact finite-sample minimax risk over all such maps, $(d-k) \mathbb{E}[1/(d+2J)]$ with $J\sim\mathrm{Pois}(κ/2)$, where the budget allows deleting $k$ directions and $κ$ is the calibration signal-to-noise ratio. Projecting out the mean calibration difference attains it without knowing $κ$ or the noise scale. This exposes a detection-repair gap: detecting the shift needs only $κ\gg\sqrt d$, whereas removing a fixed fraction of it at constant distortion needs $κ\asymp d$, as for estimating its direction. Standard linear concept erasers (MP, SAL, LEACE) remove the same calibration difference, so the formula gives, before fitting, exactly how much shift they leave on fresh data and how much calibration a target requires. The limit is robust: pairing keeps it exact for non-Gaussian shared content, the projection keeps its guarantee under anisotropic noise, and selective abstention cannot close the gap. On paired clinical and wearable sleep EEG, where differences between participants act as calibration noise, the formula predicts the device shift left in new participants, and more recordings per person soon stop helping. Together, these results tell whether a correction that falls short needs a better method, more recordings, or more participants.
Problem

Research questions and friction points this paper is trying to address.

mean shift removal
linear representation repair
minimax risk
detection-repair gap
concept erasure
Innovation

Methods, ideas, or system contributions that make the work stand out.

Minimax risk
Linear representation repair
Detection-repair gap
Concept erasure
Mean shift removal
🔎 Similar Papers
No similar papers found.
A
Anuar Aimoldin
Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)
Yankai Chen
Yankai Chen
Postdoctoral Associate, Cornell University
Information RetrievalKnowledge MiningLarge Language ModelsAgentic AI
A
Ayana Mussabayeva
Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)
X
Xue Liu
Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)