Fidelity Isn't Accuracy: When Linearly Decodable Functions Fail to Match the Ground Truth

📅 2025-06-13
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work investigates whether regression neural networks learn functions that are well-approximated by linear models, and whether linear decodability—i.e., the extent to which network outputs can be recovered via linear projection—is equivalent to predictive accuracy. To this end, we propose λ(f), a formal metric quantifying linear decodability as the R² between the neural network’s predictions and those of its optimal linear proxy. Through systematic experiments on synthetic data (x sin(x) + ε) and real-world regression benchmarks (e.g., Medical Insurance), we empirically establish—and for the first time formally define—that high λ(f) does not imply low true prediction error; thus, linear decodability ≠ high predictive accuracy. This challenges the implicit assumption that linear surrogates suffice for explaining nonlinear regression models, exposing critical limitations of such explanations in high-stakes settings. λ(f) provides a reproducible, quantitative, and comparable framework for assessing intrinsic linear structure within learned regression functions.

Technology Category

Machine Learning: Classification and RegressionNatural Language Processing: Interpretability, Analysis, and Evaluation of NLP ModelsReasoning under Uncertainty: Graphical Models

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
📝 Abstract
Neural networks excel as function approximators, but their complexity often obscures the nature of the functions they learn. In this work, we propose the linearity score $lambda(f)$, a simple and interpretable diagnostic that quantifies how well a regression network's output can be mimicked by a linear model. Defined as the $R^2$ between the network's predictions and those of a trained linear surrogate, $lambda(f)$ offers insight into the linear decodability of the learned function. We evaluate this framework on both synthetic ($y = x sin(x) + epsilon$) and real-world datasets (Medical Insurance, Concrete, California Housing), using dataset-specific networks and surrogates. Our findings show that while high $lambda(f)$ scores indicate strong linear alignment, they do not necessarily imply predictive accuracy with respect to the ground truth. This underscores both the promise and the limitations of using linear surrogates to understand nonlinear model behavior, particularly in high-stakes regression tasks.
Problem

Research questions and friction points this paper is trying to address.

Quantify linear decodability of neural network functions
Assess alignment between network predictions and linear surrogates
Evaluate limitations of linear surrogates in high-stakes regression
Innovation

Methods, ideas, or system contributions that make the work stand out.

Introduces linearity score λ(f) for interpretability
Uses linear surrogate models for evaluation
Assesses linear decodability versus ground truth accuracy
🔎 Similar Papers