Bayes-Sufficient Representations in Supervised Learning

๐Ÿ“… 2026-06-02
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work investigates the minimal input information required to achieve Bayes-optimal prediction in fixed supervised learning tasks. To this end, it introduces the notion of a *Bayes-sufficient representation*โ€”a representation that admits a predictor attaining Bayes-optimal performanceโ€”and formalizes the *Bayes-minimal representation* via the Bayes quotient, which precisely characterizes the information that must be preserved. Theoretically, the study establishes a systematic connection among representation learning, loss functions, Bayes-optimal actions, and attribute elicitation theory, revealing fundamental differences in the information requirements across distinct losses. Empirically, through neural network bottleneck analysis, controlled synthetic data, and real-world experiments on iNaturalist, the work validates the sufficiency and minimality of such representations, demonstrates the ability to distinguish redundant from essential information, and confirms that the minimal information needed for optimal prediction is jointly determined by the data distribution and the choice of loss function.
๐Ÿ“ Abstract
Representation learning is often described as preserving the information in an input that is relevant for prediction. This work asks what relevance means for a fixed supervised decision problem. A representation is defined to be Bayes-sufficient for a joint distribution and loss if some prediction head can use it to implement a Bayes-optimal action rule. This makes the target information loss-dependent. In the almost-surely unique Bayes-action case, the relevant object is a Bayes quotient, which identifies inputs that require the same Bayes-optimal action. A representation is sufficient when it refines this quotient, and Bayes-minimal when it is informationally equivalent to it. The framework connects naturally to property elicitation: zero-one loss requires the Bayes class, squared loss the conditional mean, Brier loss the conditional probability in binary prediction, and log loss or strictly proper scoring rules the predictive distribution. Controlled finite experiments, learned neural bottleneck experiments, and a real-data iNaturalist taxonomic refinement experiment illustrate the distinction between sufficiency, minimality, and retained non-required information. For a fixed supervised problem, the distribution and the loss determine the Bayes action, the Bayes action determines the quotient, and the quotient determines the minimal information required for Bayes-optimal prediction.
Problem

Research questions and friction points this paper is trying to address.

Bayes-sufficiency
representation learning
supervised learning
loss function
Bayes-optimal prediction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Bayes-sufficient representation
Bayes quotient
supervised representation learning
proper scoring rules
Bayes-minimal representation
V
Vasileios Sevetlidis
Athena Research Center, Kimmeria Campus, Xanthi, Greece; Democritus University of Thrace, Vas. Sofias Campus, Xanthi, Greece; International Hellenic University, Serres, Greece