🤖 AI Summary
This study addresses how to enhance the accuracy of standard classifiers without incurring additional inference overhead. To this end, it proposes a dimensionality-expansion reconstruction framework for supervised classification that inserts class prototypes at the semantic interface to construct a latent Gaussian mixture model. The framework is jointly trained using a quadratic consensus penalty, gradient isolation, and a sampled classification loss, while the prototypes are removed during inference to achieve structural decoupling. This mechanism significantly improves performance on both ResNet and Vision Transformer backbones without requiring modifications to their original architectures. Experiments on CIFAR benchmarks demonstrate accuracy gains of up to five percentage points over non-expanded variants, supported by comprehensive theoretical guarantees.
📝 Abstract
We propose a lifted reformulation of supervised classification that improves the final accuracy of standard classifiers without changing the architecture at inference time. A network $N=N_2\circ N_1$ is split at a single semantic interface and one learnable prototype per class is inserted there. Training combines a quadratic consensus penalty that pulls $N_1(x)$ toward the prototype of its class with a classification loss of $N_2$ evaluated on samples drawn around the prototypes, whereat no gradient crosses the interface. At inference the prototypes are discarded and the unmodified network $N_2\circ N_1$ is used. Across CIFAR-10, CIFAR-100, and TinyImageNet with ResNet and vision transformer backbones, lifted training improves test accuracy by up to five percentage points over variants without lifting under a shared tuning protocol. Moreover, we provide theoretical justification of those results.