Score
Designs, implements, and evaluates lightweight side classifiers (simple auxiliary/auxiliary/side classifiers) that are attached to intermediate layers of a model to provide additional supervisory signals during training; and specifies how their architectures, loss terms, and joint-training procedures encourage discriminative intermediate feature learning and improve final classification performance.
To address catastrophic forgetting—a critical bottleneck limiting both performance and efficiency in continual learning—this paper systematically identifies, for the first time, the strong forgetting-resistance of intermediate-layer representations. Building upon this insight, we propose a plug-and-play auxiliary classifier (AC) architecture that requires no modification to the backbone network or training pipeline. By embedding lightweight classifiers at intermediate layers and integrating an early-exit inference mechanism, our method jointly optimizes accuracy and efficiency. Under a multi-stage continual learning evaluation framework, it achieves an average relative accuracy gain of 10%, reduces inference computational cost by 10–60%, and preserves original accuracy—while remaining fully compatible with mainstream continual learning paradigms. Our core contributions are threefold: (i) uncovering the previously unrecognized forgetting-resistance of intermediate representations; (ii) introducing a modular, plug-and-play AC architecture; and (iii) establishing a new continual learning paradigm that simultaneously balances accuracy and efficiency.
This study addresses the unclear generalization mechanisms in auxiliary learning by developing an analytical nonlinear network fluctuation-dissipation theory within a teacher-student framework. We derive the online stochastic gradient descent (SGD) dynamical equations to systematically quantify the effects of task relatedness and gradient noise on generalization performance. By combining analytical solutions of differential equations with empirical validation, this work reveals the intrinsic relationship between main-auxiliary task errors and single-task errors, while elucidating the dynamical mechanism through which moderate gradient noise enhances generalization. Ultimately, this research provides a rigorous theoretical foundation for understanding implicit regularization effects in multi-task learning.
To address the insufficiency of hard-label supervision in few-shot classification, this paper proposes leveraging the distributional structure—rather than mere confidence scores—of non-ground-truth classes as auxiliary supervisory signals. Specifically, soft labels are constructed via affine combinations and jointly optimized over both direction and step size within the probability simplex. Theoretically, we establish for the first time how such distributional information fundamentally influences the convergence rate and asymptotic value of the generalization error bound. Mechanistically, we reveal its complementary role with mixing coefficients in soft-label optimization. Extensive experiments demonstrate that the proposed method significantly enhances generalization: it achieves average accuracy improvements of 2.3–5.7 percentage points across multiple standard few-shot benchmarks.
本文通过引入可学习的稀疏线性变换作为规则条件,结合梯度提升方法,解决了传统规则集成模型在缺乏精心设计特征时准确性与解释性之间的矛盾。
In model merging, representational misalignment—modeled as an orthogonal transformation—exists between fused outputs and fine-tuned classifiers in the feature space, leading to evaluation distortion and suboptimal performance. To address this, we propose a novel few-shot unsupervised classifier alignment paradigm: using only a small number of unlabeled samples, it calibrates classifier weights via orthogonal transformation to achieve feature-space alignment. Based on this, we establish a more reliable evaluation protocol for merged models. Experiments across multiple classification tasks demonstrate that our method significantly improves the accuracy of merged models and yields evaluations that more faithfully reflect the intrinsic capabilities of merging methods. This work introduces a new benchmark for model merging that jointly ensures effectiveness and evaluability.
This work addresses the challenge that privileged information—available during training but inaccessible at deployment—can mislead models when it is noisy or weakly informative. To mitigate this issue, the authors propose a joint training framework that simultaneously optimizes a teacher model leveraging privileged information and a student model restricted to inputs available at test time. Through an innovative coupling mechanism and an alternating optimization algorithm, the student selectively distills useful knowledge from the teacher while avoiding the propagation of its errors. Theoretical analysis establishes conditions under which this joint training improves accuracy and supports efficient implementation even for high-dimensional, large-scale models. Experiments on both synthetic and real-world datasets demonstrate that the proposed method significantly outperforms conventional two-stage baselines and exhibits robustness to low-quality privileged information.
Traditional differentiable decision trees for regression tasks struggle to jointly optimize internal and leaf nodes due to reliance on approximation strategies such as boundary smoothing or gradient quantization, often leading to overfitting. This work proposes DTSemNet, a novel framework that exactly represents hard oblique decision trees through a semantically equivalent and invertible neural architecture, enabling end-to-end gradient-based training without approximations. To further enhance gradient precision for regression, the method introduces an annealed Top-k mechanism that provides accurate routing signals during training. DTSemNet is the first approach to achieve approximation-free differentiable training of oblique decision trees, outperforming existing methods on both classification and regression benchmarks. Moreover, it demonstrates practical utility in reinforcement learning by serving as an interpretable, programmatic policy.
This study addresses the vanishing gradient problem caused by symmetry in constructive classifiers, where deepening soft decision trees leads to inherited parent-node distributions and a consequent loss of learning capacity. We systematically diagnose four structural growth strategies for soft decision trees, employing theoretical derivations and multi-dataset benchmarks to reveal the underlying zero-gradient defect mechanism. To overcome this limitation, we propose a perturbation-based remedy utilizing slight asymmetric initialization. Theoretically, we rigorously prove both the root cause of this defect and the efficacy of our proposed solution. Empirically, the remedied model achieves a 0.2% accuracy improvement, while sparse growth strategies approximate full-tree performance using only 23% of the splits. Furthermore, this work delineates the applicability boundaries of each growth strategy, offering practical guidance for constructing deeper and more efficient soft decision trees.
This work addresses the challenge of end-to-end machine learning inference on microcontroller-class edge devices under stringent constraints on memory, energy consumption, and latency. To bridge the gap between conventional machine learning pipelines and embedded deployment realities, the authors propose a robust design framework tailored for resource-constrained environments, encompassing data acquisition, preprocessing, model compression, and streaming deployment. The framework integrates sampling buffers, feature dimensionality reduction techniques (e.g., RMS, spectral features, MFCCs), validation strategies for class imbalance, and co-optimization of models with runtime systems to form a complete embedded ML pipeline. Experimental evaluations on two representative tasks—inertial human activity recognition and keyword spotting—demonstrate that the proposed approach enables efficient, practical, and robust on-device inference, significantly narrowing the divide between general-purpose machine learning methodologies and embedded implementation requirements.
This work addresses the limitation of conventional pooling operations—such as max and average pooling—in discarding discriminative information during downsampling. To mitigate this issue, the authors propose FlexPooling, an adaptive pooling mechanism that generalizes average pooling into a learnable weighted formulation, optimized end-to-end alongside the main network. A lightweight Simple Auxiliary Classifier (SAC) is further introduced to collaboratively guide the learning of pooling weights, thereby enhancing the preservation of salient features. Experimental results demonstrate that FlexPooling consistently improves model accuracy by 1%–3% across multiple image classification benchmarks, significantly outperforming baseline pooling strategies.