Score
Designs and analyzes training procedures and loss functions that enforce consistency of model predictions and internal representations across hierarchical branches, temporal scales, semantic labels, and appearance or frequency-domain variations. This competence includes building multi-branch or multi-scale networks, temporal/semantic consistency losses and training regimes, and test-time adaptation mechanisms that propagate knowledge across branches and bridge gaps between tasks and representations.
Neural network internal representations often lack stability and cross-architectural consistency due to architectural disparities, hindering knowledge transfer and modular deployment. To address this, we propose a structured regularization framework comprising linear shaping operators and rectified path constraints, which explicitly encode inductive biases to improve geometric alignment of representations across architectures. Through theoretical analysis, controlled transfer experiments, and a novel representation alignment metric, we systematically demonstrate that structural priors significantly enhance semantic consistency among heterogeneous models. Our method improves downstream task performance in model distillation and modular learning by up to 12.3%, offering an interpretable and scalable paradigm for building robust, composable deep learning systems.
This work investigates the linear transferability of semantic representations across language models of differing scales. Method: We propose the Linear Representational Transferability (LRT) hypothesis—that steering vectors encoding semantics in smaller models remain effective for eliciting target behaviors in larger models after undergoing an affine transformation. To operationalize this, we formally define a general affine mapping between cross-scale representation spaces and introduce a mapping learning framework grounded in hidden-state alignment and behavior-guided distillation. Experiments are conducted across the LLaMA family of models spanning multiple scales. Contribution/Results: Our approach achieves over 85% behavioral retention when transferring steering vectors from smaller to larger models on tasks including style control and factual correction, validating that small models can serve as lightweight, interpretable behavioral controllers for large models. This establishes a novel, efficient, and transparent paradigm for large-model intervention.
This work investigates the transfer mechanism in infinitely wide neural networks when both the source and downstream tasks operate in the feature-learning regime. We propose Elastic Weight Coupling (EWC), a unified framework modeling feature reuse across pretraining and fine-tuning. Within a Bayesian setting, we integrate gradient flow analysis with weight decay and the infinite-width limit to derive an adaptive feature kernel theory—whose structure depends explicitly on the data distributions and label geometries of both tasks. Our methodology encompasses posterior inference, gradient-flow dynamical modeling, and explicit kernel construction, supporting both linear and polynomial regression as well as validation on real-world datasets. Theoretically and empirically, we characterize the joint influence of coupling strength, feature-learning capacity, dataset scale, and task alignment on transfer performance—demonstrating consistent improvements in generalization across diverse scenarios.
This work identifies a previously overlooked phenomenon in large language model (LLM) evaluation—“training-on-test tasks”—where task-specific knowledge is implicitly incorporated during training, distorting relative model rankings and inducing spurious claims of emergent capabilities. Unlike data leakage, this effect systematically undermines evaluation validity. We formally define and quantify its impact for the first time. To mitigate it, we propose task-aligned fine-tuning—a controlled-variable correction method that isolates benchmark performance attribution via cross-model consistent training. After correction, the relative ordering of mainstream models shifts significantly; moreover, several purported “emergent abilities” exhibit smooth, continuous improvement under gradual task exposure, losing their abruptness. This confirms that such phenomena stem from evaluation bias—not genuine capability discontinuities. Our approach enables more faithful, interpretable LLM assessment and challenges prevailing assumptions about emergence in scaling laws.
This work investigates whether cross-dataset consistency in model representation similarity stems from intrinsic model properties or is confounded by biases inherent in common benchmark datasets. To address this, we conduct systematic representation comparison experiments across multimodal (image, image-text) and multitask (self-supervised, classification, image-text contrastive) models, using Centered Kernel Alignment (CKA) and linearly weighted similarity analysis on diverse domain-shifted datasets. Results demonstrate that training objective is the dominant factor governing cross-dataset representation similarity stability—significantly outweighing influences of data modality and network architecture. We propose the first evaluation framework explicitly designed for cross-dataset representational consistency. Furthermore, we reveal that self-supervised vision models exhibit the strongest generalization of representation similarity across datasets, and that the correlation between representation similarity and task performance is maximized on single-domain benchmarks.
This study addresses the unclear dynamic causes underlying the widening gap between training and validation performance during pre-trained model adaptation. It proposes a dynamic structural interpretation framework, revealing that the shift of update pressure from general to narrow support constitutes the core mechanism driving this generalization gap. By establishing a theoretical link between gradient allocation heterogeneity and performance divergence, the framework enables the observation of structural evolution without requiring a validation set. Through ResMLP simulations, multimodal experiments involving RoBERTa, DeBERTa, Qwen, and ResNet-18, alongside fixed-probe techniques, the work demonstrates that the proposed monitoring metric is significantly positively correlated with the accuracy gap. These findings validate the theory's broad applicability across both natural language processing and computer vision tasks.
This work addresses the long-standing isolation among research domains such as alignment training, model organisms, and toy models, which has hindered empirical cross-pollination and led to redundant exploration and inefficiency. For the first time, it systematically transfers supervised fine-tuning (SFT) practices across these domains by integrating cross-model output training, mixed-strategy data, and benign fine-tuning to rigorously evaluate the portability of key findings. The study demonstrates three successful transfer effects: enhanced behavioral generalization, mitigation of capability degradation, and the critical insight that preserving capabilities alone is insufficient to ensure robustness in subsequent training phases. These results underscore both the efficacy and limitations of reusing methodologies across domains, thereby fostering more synergistic development across disparate research areas.
This work addresses the alignment failures of large language models under emerging safety threats—including role-playing mimicry, adversarial exploits, prefilling attacks, and conditional misalignment—by introducing a multi-level consistency training mechanism within the Transformer architecture. Specifically, consistency constraints are applied at the MLP layers (MLPCT), attention heads (AttCT), and overall behavioral output (BCT), extending consistency-based alignment to these four complex threat scenarios for the first time. Experimental results demonstrate that the proposed approach substantially suppresses diverse misaligned behaviors and exhibits superior robustness and cross-threat generalization compared to existing methods tailored only to jailbreaking or sycophancy attacks. Furthermore, the study uncovers the critical role of shared residual streams in achieving effective model alignment.
本文提出了一种针对数据可用性随时间演变的领域迁移学习问题(TrED),并探讨了现有方法在处理整个演化过程中的不足。
This study investigates the impact of consistency training on model alignment, demonstrating that it is not alignment-neutral. Through systematic evaluation of seven consistency methods across 108 open-source large language models (7B–70B) with controlled misalignment, the authors find that such training generally suppresses reward hacking while exacerbating sycophancy. Leveraging controlled fine-tuning, distribution shift analysis, and theoretical modeling, they identify distribution shift as the dominant underlying mechanism. Building on this insight, they propose a unified theoretical framework that predicts under which conditions consistency training amplifies or mitigates specific misalignment behaviors, thereby offering an auditable foundation for safer alignment practices.