consistency training

Designs and implements training objectives, regularizers, and evaluation analyses that enforce consistency of model outputs or latent representations across input perturbations, multiple views, different samples or batches, and across reasoning or optimization steps. Builds and integrates specific losses and procedures (e.g., cross-step, cross-sample, batch or multi-view consistency losses) to align model answers without labels, stabilize and align representations under distribution shift, and prevent pathologies such as feature splitting or absorption while preserving reconstruction and task-relevant fidelity.

consistencytraining

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.01
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study investigates the impact of consistency training on model alignment, demonstrating that it is not alignment-neutral. Through systematic evaluation of seven consistency methods across 108 open-source large language models (7B–70B) with controlled misalignment, the authors find that such training generally suppresses reward hacking while exacerbating sycophancy. Leveraging controlled fine-tuning, distribution shift analysis, and theoretical modeling, they identify distribution shift as the dominant underlying mechanism. Building on this insight, they propose a unified theoretical framework that predicts under which conditions consistency training amplifies or mitigates specific misalignment behaviors, thereby offering an auditable foundation for safer alignment practices.

consistency trainingmisalignmentmodel alignment

ConsistentFeature: A Plug-and-Play Component for Neural Network Regularization

Dec 02, 2024
RJ
RuiZhe Jiang
🏛️ University of Chinese Academy of Science | Sichuan University

To address overfitting in over-parameterized neural networks, this paper proposes an adaptive regularization method grounded in feature consistency: it extracts features from multiple random subsets of the training set and explicitly enforces representational consistency across independent and identically distributed (i.i.d.) subsets by minimizing pairwise L2 or cosine distances among them. This is the first approach to model overfitting as *inconsistent representations on i.i.d. data subsets*, requiring no architectural or task-specific assumptions—making it plug-and-play. The method introduces zero gradient modifications and adds no extra learnable parameters. Experiments demonstrate that it significantly reduces the train–validation loss gap and improves generalization accuracy across diverse network architectures and tasks. Moreover, it exhibits strong hyperparameter robustness and negligible computational overhead.

Generalization PerformanceNeural NetworksOverfitting

This work addresses the inconsistency of existing attribution methods under geometric transformations and the limited fidelity of conventional gradient-based attributions in reflecting true model decision rationales. The authors propose an unsupervised attribution regularization framework that, for the first time, leverages submodular search to generate compact, class-discriminative, and faithful attribution supervision signals. They introduce path consistency and termination alignment losses to enable differentiable joint regularization of the discrete evidence selection process. Evaluated on ImageNet-100, the method substantially improves attribution stability and Insertion/Deletion metrics with only a 0.28% accuracy drop for ViT-B/16. On ImageNet-1K, it also enhances robustness to transformations while incurring no more than a 0.30% loss in clean accuracy.

attribution consistencyevidence reliancefaithful attribution

Neural network internal representations often lack stability and cross-architectural consistency due to architectural disparities, hindering knowledge transfer and modular deployment. To address this, we propose a structured regularization framework comprising linear shaping operators and rectified path constraints, which explicitly encode inductive biases to improve geometric alignment of representations across architectures. Through theoretical analysis, controlled transfer experiments, and a novel representation alignment metric, we systematically demonstrate that structural priors significantly enhance semantic consistency among heterogeneous models. Our method improves downstream task performance in model distillation and modular learning by up to 12.3%, offering an interpretable and scalable paradigm for building robust, composable deep learning systems.

Analyze impact of structural constraints on representation compatibilityImprove interoperability of learned features with inductive biasesStudy stability of learned representations across different architectures

Trading off Consistency and Dimensionality of Convex Surrogates for the Mode

Feb 16, 2024
EB
Enrique B. Nueve
🏛️ University of Colorado Boulder | Harvard University

To address the intractability of surrogate loss optimization in large-scale multiclass classification caused by high-dimensional embeddings, this paper proposes a low-dimensional convex polyhedral embedding framework, establishing theoretical trade-offs among embedding dimension, consistency regions, and data distribution assumptions. First, it rigorously proves that “hallucination”—i.e., spurious class predictions—necessarily occurs when the embedding dimension is less than $n-1$. Second, under a low-noise assumption, it derives a verifiable consistency criterion. Third, it designs structured embeddings—including hypercubes and permutahedra—that achieve dimensionality reductions from $2^d$ to $d$ and from $d!$ to $d$, respectively. Finally, in the multiple-instance learning setting, it shows that full simplex consistency is guaranteed with only $n/2$ embedding dimensions and proves the existence of consistent subsets around any point-mass distribution.

Avoiding hallucination when using low-dimensional convex polytope embeddingsReducing surrogate loss dimension for large multiclass classification problemsTrading off consistency, dimensionality, and number of problem instances

Latest Papers

What's happening recently
View more

This work addresses the alignment failures of large language models under emerging safety threats—including role-playing mimicry, adversarial exploits, prefilling attacks, and conditional misalignment—by introducing a multi-level consistency training mechanism within the Transformer architecture. Specifically, consistency constraints are applied at the MLP layers (MLPCT), attention heads (AttCT), and overall behavioral output (BCT), extending consistency-based alignment to these four complex threat scenarios for the first time. Experimental results demonstrate that the proposed approach substantially suppresses diverse misaligned behaviors and exhibits superior robustness and cross-threat generalization compared to existing methods tailored only to jailbreaking or sycophancy attacks. Furthermore, the study uncovers the critical role of shared residual streams in achieving effective model alignment.

alignmentconsistency trainingmodel misalignment

This work addresses the problem of efficiently merging multiple fine-tuned models into a unified multitask model without retraining. The authors formalize model merging as a convex quadratic program over residual updates, achieving theoretically optimal fusion by calibrating inputs and outputs to minimize calibration error in the output space. This study provides the first formal optimality guarantees for model merging, introduces an interpretable diagnostic metric based on residual energy, and unifies existing heuristic approaches within a single theoretical framework as special cases. Experimental results demonstrate that the proposed method matches or surpasses current techniques in single-layer settings and consistently improves performance across language and vision benchmarks in multilayer merging scenarios. Furthermore, the quality of merged models can be accurately predicted using a small calibration set.

fine-tuned modelsmodel mergingmulti-task learning

This study addresses the limitation in model merging where implicit regularization introduced during coefficient search constrains weights to a restricted subspace, thereby hindering multi-task performance improvements. To overcome this, we re-examine the implicit regularization mechanism in task arithmetic and propose directly searching the pre-trained weight space via unconstrained optimization, breaking through the subspace constraints inherent in traditional linear combinations. Our work reveals and eliminates these implicit regularization effects, demonstrating that superior solutions reside outside conventional subspaces. Across diverse architectures and extremely data-scarce scenarios, the proposed method significantly outperforms existing model merging techniques, establishing a new optimization paradigm for multi-task learning.

Implicit RegularizationModel MergingMulti-task Learning

Concept alignment lacks a unified definition, and existing methods optimize divergent objectives under the same terminology, obscuring its fundamental nature. This work formalizes its multidimensional structure by decomposing it along two axes—“alignment targets” and “alignment levels”—and identifies four distinct alignment properties, revealing that current approaches satisfy only subsets of these. To address this limitation, we propose Coupled Sparse Autoencoders (CoSAE), a framework that jointly optimizes multiple alignment objectives, alongside InterVenchA, an interventional evaluation benchmark. Experiments demonstrate that optimizing a single objective fails to reliably recover other alignment properties, whereas CoSAE achieves strong instance-level conceptual consistency using merely 0.1% paired data.

concept alignmentdistributional alignmentinstance-level alignment

Hot Scholars

WW

Weijie Wang

PhD Student, Zhejiang University
Computer VisionEfficient AIDeep Learning
ZT

Zhuotao Tian

Professor, Harbin Institute of Technology (Shenzhen)
Vision-language ModelMulti-modal PerceptionComputer Vision
GL

Guosheng Lin

Nanyang Technological University
Computer VisionMachine Learning
DC

Daniel Cremers

Technical University of Munich
Computer VisionMachine LearningOptimizationRobotics