Score
Design and implement losses, training procedures, and model components that enforce consistent, compositional latent decompositions across samples; common techniques include swapping latent factors between inputs and applying a cycle (swap-back) consistency loss (latent swapping cycle / LCC) so reconstructed outputs reflect the intended composition. Build and evaluate regularizers and metrics that promote disentanglement of compositional latent factors and stable recombination when components are exchanged.
To address training instability and degraded generation quality in latent consistency models (LCMs) caused by outlier interference, this paper proposes a robust training framework enabling high-fidelity single-step and two-step text-to-image and text-to-video generation. Our key contributions are: (1) the first adoption of Cauchy loss—replacing pseudo-Huber loss—to enhance robustness against outliers in latent-space optimization; (2) coupling early-diffusion loss with an optimal transport (OT) objective to improve alignment between predicted and target latent distributions; and (3) introducing an adaptive scaling-c scheduler and non-scaled LayerNorm to stabilize training dynamics. Experiments demonstrate that our method significantly narrows the performance gap between LCMs and diffusion models, achieving state-of-the-art fidelity and sampling efficiency across multi-scale text-conditioned generation tasks.
Diffusion models often face performance trade-offs in multi-objective alignment (e.g., multi-reward optimization) and multi-model composition, struggling to simultaneously satisfy all constraints while preserving consistency with pre-trained models. This paper proposes the first constraint-optimization framework tailored for diffusion models, unifying alignment and composition tasks via a bi-objective formulation: explicit reward constraints and model proximity constraints—marking the first integration of constrained learning into diffusion model fine-tuning. Theoretically, we characterize solutions provably satisfying all constraints; algorithmically, we design an efficient Lagrangian dual-based solver supporting both multi-reward alignment and multi-model fusion. Extensive image generation experiments demonstrate that our method significantly outperforms equal-weighted ensemble baselines, achieving superior generation quality and controllability while strictly adhering to user-specified constraints.
Intermediate checkpoints in diffusion models (DMs) and consistency models (CMs) are often underutilized, despite evidence that optimal weights frequently reside in non-convex “basins” where SGD fails to converge. Method: We propose LCSC—a learning-based checkpoint selection and combination framework—that employs evolutionary search to automatically learn linear weighting coefficients over trajectory checkpoints, integrates multi-stage weights, and synergistically combines consistency distillation with diffusion sampling optimization. Contribution/Results: LCSC establishes a generalizable checkpoint-weighted averaging paradigm that improves both generation quality and inference efficiency without increasing computational cost at deployment. On CIFAR-10 and ImageNet-64, LCSC achieves up to 23× training speedup; reduces DM sampling NFE from 15 to 9; and enables CM single-step inference to outperform the two-step baseline—demonstrating for the first time that trajectory-weighted averaging can transcend SGD’s convergence limitations, thereby introducing a novel training paradigm for generative models.
In non-injective regression, multi-output models heavily rely on pre-specified probability distributions and manually engineered prior knowledge. To address this, we propose a data-driven cycle-consistency framework that jointly optimizes a forward model Φ: X→Y and a backward model Ψ: Y→X, incorporating a cycle-consistency loss L_cycle = ℓ(Y, Φ(Ψ(Y))) to establish a generation–verification closed loop—without assuming output distributions or designing explicit rules. Our key innovation lies in dynamically compressing the solution space to enable unsupervised learning, thereby substantially reducing human intervention. Evaluated on synthetic and simulated datasets, the method achieves cycle reconstruction errors below 0.003 and improves key evaluation metrics by approximately 30% over baselines. It significantly enhances model generalizability, adaptability, and capability in modeling non-injective mappings.
This work proposes the Dual-Ended Consistency Model (DE-CM) to address the training instability and inflexible sampling of consistency models in large-scale applications. By introducing a trajectory selection mechanism, the method designs a three-segment key sub-trajectory optimization strategy and integrates continuous-time consistency objectives with flow-matching boundary regularization to enable few-step distillation. Additionally, a noise-to-noise (N2N) mapping is introduced to mitigate error accumulation at the initial step. The proposed approach significantly enhances both training stability and sampling flexibility, achieving a state-of-the-art one-step generation FID of 1.70 on ImageNet 256×256, which represents the best reported result for one-step consistency models to date.
This study addresses the unclear underlying mechanisms behind the delayed emergence of validation generalization in grokking, highlighting limitations in existing explanations. By constructing an analytical framework grounded in mode connectivity and the geometry of low-loss regions, this work reveals that the misalignment of low-loss regions induced by training-validation partitions is a critical cause of grokking. Furthermore, it proposes a symmetry-preserving data partitioning strategy to generate stable anti-grokking cases. These findings challenge prevailing correlation-based explanatory theories and demonstrate that when low-loss regions are aligned, hyperparameter tuning alone cannot induce grokking; instead, model dynamics collapse into either trainable or untrainable states.
This work addresses a critical limitation in existing factorized generative models, which only match the marginal distribution of style latent variables without enforcing independence from class information, leading to conditional style leakage. The study demonstrates for the first time that marginal distribution matching alone is insufficient for effective disentanglement and establishes that four theoretical conditions must be jointly satisfied. To systematically quantify style-class leakage, the authors introduce a comprehensive auditing framework combining maximum mean discrepancy (MMD), linear probing, clustering evaluation, and multidimensional perturbation experiments. Empirical results reveal that multiple baseline models—despite achieving near-zero marginal MMD—still enable label recovery with 74%–100% accuracy. The proposed post-processing method substantially improves generation quality, attaining external evaluation scores of 0.97 on MNIST and 0.88 on CIFAR-10.
This work addresses the feature splitting and feature absorption phenomena commonly observed in sparse autoencoders when trained with large-scale dictionaries, which undermine the atomicity and interpretability of latent representations. To mitigate these issues, the authors propose Cross-sample Consistency Regularization (C²R), a novel regularization technique that encourages semantically similar samples within a batch to activate consistent latent units while suppressing the co-activation of latent variables with similar directions. This approach systematically enhances the discreteness and semantic clarity of learned features without compromising reconstruction fidelity. Experimental results demonstrate that C²R significantly improves the atomicity and interpretability of latent representations, offering a promising direction for advancing sparse representation learning.
High-dimensional compositional data often exhibit both latent heterogeneous subpopulations and sparse effect structures, yet existing methods struggle to simultaneously perform clustering and within-cluster dimension reduction. This work proposes a Bayesian heterogeneous relative shift regression model that precisely ties cluster-specific coefficients through a projection-shrinkage prior defined on an identifiable contrast space, while incorporating a finite mixture prior to automatically infer the number of clusters. We develop a hybrid MCMC algorithm combining deterministic collapsing operators with the No-U-Turn Sampler (NUTS) for efficient posterior sampling. Theoretical results establish posterior consistency for both the latent partition and cluster-specific effect structures. Comprehensive simulations and real-data analyses demonstrate the method’s superior performance in estimation accuracy, predictive capability, and interpretability.
Existing flow matching approaches struggle to jointly model forward generation and reverse classification of multivariate data, lacking consistency in conditional inference. This work proposes a Joint Flow Matching (JFM) framework that assigns symmetric roles to variables at temporal endpoints, thereby constructing a shared joint distribution such that forward and backward integrations naturally correspond to conditional forms of the same joint distribution. JFM is the first method to enable consistent bidirectional conditional inference within continuous normalizing flows, inherently supporting confidence calibration without post-processing and providing an interpretable foundation for discriminative–generative tasks. Experiments demonstrate that JFM achieves competitive classification accuracy on conditional datasets, generates samples highly consistent with the classifier, and yields natively calibrated confidence scores.