Score
Designs and implements training objectives, loss terms, and model components that enforce that a forward mapping between two representation spaces composed with a reverse mapping reconstructs the original inputs (cycle-consistency). Builds and evaluates generators/encoders/decoders or paired mappings regularized by cycle-consistency losses or constraints, and analyzes reconstruction error, consistency trade-offs, and optimization behavior when jointly training the paired models.
In non-injective regression, multi-output models heavily rely on pre-specified probability distributions and manually engineered prior knowledge. To address this, we propose a data-driven cycle-consistency framework that jointly optimizes a forward model Φ: X→Y and a backward model Ψ: Y→X, incorporating a cycle-consistency loss L_cycle = ℓ(Y, Φ(Ψ(Y))) to establish a generation–verification closed loop—without assuming output distributions or designing explicit rules. Our key innovation lies in dynamically compressing the solution space to enable unsupervised learning, thereby substantially reducing human intervention. Evaluated on synthetic and simulated datasets, the method achieves cycle reconstruction errors below 0.003 and improves key evaluation metrics by approximately 30% over baselines. It significantly enhances model generalizability, adaptability, and capability in modeling non-injective mappings.
This study addresses the training instability, texture drift, and artifact generation inherent in CycleGAN by proposing an enhanced framework integrating Wasserstein loss with gradient penalty, perceptual loss, multi-scale discriminators, and self-attention mechanisms. Innovatively, this work categorizes these components based on computational overhead during training versus inference to establish a resource-aware deployment prioritization strategy, specifically recommending deferred adoption of self-attention. Evaluated on horse-to-zebra translation, the proposed model significantly improves FID and KID metrics while effectively mitigating training collapse and reconstruction artifacts. Consequently, this research provides both theoretical foundations and practical guidelines for optimizing generative models in resource-constrained scenarios, balancing performance gains with computational efficiency through strategic component scheduling.
This work proposes an unpaired image translation method based on CycleGAN to address the incompatibility of fluorescence microscopy images with standard histopathological workflows that rely on hematoxylin and eosin (H&E) staining. By fusing two fluorescence channels (C01 and C02) into an RGB input, the model translates multi-channel fluorescence images into virtual H&E-like images exhibiting realistic color characteristics. The architecture employs a ResNet-based generator and a PatchGAN discriminator, trained with a combination of adversarial loss, cycle-consistency loss, and identity loss. The generated images preserve the original tissue morphology while accurately mimicking the chromatic appearance of conventional H&E-stained slides, thereby significantly enhancing compatibility with existing pathological analysis pipelines and facilitating multimodal data integration for clinical applications.
CycleGAN suffers from poor generalization in unpaired image translation, yet its generalization risk remains theoretically uncharacterized. Method: This work systematically decomposes and quantifies the total generalization risk into approximation error and estimation error. Using optimal transport theory, we derive a tight upper bound on the approximation error; leveraging Rademacher complexity analysis, we bound the estimation error. We further reveal that while cycle-consistency constraints improve structural plausibility, they exacerbate the tension between model capacity and finite-sample size, thereby degrading generalization. Contribution/Results: We propose an interpretable risk decomposition framework that explicitly links architectural choices—such as generator capacity and cycle-loss weighting—to generalization performance. This framework establishes a novel theoretical paradigm for analyzing unpaired generative models and advances the interpretability of adversarial training by grounding it in rigorous statistical learning principles.
This work investigates the intrinsic relationship between deep model merging and neural network interpretability, addressing the lack of a unified geometric foundation for understanding their interplay. Method: We systematically analyze empirical phenomena in model merging and identify four geometric properties of the loss landscape—mode convexity, determinacy, directionality, and connectivity—as shared structural determinants of merging efficacy and generalization robustness. We develop a unified framework integrating loss landscape analysis, neural representation modeling, and interpretability evaluation. Contribution/Results: Theoretically, we establish how landscape geometry governs the interpretability of internal representations and their stability across tasks. Our framework provides the first geometric interpretation paradigm for model merging and pioneers a novel interdisciplinary research direction at the intersection of “landscape geometry–interpretability–robustness,” offering principled insights into both model composition and trustworthy AI.
本文提出了一种双向神经框架CyclOT,通过同步前向后向插值来学习高维未配对样本的二次最优传输映射,结合了双向二次动作、判别器限制的端点目标和双侧循环一致性。
Existing video frame interpolation methods often suffer from motion drift, directional ambiguity, and boundary misalignment due to unidirectional generation, and they lack temporal consistency over long sequences. This work proposes a bidirectionally cycle-consistent video diffusion interpolation framework that employs learnable directional tokens to guide a shared backbone network, jointly optimizing forward synthesis and backward reconstruction within a unified architecture to achieve logically invertible motion trajectories. During training, bidirectional cycle consistency is enforced as a regularizer, complemented by a curriculum learning strategy that progressively optimizes from short to long sequences. At inference, the model requires only a single forward pass. The proposed method significantly outperforms strong baselines on 37- and 73-frame interpolation tasks, achieving state-of-the-art performance in image quality, motion smoothness, and dynamic control without incurring additional computational overhead.
This work addresses the challenges of representation drift and catastrophic forgetting in exemplar-free class-incremental learning, where historical data cannot be stored. To this end, the authors propose BiCyc, a novel method that introduces, for the first time, a bidirectional projection alignment mechanism with stop-gradient gating and a cycle-consistency loss. These components jointly optimize the mapping between old and new feature spaces, enabling co-evolution of transfer and representation. Theoretical analysis demonstrates that this design contracts the singular spectrum in the whitened space and reduces perturbations in classification logit outputs. Empirical results show that BiCyc significantly lowers forgetting rates on standard EFCIL benchmarks and achieves superior performance under both from-scratch and pretrained fine-grained settings.
This work addresses the challenge of achieving strict idempotence in generative models under repeated application, where output drift arises due to geometric inconsistencies between the data manifolds learned by the encoder and decoder. The study identifies this manifold misalignment as the key cause of idempotence failure—a previously unexamined issue—and introduces a novel training framework that explicitly aligns the geometric structures of both components. By enforcing the encoder’s projection and the decoder’s reconstruction to share a common underlying manifold during training, the proposed method substantially reduces idempotence error, yielding perfectly consistent outputs across repeated generations. Empirical results demonstrate significant improvements in identity preservation and information stability for image generation and editing tasks.
Existing flow matching approaches struggle to jointly model forward generation and reverse classification of multivariate data, lacking consistency in conditional inference. This work proposes a Joint Flow Matching (JFM) framework that assigns symmetric roles to variables at temporal endpoints, thereby constructing a shared joint distribution such that forward and backward integrations naturally correspond to conditional forms of the same joint distribution. JFM is the first method to enable consistent bidirectional conditional inference within continuous normalizing flows, inherently supporting confidence calibration without post-processing and providing an interpretable foundation for discriminative–generative tasks. Experiments demonstrate that JFM achieves competitive classification accuracy on conditional datasets, generates samples highly consistent with the classifier, and yields natively calibrated confidence scores.