Score
Design U‑Net–style encoder–decoder convolutional architectures that encode and fuse multiscale spatial features using downsampling and upsampling paths with skip connections to preserve fine detail. Implement output heads that produce corrective residual fields and choose normalization, conditioning, or iterative-update blocks to stabilize the network’s iterative update dynamics.
In medical image segmentation, U-Net’s skip connections suffer from insufficient cross-scale feature interaction and simplistic fusion strategies (e.g., concatenation or addition). To address these limitations, this work pioneers modeling skip connections as discrete nodes of an ordinary differential equation (ODE), and introduces an adaptive ODE-based fusion mechanism grounded in linear multistep methods. This enables continuous, learnable, and differentiable multi-scale feature integration along the decoding path. The proposed method is architecture-agnostic—decoupled from encoder-decoder design—and thus universally pluggable. Evaluated on five benchmark datasets (ACDC, KiTS2023, MSD Brain Tumor, ISIC), it consistently improves Dice scores by 1.2–2.8%, reduces parameter count by 12–18%, and significantly enhances feature utilization. These results empirically validate the effectiveness and generalizability of ODE modeling for optimizing skip connections in U-shaped architectures.
This work addresses the limitations of standard U-Net’s concatenation-based skip connections, which often induce channel dimension explosion and fail to suppress noise propagation, thereby compromising denoising performance and interpretability. To overcome these issues, the authors propose Additive U-Net, which replaces conventional concatenation with additive skip connections modulated by learnable non-negative gating scalars. This lightweight and interpretable mechanism explicitly controls the contribution of encoder features, effectively avoiding channel expansion while revealing a natural progression of multi-scale features from high- to low-frequency components. Evaluated on the Kodak-17 benchmark, the model achieves competitive PSNR and SSIM across noise levels (σ = 15, 25, 50) and demonstrates strong robustness to variations in network depth and convolutional kernel scheduling.
Conventional decoders for dense prediction tasks suffer from outdated architectural designs and insufficient cross-layer contextual sharing, limiting feature propagation efficiency and spatial consistency. Method: We propose a novel decoder architecture centered on a learnable, shared “bank”—a parameterized module dynamically resampled and fused across multiple scales to enable explicit cross-layer contextual reuse during decoding, thereby departing from traditional serial, layer-wise independent decoding paradigms. Built upon a Transformer backbone, the bank is jointly optimized end-to-end. Contribution/Results: Our approach significantly improves decoding efficiency and spatial coherence. On both natural and synthetic image depth estimation benchmarks, it substantially outperforms state-of-the-art methods, achieving superior accuracy and generalization under large-scale training. To our knowledge, this work presents the first systematic design and empirical validation of a universal, decoder-level contextual sharing mechanism.
Existing methods for unified image restoration under diverse degradations (e.g., noise, blur, adverse weather) rely on complex architectures—such as Mixture-of-Experts (MoE) or diffusion models—and require precise degradation priors, resulting in high computational cost and poor deployability. To address this, we propose SymUNet, a symmetric U-Net that explicitly models and propagates degradation information across scales via scale-aligned feature aggregation and cross-layer feature propagation, eliminating the need for expert modules or diffusion-based priors. Furthermore, we introduce SE-SymUNet, a CLIP-semantic-enhanced variant: its frozen CLIP backbone extracts degradation-relevant semantics, which are injected via lightweight cross-attention. Extensive experiments demonstrate that both SymUNet and SE-SymUNet surpass state-of-the-art methods across multiple benchmarks—achieving superior performance with significantly fewer parameters and faster inference speed. Code is publicly available.
This work addresses the limited reconstruction performance in U-Net decoders caused by imprecise fusion of high- and low-level features. To overcome this, the authors propose a novel difference-driven adaptive gating mechanism that, for the first time, leverages the discrepancy between high- and low-level feature streams—rather than their content or correlation—to generate coupled gating maps that precisely modulate the fusion process. Two specific implementations are introduced: Feature Difference-based Gating (FDG), which uses the absolute difference of features for local refinement, and Entropy Difference-based Gating (EDG), which exploits signed entropy differences to enable global optimization. Extensive experiments demonstrate that the proposed approach consistently outperforms existing attention-based fusion strategies across diverse tasks, including medical image segmentation, remote sensing cloud removal, and speech separation, with EDG achieving the best overall performance.
We investigate additive skip fusion in U-Net architectures for image denoising and denoising-centric multi-task learning (MTL). By replacing concatenative skips with gated additive fusion, the proposed Additive U-Net (AddUNet) constrains shortcut capacity while preserving fixed feature dimensionality across depth. This structural regularization induces controlled encoder-decoder information flow and stabilizes joint optimization. Across single-task denoising and joint denoising-classification settings, AddUNet achieves competitive reconstruction performance with improved training stability. In MTL, learned skip weights exhibit systematic task-aware redistribution: shallow skips favor reconstruction, while deeper features support discrimination. Notably, reconstruction remains robust even under limited classification capacity, indicating implicit task decoupling through additive fusion. These findings show that simple constraints on skip connections act as an effective architectural regularizer for stable and scalable multi-task learning without increasing model complexity.
Traditional neural networks struggle to model higher-order topological structures such as nodes, edges, faces, and hyperedges, often losing critical information when simplifying data into graphs or sequences. This work proposes a general U-Net architecture grounded in combinatorial complexes, replacing conventional spatial scales with “rank” as the hierarchical dimension. Cross-scale feature propagation is achieved through cells, incidence maps, and rank-wise pathways, while a bottleneck support ratio is introduced to quantify compression severity. The framework enables cohomological lifting and rank-matched skip connections across diverse topological domains, revealing the structural role of skip connections under high compression. Empirical results demonstrate that the model achieves state-of-the-art average accuracy on six out of eight node classification datasets and four out of five hypergraph benchmarks, with particularly pronounced gains on heterophilic graphs.
This study addresses the challenge that increasing resolution in imaging inverse problems often degrades the generalization performance of deep networks. The authors systematically investigate the generalization behavior of U-Net and its neural operator variants across varying discretization resolutions. Through interpretable one-dimensional models and two-dimensional limited-angle computed tomography reconstruction experiments, they find that although neural operator-based U-Nets are theoretically resolution-invariant, conventional U-Nets exhibit superior robustness and practical generalization. This work highlights a notable gap between theoretical resolution invariance and empirical performance, offering new insights for architecture selection in high-resolution inverse problem solving.
This work addresses the representational degradation and encoder–decoder imbalance in Diffusion Transformers (DiT) caused by fixed downsampling operations incompatible with the Transformer architecture. To resolve this, the authors propose the UDT framework, which introduces, for the first time in DiT, a data-adaptive token merging mechanism that replaces conventional downsampling, enabling efficient up- and down-sampling while preserving consistent token dimensions. UDT synergistically combines the multi-scale encoder–decoder strengths of U-Net with the powerful representational capacity of DiT, further enhanced by REPA regularization and a VA-VAE decoder. On ImageNet at 256×256 resolution, the XL variant achieves an FID of 7.9 after only 40 training epochs—matching the performance of SiT trained for 1,400 epochs, yielding a ~40× speedup. With classifier-free guidance and the VA-VAE decoder, UDT attains a state-of-the-art FID of 1.35 within 500 epochs.
U-Net’s static skip connections suffer from two key limitations: (i) inter-feature constraints—lacking content-aware cross-layer feature interaction—and (ii) intra-feature constraints—insufficient multi-scale feature aggregation. To address these, we propose the Dynamic Skip Connection Module (DSCM), which jointly incorporates test-time training (TTT) and dynamic multi-scale kernels (DMSK) to enable content-adaptive fusion of high- and low-level features and global-context-guided multi-scale interaction. DSCM is architecture-agnostic and seamlessly integrates into diverse U-shaped backbones—including CNNs, Transformers, and Mamba-based models. Evaluated across multiple medical image segmentation benchmarks, DSCM consistently improves segmentation accuracy while demonstrating strong generalizability and robustness to domain shifts and annotation noise.