Score
Designs and evaluates algorithms and models that represent, predict, and manipulate residual (difference) signals in volumetric data, including residual network and stream architectures, residual regression, resampling, and representation learning. Builds memory‑efficient diffusion/transformer variants and statistical residual diagnostics to analyze model residuals and to preserve or refine high‑frequency voxel/detail information.
To address the fundamental limitation in Transformers—where residual flow capacity is tightly constrained by computational cost and model parameters, hindering efficient scaling—this paper proposes the Residual Matrix Transformer (RMT). RMT replaces conventional scalar/vector residual connections with independently scalable outer-product memory matrices, thereby redefining information storage and retrieval. It further introduces variance-aware propagation rules and a theory-guided training dynamic optimization strategy. Crucially, RMT achieves the first decoupling of residual capacity from both FLOPs and parameter count. Experiments demonstrate that, at equivalent loss, RMT reduces FLOPs by 58%, parameters by 25%, and training tokens by 41%. Across diverse downstream tasks, RMT consistently outperforms standard Transformers, delivering substantial gains in both training efficiency and generalization performance.
To address resampling artifacts and information loss arising from highly heterogeneous voxel sizes in multicenter MRI, this work proposes a resampling-free framework for multiple sclerosis lesion segmentation. The core innovation lies in a physically grounded, radius-fixed spherical harmonic parameterization of convolutional kernels, integrated with E(3)-equivariant convolutions and spherical harmonic expansions—enabling voxel-size-invariant and resolution-adaptive feature learning. By preserving the original spatial geometry and physical constraints, the method directly models local structural patterns under variable voxel resolutions. Evaluated on multiple publicly available and internal multicenter MS datasets exhibiting high inter-scanner heterogeneity, our approach significantly outperforms both resampling-based and resampling-free U-Net baselines. It achieves consistent improvements across 2D and most 3D evaluation metrics and demonstrates superior generalizability.
This work addresses gradient explosion during backpropagation in residual network training and the high memory consumption and communication overhead inherent in distributed training. We propose a serial and parallel optimization framework based on proximal (linearized) ADMM. Theoretically, we establish, for the first time without assumptions on network width, depth, or dataset size, the R-linear convergence of the algorithm. Practically, we design a low-communication, low-memory distributed coordination protocol enabling localized parameter updates. Experiments demonstrate rapid and stable convergence, improved test accuracy, and significantly reduced per-node memory usage; the parallel variant substantially enhances scalability and training efficiency. Our core contribution lies in systematically integrating ADMM into residual network training—achieving both rigorous theoretical guarantees and practical engineering viability.
Why do residual architectures (e.g., ResNet, Transformer) consistently improve performance with increased depth? This paper addresses this fundamental question from a functional perspective. We propose and rigorously prove the *Residual Expansion Theorem*, establishing that depth growth is equivalent to an exponential expansion of implicit ensemble capacity: each added layer introduces new computational paths, inducing combinatorial path explosion and yielding a hierarchical ensemble mechanism. This mechanism critically relies on normalization layers to suppress signal explosion, while depth itself implicitly imposes regularization that governs model complexity. Based on this insight, we provide the first theoretical foundation for normalization-free residual architectures and derive the *module scaling principle*—a theoretically grounded strategy for stabilizing deep-network training. Our approach integrates analytical modeling, combinatorial mathematics, and function-space analysis to unify the interplay among depth, ensembling, and regularization.
The scaling factor in residual connections of ResNets critically influences generalization, yet its mechanistic role and robustness across hyperparameter configurations remain poorly understood. Method: We establish the first finite-width field-theoretic framework for ResNets and analytically derive the input response function to characterize signal propagation. Contribution/Results: Our theory reveals that the empirically optimal scaling interval corresponds to the regime of maximal input sensitivity; moreover, the optimal scaling value depends only weakly on network depth and weight variance—explaining its empirical stability across diverse hyperparameter settings. This work provides the first analytical solution for the residual scaling factor and yields interpretable, theoretically grounded guidelines for its selection, thereby bridging empirical practice with rigorous understanding of signal propagation in deep residual networks.
This work addresses a key limitation in conventional heterogeneous representation learning, where feature concatenation obscures the origin of prediction errors and impedes attribution of individual representations to model performance. To overcome this, the authors propose a novel paradigm that models representations as typed objects endowed with coordinate systems and unresolved residuals, preserving their identities through ordered operator composition. The framework explicitly captures residual attribution and marginal gains via a reflective introspection operator and an orthogonal projection mechanism, thereby eliminating the need for grid search. Building upon Fold, it constructs a conditional mean field and integrates FPRC-PQ to realize a relaxation-aggregation-closure algebraic pipeline, enhanced by controlled-variable interfaces and first-order coupling path orthogonality optimization. Evaluated on 3.67 million daily A-share records, the method boosts net cost-adjusted returns from 13.52% to 19.10% and increases the Sharpe ratio from 1.42 to 2.09, substantially outperforming baseline models.
This work addresses key challenges in real-world image restoration with diffusion models—namely slow inference, insufficient fidelity, and ineffective utilization of pretrained generative priors. The authors propose a scalable restoration framework built upon a pretrained text-to-image Rectified Flow model, introducing a residual vector field to construct a residual Rectified Flow that enables efficient transport starting from degraded images rather than pure noise. This approach preserves consistency with the original pretrained model while allowing parameter-efficient fine-tuning. A knowledge distillation strategy is further integrated to reduce sampling costs. Extensive experiments demonstrate state-of-the-art performance across multiple real image restoration tasks, significantly accelerating inference and validating the practicality and scalability of large-scale pretrained diffusion models for restoration applications.
This work investigates whether privileged directions with semantic or functional significance exist within the residual stream of Transformers. Introducing prediction directions as novel geometric anchors, we systematically analyze residual stream structure across diverse Transformer scales and architectures through geometric analysis, directional perturbations, variance decomposition, and cross-model validation. Our findings reveal a hierarchical organization of the residual stream based on proximity to the prediction direction: regions near this direction exhibit strong structural coherence and dominate task-critical decisions, while distal regions—though weakly aligned—provide essential causal and temporal support. This consistent geometric–functional correspondence is observed across 18 distinct models, underscoring a fundamental principle underlying residual stream organization in Transformers.
This work addresses the challenge that existing neural network width theories fail to guarantee the generalization of widening directions identified during training. Focusing on function-preserving residual expansions, the study investigates the alignment between training and test gradients, introducing the notion of “effective alignment dimension” to characterize the signal-to-noise geometric structure of activation gradients. For the first time, this measurable quantity enables a high-probability guarantee of improved test risk under finite samples—without requiring assumptions on covariance spectra or predetermined width growth rates. The theoretical analysis leverages mean-variance decomposition of inner products of activation gradients within a residual expansion framework. Experiments on LLaMA-style Transformers, Pythia, and ResNet-20 demonstrate that wider models exhibit higher effective alignment dimensions and lower empirical misalignment, with this metric accurately predicting both the direction and magnitude of held-out loss changes.