Score
Designs and implements encoders and embedding pipelines that represent the residual signal (the difference between a base prediction and the desired target) and builds conditioning mechanisms that feed those residual embeddings into reconstruction or generative processes. Analyzes and optimizes how residual conditioning reduces inversion noise, preserves global consistency and identity, and enables precise localized edits and relighting by improving fidelity of the final output.
This work addresses the challenge of simultaneously preserving identity consistency and structural stability in generative image editing when no paired fine-tuning data is available. To this end, the authors propose a decoupled editing framework based on residual image encoding and gradient reversal optimization. By incorporating residual embeddings as an additional conditioning signal in the diffusion model, the method enhances reconstruction fidelity and reduces reliance on weakly conditioned inversion. A gradient reversal layer is further employed to effectively disentangle the edited attributes from the original content. The approach demonstrates significant improvements in editing fidelity and consistency across tasks such as relighting and text-guided manipulation, surpassing the performance limitations of existing training-free image editing methods.
To address the challenges of few-shot learning, sparse labeling, heterogeneous (numerical and categorical) conditioning variables, and constrained computational resources in engineering applications, this paper proposes a masked conditional generative paradigm. We design a unified learnable embedding to jointly model heterogeneous conditions and introduce a masked conditional scheduling mechanism that explicitly simulates missing conditions during training to enhance robustness to incomplete inputs. Furthermore, we construct a lightweight collaborative architecture integrating a variational autoencoder and a latent diffusion model, coupled with knowledge distillation from pre-trained large models. Experiments on 2D point cloud and engineering image datasets demonstrate that the method enables efficient training with only a small number of labeled samples; achieves a 32% reduction in Fréchet Inception Distance (FID); significantly improves conditional fidelity; and simultaneously ensures strong controllability and high generation quality.
Why do residual architectures (e.g., ResNet, Transformer) consistently improve performance with increased depth? This paper addresses this fundamental question from a functional perspective. We propose and rigorously prove the *Residual Expansion Theorem*, establishing that depth growth is equivalent to an exponential expansion of implicit ensemble capacity: each added layer introduces new computational paths, inducing combinatorial path explosion and yielding a hierarchical ensemble mechanism. This mechanism critically relies on normalization layers to suppress signal explosion, while depth itself implicitly imposes regularization that governs model complexity. Based on this insight, we provide the first theoretical foundation for normalization-free residual architectures and derive the *module scaling principle*—a theoretically grounded strategy for stabilizing deep-network training. Our approach integrates analytical modeling, combinatorial mathematics, and function-space analysis to unify the interplay among depth, ensembling, and regularization.
This work addresses the limitation of conventional residual connections, which sum sublayer updates with fixed coefficients and cannot dynamically assess the reliability of proposed updates. To overcome this, the authors propose Review Residuals—a novel mechanism that explicitly incorporates conditional dependence on the proposed update within the residual gating function. By employing a learnable sigmoid gate conditioned on two inputs via RMSNorm, the method dynamically scales the residual term while preserving the identity additive structure, thereby balancing training stability and representational capacity. The approach integrates seamlessly into standard Transformers and demonstrates statistically significant improvements (p<0.05) over both standard residual connections and Highway gating in models of 590M parameters and larger, with performance gains increasing with model scale. It also enables stable training of extremely deep networks.
This study addresses the inefficiency of existing learning-based methods in high-fidelity scientific data compression (block-level NRMSE of 10⁻⁶–10⁻⁴), which stems from excessive bitrate consumption by global residual correction. The work introduces, for the first time, a residual-centric modeling framework that treats residuals as an independent signal type and proposes two specialized coding schemes: a training-free deterministic LBRC encoder and an NGLR encoder incorporating causal neural prediction. The approach integrates 3D Lorenzo differencing, Zigzag mapping, bitplane decomposition, and entropy coding, augmented with adaptive quantization and integer-based residual processing. Evaluated on E3SM, JHTDB, and ERA5 datasets, LBRC achieves 30–60% higher compression ratios than GAE, while NGLR further improves this by 10–40%, outperforming the SZ compressor in high-fidelity regimes.
This work addresses the challenges in diffusion-based image editing, particularly the insufficient inversion accuracy and the trade-off between editing fidelity and background preservation. The authors propose SimEdit, a framework that investigates how textual conditions influence the geometry of the diffusion velocity field and cross-branch attention consistency, thereby revealing the critical role of condition precision in inversion stability. Building on this insight, they design a condition-aware dual-component editing mechanism that integrates condition refinement with token-wise cross-branch attention control. This approach substantially improves both image reconstruction quality after inversion and overall editing performance, outperforming existing attention-manipulation methods on the PIE-Bench benchmark.
This work investigates whether Rectified Flows implicitly leak private information from their training data. By analyzing the interpolation paths they rely on, defined as $X_\lambda = (1-\lambda)X_0 + \lambda X_1$, the study reveals for the first time that reconstruction error exhibits a bell-shaped distribution along $\lambda$. Under Gaussian assumptions, the authors derive a closed-form expression for the location of this peak. Leveraging this $\lambda$-resolved signal, they propose a novel membership inference attack. Experiments on image and audio datasets demonstrate the universality of the bell-shaped error structure and the accuracy of the predicted peak location. The resulting attack significantly outperforms existing baselines, offering a new perspective on privacy risks in generative modeling.
This work addresses the inherent ambiguity in conventional feature inversion methods, which often fail to establish a unique correspondence between inverted outputs and original inputs due to a lack of sample specificity. To overcome this limitation, the authors propose a source-anchored feature inversion approach that explicitly links features to the local network geometry of their source inputs, eliminating reliance on generic image priors. By employing a closed-form Wiener filtering mapping to reconstruct the accompanying signal and integrating a Jacobian-vector product (JVP)-based forward consistency residual, the method achieves precise inversion within a single backward pass. The framework enables zero-intercept mappings across architectures, depths, and channels, successfully generating images that simultaneously align with both the original input and target features in both CNNs and Transformers. Validation via predicted conditional feature maps demonstrates its efficacy in revealing the true influence of internal representations on model decisions.
Existing generative models struggle to synthesize physically consistent camera raw images, hindering progress in low-level vision tasks. This work proposes RawGen, the first diffusion framework capable of both text-to-raw image generation and sRGB-to-raw inverse mapping. RawGen leverages the generative priors of large-scale sRGB diffusion models and integrates multi-parameter ISP simulation, a conditional denoiser, and a dedicated decoder to jointly produce physically plausible linear raw images in both latent and pixel spaces. To overcome the limitation of fixed ISP assumptions, the authors construct a many-to-one inverse ISP dataset. Experiments demonstrate that RawGen significantly outperforms existing methods in raw reconstruction quality, and its synthetic data effectively enhances performance on downstream vision tasks.