Score
Design and implement convolutional encoder–decoder neural networks whose encoder and decoder mirror each other’s layer structure and parameterization, and that include constraints or losses to enforce representation consistency across inputs and geometric properties of convolutional filters. Use these symmetric convolutional autoencoders to build and analyze stable low-dimensional latent representations and to evaluate reconstruction quality and downstream predictive performance for tasks like dimensionality reduction, compression, or surrogate modeling.
Deep autoencoders lack rigorous theoretical foundations for their representational capacity. Method: This paper focuses on symmetric autoencoders—a canonical architecture—and establishes, for the first time, a rigorous mathematical connection between their reconstruction error and the Eckart–Young–Schmidt (EYS) theorem, revealing them as nonlinear generalizations of optimal low-rank approximation. We propose an EYS initialization strategy based on iterative singular value decomposition (SVD) to bridge classical linear approximation and deep nonlinear modeling. Orthogonal regularization, symmetric architectural design, and theoretical error analysis jointly characterize the expressivity bounds of diverse symmetric structures. Results: Extensive benchmark experiments demonstrate that EYS initialization significantly accelerates convergence and improves reconstruction accuracy, empirically validating the effectiveness and practicality of theory-driven model design.
This work addresses the joint problem of intrinsic dimension estimation and geometry-invariant embedding learning for nonlinear manifold-structured data. We propose an autoencoder framework incorporating orthogonality constraints on hidden-layer gradients. Methodologically, we establish, for the first time, a theoretical connection between gradient orthogonality in neural network latent spaces and the local tangent space dimension of the underlying manifold; this enables simultaneous intrinsic dimension estimation, learning of invertible embedding mappings, and construction of coordinate-invariant representations under local Lie group actions on low-dimensional submanifolds. Our key contribution lies in unifying gradient orthogonality with differential-geometric structure, thereby extending invariant representation learning to continuous group actions. Experiments on standard benchmarks demonstrate accurate intrinsic dimension estimation, disentangled representations, and robust group-invariant embeddings, validating both theoretical soundness and algorithmic robustness.
A key biological implausibility of backpropagation lies in its requirement for precise weight symmetry between forward and backward pathways. To address this, we propose Product Feedback Alignment (PFA), a novel learning algorithm that replaces the fixed random feedback matrix in classical Feedback Alignment with a dynamic, multiplicative feedback mechanism. PFA is the first method—both theoretically and empirically—to achieve high-fidelity approximation of standard backpropagation gradients while fully eliminating the weight symmetry constraint. The algorithm is natively compatible with convolutional neural networks (CNNs) and mainstream optimizers. On standard image recognition benchmarks—including ImageNet—it matches backpropagation’s accuracy, substantially outperforms classical Feedback Alignment, and exhibits superior convergence stability and generalization performance in deep architectures. By reconciling gradient-based learning with biologically plausible circuitry, PFA advances the frontier of biologically interpretable deep learning.
Approximate equivariant neural networks suffer from excessive parameter counts and limited flexibility in modeling symmetries. Method: This paper introduces a structured parameterization framework based on group matrices (GMs), unifying low-displacement-rank (LDR) structures for arbitrary finite groups. By integrating group representation theory with structured matrix modeling, the approach naturally encodes approximate equivariance and generalizes fundamental CNN operations beyond cyclic groups. Contribution/Results: The method preserves approximate equivariance while drastically improving parameter efficiency—reducing parameter counts by one to two orders of magnitude compared to state-of-the-art approximate equivariant networks and structured models. It achieves competitive performance across multiple tasks, offering both theoretical unification—bridging symmetry-aware learning and structured linear algebra—and practical computational benefits.
To address the uncontrollable topology of latent spaces (LS) in autoencoders (AEs), this paper proposes a geometric-loss-guided supervised AE co-optimization framework—the first to explicitly configure LS topology in supervised AEs. The method jointly optimizes encoder architecture and geometric constraint losses, enabling user-defined cluster positions and shapes, decoder-free label prediction, and cross-sample similarity assessment. Key innovations include zero-shot cross-dataset generalization, similarity estimation for unseen classes, and text-driven image retrieval without classifiers or language models. Experiments demonstrate 12–19% improvements in zero-shot accuracy on LIP, Market-1501, and WildTrack, and achieve 78.3% mAP in cross-modal retrieval.
This work addresses the limitations of conventional convolutional autoencoders in reduced-order modeling, which often fail to preserve the essential geometric structure of solution manifolds governed by parametric partial differential equations, leading to unstable latent trajectories and high reconstruction errors. To overcome this, the authors propose a symmetric convolutional autoencoder that, for the first time, incorporates the principle of representation consistency into convolutional layers by embedding differential-geometric priors. This design explicitly preserves the symmetry and intrinsic geometric properties of the solution manifold. Evaluated on three one-dimensional benchmark problems, the proposed method significantly outperforms standard convolutional autoencoders, achieving lower reconstruction errors, improved accuracy of latent trajectories, and enhanced model robustness.
This work addresses the growing complexity and lack of interpretability in deep image compression autoencoder models, which hinder the design of efficient architectures. For the first time, it systematically employs Jacobian analysis to examine the internal transformations of unbiased autoencoders, uncovering consistent and interpretable operational patterns that are prevalent across high-dimensional compression models. Building on these insights, the study identifies multiple semantically meaningful internal operations shared across diverse models and demonstrates their utility in constructing lightweight architectures that simultaneously achieve high compression performance and low computational complexity. This approach establishes a new paradigm for designing interpretable and efficient compression models grounded in analytically derived internal mechanisms.
This work addresses the geometric mismatch arising from mainstream optimizers like Adam, which disregard the symmetry and equivariance structures inherent in neural network parameter spaces. The authors propose a principled framework for designing symmetry-compatible optimizers, tailoring gradient update rules to respect the equivariance of specific architectural components—such as embedding layers, language model heads, SwiGLU projections, and MoE routers—and assembling them into an end-to-end hierarchical optimizer stack. For the first time, this approach is systematically applied beyond generic matrix layers, encompassing permutation and shared translational symmetries, thereby unifying and extending equivariant optimization methods. Efficient compatibility is achieved through techniques including one-sided spectral updates, row/column-aware normalization, and centering. In pretraining both dense and sparse MoE language models, the proposed optimizer consistently outperforms AdamW, yielding lower validation loss and enhanced training stability.
This work addresses the high computational and memory costs of group convolutional neural networks when processing 3D geometric data, which arise from dense sampling over transformation groups. To mitigate this, the authors propose a feature-space sparse sampling strategy that selects representative samples based on feature similarity, replacing conventional dense geometric sampling. This approach strictly preserves equivariance while substantially reducing computational overhead. By decoupling geometric resolution from computational cost, the method enables flexible trade-offs between accuracy and efficiency and further accelerates training through precomputed geometric similarities. Experimental results demonstrate that even with coarse-grained sampling, the model maintains high classification accuracy and significantly speeds up the training of 3D equivariant classifiers.