Score
Designs and implements coordinate-based implicit neural fields that map continuous input coordinates to target features, building architectures, activation schemes (including periodic activations) and latent encodings to represent shapes, signals or temporal fields as compact neural functions. Builds and evaluates training and fitting procedures and regularization to handle undersampled measurements, reduce optimization instabilities, encode exemplar geometries as variables, and analyze the properties and comparability of the resulting implicit neural representations.
This paper systematically surveys implicit neural representations (INRs), identifying key performance bottlenecks—particularly limited expressivity and scalability—in inverse problems such as audio synthesis, image reconstruction, 3D scene modeling, and high-dimensional data generation. To address these challenges, we propose the first unified four-dimensional taxonomy—spanning activation functions, positional encodings, joint encoding strategies, and network architectures—and uncover a fundamental trade-off between local bias suppression and fine-grained detail modeling. Through cross-modal benchmarking and gradient-guided structural optimization, we quantitatively evaluate reconstruction fidelity, memory efficiency, and generalization capability. Our contributions include: (i) an open, reproducible experimental benchmark; (ii) identification of three critical research frontiers—activation expressivity, positional encoding robustness, and high-dimensional scalability; and (iii) theoretically grounded, practice-oriented guidance for INR method selection and future development.
Traditional discrete representations suffer from resolution dependency, modality coupling, and poor generalization in data reconstruction. To address these limitations, this paper establishes a unified framework for continuous representation (CR), which maps spatial coordinates to continuous functions—enabling resolution-agnostic modeling for tasks such as image reconstruction and novel-view synthesis. Methodologically, we systematically formalize the CR paradigm along three dimensions: algorithmic design, theoretical foundations, and cross-domain applications—constituting the first comprehensive taxonomy. We identify and characterize three core properties: implicit regularization, cross-modal adaptability, and controllable approximation error. The framework encompasses basis-function expansions, statistical modeling, tensor decomposition, and implicit neural representations, supported by convergence proofs and generalization bounds. Furthermore, we release Continuous-Representation-Zoo, an open-source knowledge repository spanning computer vision, graphics, bioinformatics, and remote sensing—advancing the systematic development of continuous representation research.
This work addresses the limitations of traditional signal modeling, which relies on discrete sampling and struggles to unify multimodal continuous signals or support analytical operations. Viewing implicit neural representations (INRs) through the lens of signal processing, the study models images, audio, 3D geometry, and other modalities as continuous coordinate-based functions, realized via differentiable neural networks. The authors systematically analyze the spectral properties, sampling theory, and multiscale mechanisms underlying INRs and propose structured representation strategies—integrating periodicity, localization, adaptive activation functions, and hash-grid encodings—to reshape the approximation space for enhanced spatial adaptivity and computational efficiency. The resulting framework demonstrates superior performance in inverse problems such as medical and radar imaging, signal compression, and 3D scene reconstruction, advancing both the theoretical understanding and practical applicability of INRs as learnable continuous signal models.
This paper addresses the poorly understood dynamical mechanisms underlying neural network latent spaces. We model post-trained networks as dynamical systems defined on latent manifolds and, for the first time, analytically derive a data-agnostic, training-free implicit vector field directly from model parameters. Methodologically, we derive the latent flow field via iterative encoder-decoder mappings and combine attractor detection with trajectory evolution analysis to uncover the intrinsic dynamical origins of generalization and memorization. Our key contributions are: (1) establishing the first zero-shot prior knowledge extraction paradigm grounded in manifold dynamics; (2) enabling interpretable, step-by-step analysis of generalization and memorization mechanisms throughout the entire training process; and (3) achieving high-accuracy out-of-distribution (OOD) sample detection on vision foundation models—without additional training or labeled data.
Existing conditional neural fields (CNFs) suffer from limited performance on fine-grained geometric reasoning tasks—such as classification, segmentation, and reconstruction—due to the lack of explicit modeling of local geometry (e.g., locality, orientation) in their latent spaces. To address this, we propose Equivariant Neural Fields (ENFs), the first CNF framework incorporating *implicit geometric equivariance*. ENFs achieve explicit geometric alignment and equivariant mapping between latent space and continuous signals via geometry-aware cross-attention, coupling neural field decoding with point-cloud–based geometric latent variables. These variables exhibit interpretable rotation/translation covariance, enabling geometric reasoning and local weight sharing. Our method encompasses geometric latent variable modeling, equivariant cross-attention, neural field conditioning, and joint point-cloud–field optimization, efficiently implemented in JAX. Experiments demonstrate that ENFs consistently outperform geometry-agnostic baselines across classification, segmentation, prediction, reconstruction, and generation tasks, significantly improving geometric fidelity of latent representations and generalization efficiency.
Existing Coordinate-MLPs for audio implicit representation lack systematic investigation and suffer from sensitivity to hyperparameters, reliance on complex positional encodings, and fragile initialization schemes. This work establishes the first comprehensive benchmark for audio signals using Coordinate-MLPs, evaluating combinations of three positional encoding strategies and sixteen activation functions. Furthermore, it introduces Fourier-ASR, a novel framework grounded in Fourier series and the Kolmogorov–Arnold representation theorem, which incorporates a Fourier-KAN network and a frequency-adaptive learning strategy (FaLS) to achieve robust audio representation without any additional positional encoding. Experiments demonstrate that the proposed method significantly outperforms conventional Coordinate-MLPs on both speech and music datasets, effectively modeling high-frequency components and mitigating low-frequency overfitting—all without requiring meticulous hyperparameter tuning.
This work addresses the joint problem of intrinsic dimension estimation and geometry-invariant embedding learning for nonlinear manifold-structured data. We propose an autoencoder framework incorporating orthogonality constraints on hidden-layer gradients. Methodologically, we establish, for the first time, a theoretical connection between gradient orthogonality in neural network latent spaces and the local tangent space dimension of the underlying manifold; this enables simultaneous intrinsic dimension estimation, learning of invertible embedding mappings, and construction of coordinate-invariant representations under local Lie group actions on low-dimensional submanifolds. Our key contribution lies in unifying gradient orthogonality with differential-geometric structure, thereby extending invariant representation learning to continuous group actions. Experiments on standard benchmarks demonstrate accurate intrinsic dimension estimation, disentangled representations, and robust group-invariant embeddings, validating both theoretical soundness and algorithmic robustness.
Traditional grid-based methods struggle to effectively reconstruct continuous environmental fields from sparse and irregular ecological observation data. This work proposes implicit neural representations (INRs) as a coordinate-driven modeling framework, leveraging coordinate-based neural networks to directly learn spatial or spatiotemporal continuous fields. This approach inherently supports resolution-agnostic querying, preserves spatial coherence, and offers controllable computational costs. The method integrates seamlessly into existing ecological analysis pipelines and achieves stable, high-fidelity reconstructions in tasks such as species distribution modeling, phenological dynamics, and morphological segmentation. Empirical results demonstrate that INRs match or exceed the performance of classical smoothing techniques and tree-based models, highlighting their strong scalability and practical utility for ecological applications.
This work addresses the limited expressivity of positional encoding in conventional implicit neural representations and the high-resolution requirements of existing grid-based encodings for effective learning. The authors introduce a novel approach that models positional encoding as a series of projected points governed by specific motion patterns across different frequencies. Building upon this formulation, they propose a basis decomposition strategy to construct a learnable, grid-based positional encoding scheme. The method consistently outperforms state-of-the-art techniques across multiple tasks—including image representation, texture compression, and signed distance function modeling—achieving comparable or superior reconstruction and rendering accuracy with approximately 25% fewer parameters.
Traditional neural fields exhibit slow convergence and limited scalability when modeling high-dimensional scientific signals, particularly in efficiently handling spatiotemporal and multivariate data. This work proposes a transferable neural field feature mechanism that integrates implicit neural representations with amortized optimization strategies to enable rapid fitting across time steps and ensemble simulations. The approach dramatically accelerates the reconstruction process, reducing the required number of iterations by an order of magnitude while improving early-stage reconstruction quality by over 10 dB. Furthermore, it consistently enhances the accuracy of key physical quantities—such as density gradients and vorticity—across diverse scientific scenarios including turbulence, fluid–material interactions, and astrophysical simulations.
This work addresses the high computational cost and limited temporal modeling capability of traditional implicit neural representations (INRs) when handling time-varying volumetric data, which typically rely on dense spatiotemporal coordinate sampling. The authors propose reformulating time-varying volumes as collections of time series indexed by spatial locations and replacing point-wise scalar supervision with sequence-level supervision. To better capture heterogeneous temporal dynamics, they introduce a Mixture-of-Experts (MoE) architecture that adaptively allocates model capacity across different spatial regions. This approach substantially reduces training overhead while improving reconstruction quality, demonstrates compatibility with various existing INR frameworks, and outperforms current state-of-the-art methods across multiple evaluation metrics.