Score
Design and train implicit neural representations (INRs) that are conditioned on a source identity or source position to produce source-dependent directional weighting or response fields; build models that are optimized from microphone or similar measurements to capture common directional patterns across sources and to generalize directional weighting to unseen source positions.
This paper systematically surveys implicit neural representations (INRs), identifying key performance bottlenecks—particularly limited expressivity and scalability—in inverse problems such as audio synthesis, image reconstruction, 3D scene modeling, and high-dimensional data generation. To address these challenges, we propose the first unified four-dimensional taxonomy—spanning activation functions, positional encodings, joint encoding strategies, and network architectures—and uncover a fundamental trade-off between local bias suppression and fine-grained detail modeling. Through cross-modal benchmarking and gradient-guided structural optimization, we quantitatively evaluate reconstruction fidelity, memory efficiency, and generalization capability. Our contributions include: (i) an open, reproducible experimental benchmark; (ii) identification of three critical research frontiers—activation expressivity, positional encoding robustness, and high-dimensional scalability; and (iii) theoretically grounded, practice-oriented guidance for INR method selection and future development.
This work addresses the limitations of traditional signal modeling, which relies on discrete sampling and struggles to unify multimodal continuous signals or support analytical operations. Viewing implicit neural representations (INRs) through the lens of signal processing, the study models images, audio, 3D geometry, and other modalities as continuous coordinate-based functions, realized via differentiable neural networks. The authors systematically analyze the spectral properties, sampling theory, and multiscale mechanisms underlying INRs and propose structured representation strategies—integrating periodicity, localization, adaptive activation functions, and hash-grid encodings—to reshape the approximation space for enhanced spatial adaptivity and computational efficiency. The resulting framework demonstrates superior performance in inverse problems such as medical and radar imaging, signal compression, and 3D scene reconstruction, advancing both the theoretical understanding and practical applicability of INRs as learnable continuous signal models.
Implicit neural representations (INRs) suffer from suboptimal performance due to the absence of systematic, joint design principles for activation functions and initialization parameters; existing approaches rely on heuristic tuning or exhaustive search, yielding inconsistent results across modalities. Method: We propose the first unified joint optimization framework for INR configuration, modeling discrete activation families (e.g., SIREN, WIRE, FINER) and continuous initialization scales as co-optimized variables, and employing Bayesian optimization for end-to-end automated configuration. Contribution/Results: Our method replaces ad hoc empirical practices with a data-driven, standardized configuration pipeline. Evaluated on multimodal signal reconstruction and 3D shape modeling tasks, it achieves significant improvements in reconstruction accuracy and training stability, demonstrating strong generalizability and cross-modal consistency.
Implicit Neural Representations (INRs) suffer from the spectral bias of MLPs, hindering high-fidelity reconstruction of high-frequency details. To address this, we propose an inductive gradient adjustment method grounded in the empirical Neural Tangent Kernel (eNTK), which— for the first time—formally bridges spectral bias and training dynamics. Our approach dynamically designs a gradient transformation matrix to mitigate bias directionally, without altering network architecture. By establishing a linearized training dynamics model, it enables generalized gradient optimization across diverse INR architectures and tasks. Experiments demonstrate consistent improvements across multiple INR variants (e.g., SIREN, Fourier Features) and reconstruction tasks (images and videos): reconstructed outputs exhibit richer texture, sharper edges, and superior quantitative performance—achieving higher PSNR and lower LPIPS than state-of-the-art training strategies.
Implicit Neural Representations (INRs) commonly suffer from mean regression bias, leading to loss of high-frequency details and poor noise robustness—limiting their effectiveness in signal reconstruction. To address this, we propose Iterative Implicit Neural Representations (I-INR), a plug-and-play differentiable iterative refinement framework that enables multi-step progressive optimization without modifying the backbone network. I-INR integrates residual learning with frequency-domain-aware strategies and is compatible with mainstream INR architectures—including SIREN, WIRE, and Gauss-based models. Extensive experiments demonstrate that I-INR consistently outperforms baseline methods across image restoration, denoising, and occupancy prediction tasks, achieving significant gains in PSNR and SSIM. Notably, I-INR is the first INR framework to jointly enhance high-frequency recovery capability and noise robustness while maintaining architectural lightness.
This work addresses the challenge of directly classifying implicit neural representations (INRs) due to the high dimensionality and complex structure of their weight spaces, as well as the unclear mechanisms governing the distribution of discriminative information. To tackle this, the authors propose a structure-aware hierarchical mixture-of-experts (HMoE) Transformer integrated within a meta-learning framework, enabling conditional computation in weight space. They introduce, for the first time, a structure-aligned MoE architecture for INR learning and develop weight attribution and structured pruning techniques to uncover class-specific substructures, substantially enhancing model interpretability. The proposed method achieves state-of-the-art accuracy on multiple benchmarks, including ImageNet-1K, demonstrating its effectiveness across both low- and high-resolution data.
To address the high computational cost and slow convergence in training implicit neural representations (INRs) for high-resolution signal modeling—stemming from inefficient coordinate sampling—this paper proposes a neural tangent kernel (NTK)-guided dynamic coordinate selection mechanism. Our method quantifies sample importance via the NTK-calibrated norm of the loss gradient, jointly accounting for reconstruction error and inter-coordinate coupling to adaptively select coordinates that contribute most to global function updates. Unlike fixed or heuristic sampling strategies, our approach significantly improves training efficiency under standard MLP architectures. Experiments demonstrate an average 47% reduction in training time while maintaining or even improving reconstruction quality, establishing it as the new state-of-the-art (SOTA) among sampling-based INR acceleration methods.
Implicit Neural Representations (INRs) suffer from spectral bias, limiting their ability to model high-frequency visual and geometric signals. To address this, we propose Dynamic Implicit Neural Representations (DINR), the first INR framework formulated as a continuous-time dynamical system, where features evolve explicitly over time to mitigate spectral bias. Our method introduces dynamical complexity regularization to balance expressivity and generalization, supported theoretically by Rademacher complexity and the neural tangent kernel. DINR employs a differentiable continuous-dynamics architecture enabling end-to-end training. Extensive experiments on image representation, field reconstruction, and data compression demonstrate that DINR significantly improves convergence speed, signal fidelity, and generalization performance—consistently outperforming static INR baselines across all tasks.
This work addresses the lack of rigorous theoretical understanding regarding how weights in implicit neural representations (INRs) encode data semantics. By leveraging the implicit function theorem, the authors establish a differentiable mapping between the data space and the INR weight space for the first time. They further introduce a shared hypernetwork that maps instance embeddings to INR weights, thereby enabling semantic-preserving weight learning. This approach provides theoretical guarantees for semantic encoding within INRs and constructs an interpretable framework for learning in weight space. Experiments on 2D and 3D datasets demonstrate that the proposed method achieves competitive performance on downstream classification tasks compared to existing baselines, validating both its effectiveness and theoretical advantages.
This work addresses the challenges of spectral bias and inter-branch interference in implicit neural representations (INRs) for multi-scale signal modeling, where high-frequency updates often corrupt low-frequency structures. The authors propose a multi-branch INR architecture that aligns the signal spectrum to each branch’s optimal operating range through directional coordinate scaling and incorporates a directional edge-guided loss to achieve functional disentanglement. By innovatively integrating the inverse Fourier scaling theorem with a gradient-based spatially conditioned sparsity prior, the method explicitly separates multi-scale features, effectively eliminating spectral crosstalk and accelerating convergence. Experiments demonstrate significant improvements over current state-of-the-art methods in image reconstruction (+5.16 dB), denoising (+0.65 dB), audio reconstruction (50.02 dB), and 3D reconstruction (IoU 0.999).