🤖 AI Summary
This study addresses the inefficiency of meta-learning in neural fields caused by the coupling of latent representation inference and decoder optimization, as well as information loss from first-order approximations. To overcome these limitations, this work proposes MetaLF, a model that adopts an encoder-optimization perspective by introducing an equivariant Transformer self-attention mechanism to coordinate latent point clouds, enabling end-to-end second-order meta-learning. By unifying second-order differentiation, parameterization, and task supervision, MetaLF decouples internal encoding objectives from external supervision, establishing a unified framework supporting both reconstruction and semantic prediction. Experimental results demonstrate that the proposed model requires only three to five gradient updates to substantially improve reconstruction fidelity for images and 3D shapes while effectively facilitating multimodal semantic prediction.
📝 Abstract
Conditional neural fields represent signals continuously, but their effectiveness depends on how the conditional latent representations are inferred from observed data. In meta-learning, this encoding occurs through gradient updates induced by the decoder, tying representation learning directly to decoder design. We formalize this connection by interpreting latent optimization as an optimization encoder, unifying the roles of second-order differentiation, latent parameterization, and task supervision. This concept enables second-order meta-learning for end-to-end training of the encoding procedure alongside the decoder, and clarifies which learning pathway first-order approximations discard. Guided by this view, we introduce Attentive Latent Fields (MetaLF), an equivariant transformer-based neural field that contextualizes a latent pointcloud through self-attention. These interactions shape both field predictions and the updates that construct their representation, allowing local observations to inform coherent non-local structure. Disentangling the inner encoding objective from outer task supervision unifies reconstruction, classification, and segmentation within an end-to-end meta-learning framework, using reconstruction-only latent adaptation at test time. Controlled experiments on polynomial fields link latent coordination to lower effective rank and stronger alignment with the underlying function space. Across image and 3D shape reconstruction, MetaLF improves fidelity within three to five gradient updates, while supporting semantic prediction across images, shapes, and volumes. Together, these findings position the optimization encoder perspective as a unified basis for designing neural fields around how representations are constructed, coordinated, and used.