Sparse-View Interpretable 3D Animal Behavior Representations for Neural Encoding and Decoding

๐Ÿ“… 2026-09-28
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the challenges of pose estimation information loss and uninterpretable nonlinear embeddings in video-based behavioral representation by proposing the SABLE framework. This method introduces geometric inductive biases, employing a multi-view Transformer to fuse depth and pose priors. It reconstructs interpretable 3D latent variables of animal behavior from sparse binocular views without requiring 3D ground-truth labels. Experimental results demonstrate that SABLE matches or surpasses state-of-the-art methods in neural encoding and decoding performance across multiple datasets while achieving zero-shot generalization across animals. By effectively balancing representational richness with interpretability, this work establishes a novel paradigm for investigating brainโ€“behavior relationships.
๐Ÿ“ Abstract
A deeper understanding of brain function requires a precise, structured characterization of behavior.Yet, extracting behavioral representations from video in a form suitable for scientific analysis remains a fundamental challenge. Many prior studies represent behavior via pose estimation or nonlinear video embeddings. However, pose tracking discards rich information beyond predefined keypoints, while nonlinear video embeddings lack interpretability. We address this limitation with SABLE (Sparse-view Animal Behavior Latent Embeddings), a self-supervised framework that leverages a geometric inductive bias to learn behavior representations.By augmenting a multi-view transformer with priors from monocular depth and pose estimation, SABLE reconstructs 3D animal behavior from extremely sparse views while learning explicit 3D latent structure. Without ground-truth 3D labels, it reliably recovers 3D behavior from two-view videos, whereas state-of-the-art (SOTA) methods fail or yield degenerate solutions. Across the International Brain Lab and Cheese3D datasets, we demonstrate that SABLE learns 3D representations that match or exceed prior SOTA performance in neural encoding and decoding. Once pretrained across animals, SABLE serves as an off-the-shelf model that generalizes zero-shot to unseen animals without animal-specific calibration or retraining. Our method establishes 3D-aware video embeddings that capture complex behavior, opening new avenues for studying brain-behavior relationships.
Problem

Research questions and friction points this paper is trying to address.

behavior representation
3D animal behavior
sparse-view reconstruction
interpretability
neural encoding and decoding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-supervised 3D reconstruction
Sparse-view Transformer
Geometric inductive bias
Interpretable latent embeddings
Zero-shot generalization
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
X
Xinming Dai
Columbia University
Qihang Jin
Qihang Jin
University of Science and Technology of China
T
Tianshu Tan
Harvard University
B
Baiyuan Chen
University of Cambridge
H
Hanrui Lyu
Northwestern University
L
Lenny Aharon
Columbia University
K
Kyle Daruwalla
Cold Spring Harbor Laboratory
X
Xun Helen Hou
Cold Spring Harbor Laboratory
M
Matthew R. Whiteway
Columbia University
Liam Paninski
Liam Paninski
Columbia University
Neural data science
Y
Yizi Zhang
Columbia University