MotionJEPA: Preventing Temporal Feature Collapse by Capturing Visual Changes in Latent Space
本文针对JEPA训练中特征抑制和潜在表示崩溃问题,提出了一种名为DISReg的新正则化方法,并将其整合进MotionJEPA架构,以促进静态和动态特征的平衡学习。
本文针对JEPA训练中特征抑制和潜在表示崩溃问题,提出了一种名为DISReg的新正则化方法,并将其整合进MotionJEPA架构,以促进静态和动态特征的平衡学习。
This work proposes an open-source, modular data collection pipeline to address the labor-intensive, costly, and poorly reproducible nature of dataset construction from heterogeneous online sources in computational social science. The system employs lightweight natural language instructions to configure workflows, decomposing tabular data curation into entity-level search and structured extraction tasks. By leveraging large language model–based intelligent agents, it achieves a task-agnostic, reusable architecture that operates effectively even without predefined entity lists. Evaluated across multiple representative tasks, the approach attains high accuracy while reducing data acquisition costs by an order of magnitude compared to manual methods, substantially lowering both technical and human resource barriers to entry.
This work addresses multi-task learning by proposing a neuro-modulation-inspired approach to dynamic parameter modeling. Instead of conventional context-conditioning via input concatenation, it constructs smooth manifolds—endowed with predefined topologies (e.g., line, ellipse, torus)—directly in weight space, enabling continuous parameter evolution across tasks. It is the first to formalize neural modulation as a manifold-constrained optimization problem, leveraging topological priors to guide task relationship modeling and generalization. The method integrates variational manifold optimization, constraint-aware volume-minimization loss, and low-dimensional implicit weight parameterization. Experiments demonstrate that linear and elliptical manifolds significantly outperform input-concatenation baselines in noise robustness and image rotation generalization, while also improving out-of-distribution accuracy.
本文针对JEPA训练中特征抑制和潜在表示崩溃问题,提出了一种名为DISReg的新正则化方法,并将其整合进MotionJEPA架构,以促进静态和动态特征的平衡学习。
This work proposes an open-source, modular data collection pipeline to address the labor-intensive, costly, and poorly reproducible nature of dataset construction from heterogeneous online sources in computational social science. The system employs lightweight natural language instructions to configure workflows, decomposing tabular data curation into entity-level search and structured extraction tasks. By leveraging large language model–based intelligent agents, it achieves a task-agnostic, reusable architecture that operates effectively even without predefined entity lists. Evaluated across multiple representative tasks, the approach attains high accuracy while reducing data acquisition costs by an order of magnitude compared to manual methods, substantially lowering both technical and human resource barriers to entry.
This work addresses multi-task learning by proposing a neuro-modulation-inspired approach to dynamic parameter modeling. Instead of conventional context-conditioning via input concatenation, it constructs smooth manifolds—endowed with predefined topologies (e.g., line, ellipse, torus)—directly in weight space, enabling continuous parameter evolution across tasks. It is the first to formalize neural modulation as a manifold-constrained optimization problem, leveraging topological priors to guide task relationship modeling and generalization. The method integrates variational manifold optimization, constraint-aware volume-minimization loss, and low-dimensional implicit weight parameterization. Experiments demonstrate that linear and elliptical manifolds significantly outperform input-concatenation baselines in noise robustness and image rotation generalization, while also improving out-of-distribution accuracy.