EgoPhys: Estimating Peak Contact Force and Mechanical Work from Egocentric Manipulation Video

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of estimating peak contact force and mechanical work during physical interactions from first-person videos, where available cues are inherently local and indirect. To this end, it proposes EgoPhys, a framework that predicts force and work using only RGB video input. The core innovations include Contact-Aware Spatial Aggregation (CASA), which integrates appearance and geometric features to localize interaction regions, and Target-specific Multi-expert Temporal Routing (TMTR), which dynamically models semantic, event-level, and kinematic cues over time. Evaluated on the HOI! dataset, EgoPhys significantly improves prediction accuracy, reducing the mean absolute error for peak contact force and mechanical work to 5.205 N and 0.894 J, respectively.
📝 Abstract
Physically grounded manipulation of articulated objects requires understanding both the maximum forces encountered during contact and the work performed as their parts move. Peak contact force and mechanical work quantify these complementary aspects, but estimating them from egocentric video is challenging because physical interaction cues are local and indirect. Moreover, peak force is associated with brief contact events, whereas mechanical work depends on force-motion coupling throughout the contact duration. To address these challenges, we propose EgoPhys, an RGB-only framework comprising Contact-Aware Spatial Aggregation (CASA) and Target-Specific Multi-Expert Temporal Routing (TMTR). CASA integrates appearance and geometry features to emphasize interaction-relevant cues, while TMTR models semantic, event, and motion cues with specialized temporal experts and routes them separately for force and work prediction. On the test split from Hoi! dataset, EgoPhys substantially improves predictions of peak force and mechanical work, achieving MAEs of \(5.205 \pm 0.584\) $N$ and $0.894 \pm 0.081$ $J$, respectively.
Problem

Research questions and friction points this paper is trying to address.

Egocentric video
Peak contact force
Mechanical work
Articulated object manipulation
Physical interaction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Egocentric Video
Contact-Aware Spatial Aggregation
Multi-Expert Temporal Routing
Peak Contact Force
Mechanical Work
💼 Related Jobs
No related jobs found.
Z
Zhuo Dong
School of Artificial Intelligence, Shandong University, Jinan, China
J
Jianhua Yang
Institute of Automation, Chinese Academy of Sciences, Beijing, China
H
Haohao Li
School of Mechanical Engineering, Tianjin University, Tianjin, China
Y
Yumeng Zhao
School of Artificial Intelligence, Shandong University, Jinan, China
Keji He
Keji He
SDU << CASIA & NUS
Cross-modal LearningEmbodied AI
Yan Huang
Yan Huang
Institute of Automation, Chinese Academy of Sciences
computer visiondeep learningmultimodal learning
Liang Wang
Liang Wang
Institute of Psychology, Chinese Academy of Sciences
ECoGfMRINeuronal oscillationsBrain networksSpatial attention