🤖 AI Summary
This study addresses the challenge of estimating peak contact force and mechanical work during physical interactions from first-person videos, where available cues are inherently local and indirect. To this end, it proposes EgoPhys, a framework that predicts force and work using only RGB video input. The core innovations include Contact-Aware Spatial Aggregation (CASA), which integrates appearance and geometric features to localize interaction regions, and Target-specific Multi-expert Temporal Routing (TMTR), which dynamically models semantic, event-level, and kinematic cues over time. Evaluated on the HOI! dataset, EgoPhys significantly improves prediction accuracy, reducing the mean absolute error for peak contact force and mechanical work to 5.205 N and 0.894 J, respectively.
📝 Abstract
Physically grounded manipulation of articulated objects requires understanding both the maximum forces encountered during contact and the work performed as their parts move. Peak contact force and mechanical work quantify these complementary aspects, but estimating them from egocentric video is challenging because physical interaction cues are local and indirect. Moreover, peak force is associated with brief contact events, whereas mechanical work depends on force-motion coupling throughout the contact duration. To address these challenges, we propose EgoPhys, an RGB-only framework comprising Contact-Aware Spatial Aggregation (CASA) and Target-Specific Multi-Expert Temporal Routing (TMTR). CASA integrates appearance and geometry features to emphasize interaction-relevant cues, while TMTR models semantic, event, and motion cues with specialized temporal experts and routes them separately for force and work prediction. On the test split from Hoi! dataset, EgoPhys substantially improves predictions of peak force and mechanical work, achieving MAEs of \(5.205 \pm 0.584\) $N$ and $0.894 \pm 0.081$ $J$, respectively.