🤖 AI Summary
This study addresses the challenge of identifying reliable pose cues for social robots anticipating proxemic touch behaviors across datasets by proposing a fixed-reference-pose residual model. Methodologically, it integrates bounding boxes, mask-based geometric predictors, and temporal networks to innovatively decouple predictions into geometric and pose components via an additive correction mechanism, thereby quantifying cue transferability across different robots. The research reveals asymmetric transfer effects and cross-dataset inversions in head pitch orientation, demonstrating that frozen strategies offer no advantage while simpler baselines prove more competitive. Ultimately, average precision improves from 0.277 to 0.321. Furthermore, findings indicate that neural models detect at most 17% of target interactions, offering novel perspectives for cross-domain measurement in human-robot interaction.
📝 Abstract
Social and service robots in public spaces need to anticipate which nearby person is about to approach and touch them, so that a response can be prepared before contact. It is largely unknown which cues support this anticipation when a model trained with one robot is used on another robot at a different site. We study this question with a fixed-reference pose residual (FRPR) model: a geometry predictor built from the person's bounding box and mask is trained and frozen, and a temporal network then learns from body pose an additive correction to its logit, so that every prediction splits exactly into a geometry term and a pose term. Between two public egocentric datasets recorded by different robots, HUI360 and SSUP-A, with every choice made on source data, the pose correction raised average precision (AP) from 0.277 to 0.321 from SSUP-A to HUI360 and gave no measurable gain in the opposite direction; the same asymmetry held over a stronger, source-selected geometry reference. Freezing gave no AP advantage over joint training, and simple geometric baselines and tree ensembles remained competitive or better, so the construction serves measurement rather than prediction accuracy. A head-orientation residual added small gains in both directions. Post hoc, whether a person faces the camera kept its discriminative direction across datasets, whereas head pitch reversed. With thresholds chosen on source data, the neural models that use geometry detected at most 17% of target interactions. Code and processed data are available at https://github.com/WeiZhou96/FRPR-interaction-anticipation.