HEIR: Learning Human-Entity Interactions with Functional Roles

๐Ÿ“… 2026-09-28
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the limitations of conventional Human-Object Interaction (HOI) detection metrics, which overlook complete event structures and role sharing, by introducing the HEIR benchmark and the CoRISP method. CoRISP pioneers a comprehensive event evaluation framework encompassing multiple actions and roles, coupling the assignment process with cardinality and role-multiplicity potential functions. Furthermore, it constructs a compositional role-aware interaction set prediction model that regularizes participantโ€“role sets through shared entity identity prediction, thereby strengthening event-level compositional understanding. Experiments demonstrate that CoRISP improves Set mAP by 2.87 and 3.82 points in scenarios involving repeated roles and shared participants, respectively. Achieving state-of-the-art performance on the V-COCO dataset, this work enables precise modeling of complex multi-agent interactions.
๐Ÿ“ Abstract
Understanding human-entity interactions requires recovering each person-action event's participants, roles, and shared identities. This structure can support embodied agents by clarifying who acts on which entities and how, informing anticipation and coordination in shared environments. Standard HOI metrics score individual links, leaving complete event composition undermeasured. We introduce HEIR (Human-Entity Interactions with Functional Roles), an image benchmark for complete grounded participant-role sets across object, interpersonal, and self-directed interactions. It contains 18,730 images, six roles, 105 actions, and 437 nouns, with shared entities, role changes, and repeated fillers; 51.6% of images contain multiple actors and 62.1% contain multiple actions. HEIR pairs relation AP with complete-set AP and structural evaluation. We also introduce CoRISP (Compositional Role-aware Interaction Set Prediction), which uses shared entity identities to combine role-conditioned evidence and predict normalized participant-role sets. Cardinality and role-multiplicity potentials couple assignments through event size and role composition, with exact per-event normalization. Across 16 baselines, relation and complete-event rankings diverge even after aligning action weights. CoRISP leads the evaluated systems on repeated-role events and shared-participant images in HEIR by 2.87 and 3.82 Set mAP points, respectively. On V-COCO, CoRISP achieves 73.72/76.23 role AP and 61.06/68.59 complete-set AP on two-slot actions under Scenarios 1/2. These results show the value of learning and evaluating event composition alongside individual relations. The code and dataset are publicly available at https://github.com/Kratos-Wen/HEIR.
Problem

Research questions and friction points this paper is trying to address.

Human-Entity Interactions
Functional Roles
Event Composition
HOI Evaluation
Participant-Role Sets
Innovation

Methods, ideas, or system contributions that make the work stand out.

Human-Entity Interactions
Functional Roles
Compositional Role-aware Interaction Set Prediction
Complete-set Evaluation
Shared Entity Identities
๐Ÿ”Ž Similar Papers
2024-08-202024 2nd International Conference on Computer, Vision and Intelligent Technology (ICCVIT)Citations: 2
Di Wen
Di Wen
Karlsruhe Institute of Technology
Fine-grained Action UnderstandingAnomaly DetectionRobustnessUncertainty
W
Wenhao Guo
Karlsruhe Institute of Technology (KIT)
Y
Yuedong Tan
Institute for Computer Science, Artificial Intelligence and Technology (INSAIT)
Y
Yun Huang
Karlsruhe Institute of Technology (KIT)
M
Minheng Wu
Karlsruhe Institute of Technology (KIT)
Z
Zhihang Chen
Karlsruhe Institute of Technology (KIT)
H
Haiwen Sun
Karlsruhe Institute of Technology (KIT)
Fei Teng
Fei Teng
Reader in Intelligent Energy Systems, Imperial College London
Stability-constrained OptimisationCyber-resilient System OperationData Privacy and Trading
Z
Zhiyuan Gao
University of Bremen
Y
Yufeng Zhang
Karlsruhe Institute of Technology (KIT)
Y
Yuanhao Luo
Karlsruhe Institute of Technology (KIT)
J
Jingqi Zhang
Karlsruhe Institute of Technology (KIT)
Yufan Chen
Yufan Chen
Karlsruhe Institute of Technology
Document AnalysisComputer VisionRobust Deep Learning
Junwei Zheng
Junwei Zheng
CV:HCI, KIT; CVG, ETH Zurich
Visual LocalizationScene UnderstandingAssistive Technology
R
Ruiping Liu
Karlsruhe Institute of Technology (KIT)
J
Jiale Wei
Karlsruhe Institute of Technology (KIT)
Kailun Yang
Kailun Yang
Professor. School of Artificial Intelligence and Robotics, Hunan University (HNU); KIT; UAH; ZJU
Computer VisionComputational OpticsIntelligent VehiclesAutonomous DrivingRobotics
Kunyu Peng
Kunyu Peng
Karlsruhe Institute of Technology
video understandingopen set recognitiongeneralizable deep learning