Harnessing Multimodal Large Language Models for Training-Free Human-Object Interaction Detection

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of existing training-free human-object interaction (HOI) detection methods, specifically their constrained visual exploration and semantic circularity that compromise accuracy. To overcome these challenges, this work proposes HarnessHOI, a novel framework that pioneers a mechanism for feeding interaction hypotheses back into the visual space. By incorporating an interaction-guided perception module, it transforms passive inference into an active closed-loop process. Furthermore, a relation-agnostic geometric arbitration module is introduced to establish a unified criterion across multi-source evidence, effectively disrupting self-reinforcing semantic loops. Built upon multimodal large language models for training-free reasoning, the proposed method achieves state-of-the-art performance under the training-free paradigm on both the HICO-DET and V-COCO datasets.
📝 Abstract
Human-object interaction (HOI) detection aims to localize human-object pairs and recognize their interactions. Traditional supervised methods perform strongly but rely on task-specific training. Recent multimodal large language models (MLLMs) offer a promising route to training-free HOI detection through their broad visual-semantic knowledge and versatile perceptual and reasoning capabilities. However, existing approaches largely invoke these capabilities through loosely coordinated inference stages. This fragmented execution restricts the role of interaction hypotheses in guiding visual exploration, leaving key participants overlooked and local ambiguities unresolved. Furthermore, propagating early semantic assumptions through subsequent visual grounding and relation prediction induces self-reinforcing semantic circularity. To resolve these challenges, we propose HarnessHOI, a training-free framework that transforms passive MLLM inference into an active interaction-centric harness. Specifically, we introduce an interaction-guided perception mechanism that projects emerging interaction hypotheses back into the visual space to discover missing participants and refine ambiguous evidence through targeted observation. Furthermore, a relation-agnostic geometric adjudication module reconciles multi-source evidence to establish a unified spatial basis for grounded interaction reasoning across multiple actions and semantic roles. Extensive experiments on HICO-DET and V-COCO demonstrate that HarnessHOI achieves state-of-the-art performance among training-free methods, confirming the effectiveness of the proposed harness for complex interaction understanding. Code will be released upon publication.
Problem

Research questions and friction points this paper is trying to address.

Human-Object Interaction Detection
Multimodal Large Language Models
Training-Free
Semantic Circularity
Visual Grounding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Training-Free HOI Detection
Multimodal Large Language Models
Interaction-Guided Perception
Relation-Agnostic Geometric Adjudication
HarnessHOI
🔎 Similar Papers
No similar papers found.