🤖 AI Summary
This study addresses the challenges of generating human-object interactions (HOIs) in complex 3D scenes, where paired data are scarce and jointly optimizing environment awareness with motion synthesis remains difficult. To this end, it proposes a factorized decoupling framework based on explicit motion affordances. By decomposing the architecture into a scene-conditioned prediction model and an affordance-conditioned HOI generation model, the approach decouples scene understanding from action synthesis. These components are trained independently using complementary supervision sources, enabling the learning of interaction feasibility and dynamic characteristics without requiring human-object-scene triplet paired data. Consequently, this work significantly mitigates object penetration artifacts in complex indoor environments and effectively enhances interaction quality, producing scene-aware human-object interactions that are both physically plausible and highly realistic.
📝 Abstract
Generating realistic human-object interactions (HOI) in complex 3D scenes requires two complementary capabilities: reasoning about interaction feasibility in the environment and synthesizing realistic human-object motion. However, supervision for these capabilities is rarely available jointly at scale. Human-scene datasets provide rich information about environment-aware motion, while human-object datasets capture detailed interaction dynamics, yet paired human-object-scene data remain scarce. We present MAMHOI, an affordance-mediated factorization for scene-aware human-object interaction generation. MAMHOI factorizes scene-aware HOI generation through an explicit motion-affordance interface between scene understanding and motion synthesis: a scene-conditioned model first predicts where and how an interaction can be feasibly executed, and an affordance-conditioned HOI model then generates the corresponding human-object motion. This factorization allows scene understanding and interaction dynamics to be learned from complementary sources of supervision without requiring paired human-object-scene data. Experiments in complex indoor environments show that MAMHOI reduces object--scene penetration while better preserving human--object interaction quality, yielding more realistic and physically feasible scene-aware interactions. Project page: https://leimingyuan.github.io/MAMHOI-project-page/