COSMI: COmpositional Synthesis of Multi-object Interactions

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the scarcity of multi-object interaction data arising from high acquisition costs by proposing a combinatorial data augmentation strategy grounded in interaction locality. Methodologically, large-scale datasets are constructed by concatenating single-object interaction segments to train a text-driven diffusion Transformer. Furthermore, a shared-weight object slot generation mechanism is introduced, integrated with language model filtering, geometric consistency verification, and mirror augmentation techniques to enhance generation quality. The resulting dataset comprises 222,000 sequences. Experimental evaluations demonstrate that the proposed approach significantly outperforms existing baselines in generalizing to unseen object and interaction combinations.
📝 Abstract
Generative models of human-object interaction are bounded by the data that exists: everyday activities involve several objects, but most captured datasets record one at a time, as multi-object capture is combinatorially expensive. Our observation is that interactions are local, so single-object captures already contain the parts of multi-object activities. We compose them: contact-consistent clips of single interactions, mirrored to balance the hands, transfer between bodies, and a language model and geometric checks admit only the pairings that are plausible, semantically and physically. Therefore, the dataset grows combinatorially with the clips rather than recording time. The COSMI dataset holds 222k sequences and 275 hours with up to five objects, nearly thirty times the largest multi-object capture, and can be extended by adding datasets or even hand-object recordings. On this data we train the COSMI method, a text-to-interaction diffusion transformer that follows how the data is built: weight-shared object slots generate a variable number of objects, predicted relative to the body parts that move them. On a benchmark with an unseen object and unseen interaction combinations, models trained on the dataset generalize to the unseen combinations. COSMI outperforms baselines in text alignment and contact accuracy, where its margin is largest on the unseen object. Code, models, and the dataset pipeline will be released on the project page: https://ptrvilya.github.io/cosmi.
Problem

Research questions and friction points this paper is trying to address.

human-object interaction
generative models
multi-object interactions
data scarcity
compositional synthesis
Innovation

Methods, ideas, or system contributions that make the work stand out.

Compositional Synthesis
Multi-object Interaction
Diffusion Transformer
Weight-shared Object Slots
Zero-shot Generalization
🔎 Similar Papers