🤖 AI Summary
This study addresses the challenge of transferring user action sequences across heterogeneous environments, where variations in object configurations and spatial constraints pose significant difficulties. To this end, we propose a multi-level representation-based adaptive transfer mechanism. By integrating large language models with scene graph representations, our method achieves generalized modeling of user activities through multi-granularity action abstraction. This enables the prediction and adaptation of action sequences tailored to target environments, thereby overcoming fixed-environment limitations and facilitating cross-spatial semantic alignment. Experimental results demonstrate that the proposed system can generate effective and compliant action sequences in spaces with vastly different object layouts. Ultimately, this work establishes a novel paradigm for cross-environment activity transfer.
📝 Abstract
We present an action sequence transfer system that adaptively transfers user action sequences across different target spaces. Given an input action sequence from a source space and scene graph representations of both the source and target environments, our system predicts a corresponding action sequence in the target space by adapting to the spatial and object constraints of the new environment. To achieve this, we leverage multi-level representations of user activity to generalize actions at varying levels of abstraction. To demonstrate our system, we collect a new scene graph-based dataset derived from the Ego4D GoalStep dataset for evaluation. Results indicate that our system can generate valid action sequences even between spaces with drastically different object configurations.