🤖 AI Summary
This study addresses the lack of contextual adaptability and realism in activity generation within virtual environments. We propose a structured framework leveraging multimodal large language models (MLLMs) to jointly parse scene layout, semantics, and object identities, enabling vision-language co-reasoning. The resulting abstract scene representation is dynamically coupled with a human activity knowledge base to drive context-aware, multi-agent behavior generation and real-time optimization. We introduce, for the first time, a three-level coupled mechanism—“scene understanding → knowledge injection → behavior alignment”—which significantly enhances situational consistency, plausibility of multi-agent interactions, and dynamic responsiveness. Experimental results demonstrate that our approach outperforms existing methods in both perceptual realism and contextual fidelity.
📝 Abstract
In this paper, we investigate the use of multimodal large language models (MLLMs) for generating virtual activities, leveraging the integration of vision-language modalities to enable the interpretation of virtual environments. Our approach recognizes and abstracts key scene elements including scene layouts, semantic contexts, and object identities with MLLMs' multimodal reasoning capabilities. By correlating these abstractions with massive knowledge about human activities, MLLMs are capable of generating adaptive and contextually relevant virtual activities. We propose a structured framework to articulate abstract activity descriptions, emphasizing detailed multi-character interactions within virtual spaces. Utilizing the derived high-level contexts, our approach accurately positions virtual characters and ensures that their interactions and behaviors are realistically and contextually appropriate through strategic optimization. Experiment results demonstrate the effectiveness of our approach, providing a novel direction for enhancing the realism and context-awareness in simulated virtual environments.