A.D.A.M.O. (Agent for language-Driven Actions with Multimodal Observations): A Visual-Symbolic Framework for Virtual Humans

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of synergistically integrating perception, reasoning, and action for virtual humans in 3D environments by proposing a dual visual-symbolic world model. The method leverages pretrained vision-language models to unify the control loop, enabling language-driven behavior generation through tool invocation and multimodal observation fusion. Furthermore, it establishes a controlled task evaluation framework based on the CD taxonomy. Experimental results demonstrate that semantic annotation significantly reduces perceptual ambiguity, effectively shifting failure modes toward downstream execution stages. Consequently, this work provides an interpretable and unified framework for embodied agents operating in complex three-dimensional settings.
📝 Abstract
Creating believable vh requires the coherent integration of perception, reasoning, and action mediated by language. A central challenge is to combine these components into a control loop grounded in interactive 3D environments. To this end, we present A.D.A.M.O. (Agent for language-Driven Actions with Multimodal Observations), a visual-symbolic framework for language-driven vh that leverages a pretrained vlm with tool calling to unify perception, reasoning, and action within a single control loop. A.D.A.M.O. maintains a dual visual-symbolic world model that combines egocentric visual input and synchronized symbolic state to support grounded task-oriented behavior from natural language prompts. To support diagnostic evaluation, we introduce a controlled task suite organized by a cd taxonomy that breaks down spatial tasks into procedural and linguistic complexity. Experiments in controlled scenes show that semantic labeling strongly influences task completion and failure modes, reducing perceptual ambiguity while shifting failures toward downstream execution, whereas reasoning errors remain comparatively rare.
Problem

Research questions and friction points this paper is trying to address.

Virtual Humans
Multimodal Perception
Reasoning and Action Integration
3D Interactive Environments
Control Loop
Innovation

Methods, ideas, or system contributions that make the work stand out.

Visual-Symbolic Framework
Vision-Language Model
Tool Calling
World Model
Virtual Humans
🔎 Similar Papers
2024-03-15International Symposium on Mixed and Augmented RealityCitations: 0
💼 Related Jobs
No related jobs found.