Robot Learning to Communicate through Projected Visual Abstractions

📅 2026-07-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge that existing robots struggle to leverage visually abstract projections of their bodies—such as shadows—for effective communication. The authors propose a dynamic shadow generation method based on a high-degree-of-freedom soft robotic hand, which jointly optimizes hand poses through a differentiable self-model integrated with physical simulation, thereby enabling robot-driven dynamic shadow expression for the first time. Key contributions include the formulation of expressive region objectives, temporal smoothness regularization, and a keyframe-based optimization strategy, combining soft-body structural modeling with gradient-based optimization. The approach is successfully demonstrated on both simulated and physical platforms, generating sign language gestures, shadow puppetry sequences, and animal motion imitations, thereby validating the feasibility of using robotic shadows for visual storytelling and nonverbal communication.
📝 Abstract
Humans routinely communicate through abstractions of their bodies, including shadows, silhouettes, and reflections. Yet robots remain largely confined to expressing themselves through their physical morphology. Enabling robots to communicate through such projected visual abstractions requires reasoning not only about bodily motion but also about how that motion is transformed into an external representation perceived by an observer. Among these abstractions, shadows provide a particularly compelling example because they emerge directly from the robot's embodiment while remaining visually distinct from the body itself. Here, we present a robotic system capable of dynamic shadow expression using a 21-degree-of-freedom dexterous hand with compliant soft skin and a learned shadow self-model. The soft-skinned embodiment reduces light leakage to produce visually continuous silhouettes, while the differentiable self-model learns the mapping between hand configurations and projected shadow appearance through task-agnostic self-exploration. Given a target shadow image or video, the robot optimizes its hand configurations through gradient-based search over 1 the learned self-model and refines the solution through collision-aware simulation to obtain physically feasible motions. For dynamic shadow performance, we further introduce expressive-region objectives, temporal smoothness regularization, and keyframe-based optimization to preserve visually important motion cues while reducing optimization complexity. We demonstrate robotic shadow expression across sign-language gestures, hand-shadow puppetry, and animal motion imitation in both simulation and physical experiments. These results establish a framework for enabling robots to manipulate projected visual abstractions of themselves for communication and visual storytelling.
Problem

Research questions and friction points this paper is trying to address.

robot communication
visual abstraction
shadow expression
embodied intelligence
projected representation
Innovation

Methods, ideas, or system contributions that make the work stand out.

projected visual abstractions
shadow expression
differentiable self-model
soft-skinned robotic hand
gradient-based motion optimization