Institution profile

Disney Research

Industry researchnorthamerica · us
Official website
Research library36linked papers
Opportunities0open roles
Selected work

Representative Papers

Design and Control of a Bipedal Robotic Character

Jul 15, 2024Robotics

To address the limited expressiveness and poor terrain adaptability of bipedal robots in entertainment applications, this paper proposes a dynamic control framework tailored for humanoid stage performances. Methodologically: (1) it introduces a character-driven mechanical design that jointly optimizes artistic expressivity and locomotion robustness; (2) it develops an action-conditioned reinforcement learning controller enabling real-time synthesis and blending of multi-source motion animations; and (3) it integrates online dynamic gait planning with an intuitive human–robot interaction interface. Experimental results demonstrate stable locomotion over complex terrains and high-fidelity, low-latency stage performances with enhanced expressivity. The system significantly improves affective human–robot connection and audience immersion. This work establishes a novel paradigm for entertainment robotics that unifies artistic expression with adaptive motor intelligence.

4 citationsRead paper

RePlay: Retrieval-Based Voice Playback for Multi-Turn spoken dialogue

Sep 25, 2026

This study addresses the challenge of simultaneously achieving content controllability and low latency in full-duplex spoken dialogue systems by proposing PersonaPlex, a retrieval-based speech playback framework. The method leverages internal hidden states as retrieval queries to precisely invoke pre-recorded audio segments for multi-turn conversations via semantic matching. By streamlining critical network layers, replacing generative modules with a lightweight retrieval head, and integrating an efficient turn-taking mechanism with end-to-end speech modeling, PersonaPlex effectively balances response speed and content controllability. Experimental results demonstrate that the system achieves a median latency of merely 383ms, yielding a 3–7× speedup over cascaded systems of comparable quality. Furthermore, user preference evaluations indicate that PersonaPlex significantly outperforms fast cascaded baselines while approaching the performance of high-quality, slower systems.

0 citationsRead paper

Interactive Generative Motion Editing via Scheduled Inpainting

Jul 31, 2026

Existing motion editing methods struggle to simultaneously achieve large-scale structural modifications and faithful preservation of original motion, while generative models often lack interactive editing capabilities. This work proposes scheduled inpainting—a novel approach that dynamically controls the balance between preserving and generating motion in specific spatiotemporal regions during inference of a generative model. For the first time, this method unifies generative motion synthesis with interactive editing. It enables flexible operations such as extension, stitching, and composition while maintaining natural motion quality, and achieves high-precision editing through fine-grained spatiotemporal control. Experiments demonstrate that the proposed method outperforms four baselines across multiple tasks, and both ablation studies and user evaluations confirm its effectiveness and practical utility.

0 citationsRead paper
Recent publications

Latest Papers

RePlay: Retrieval-Based Voice Playback for Multi-Turn spoken dialogue

Sep 25, 2026

This study addresses the challenge of simultaneously achieving content controllability and low latency in full-duplex spoken dialogue systems by proposing PersonaPlex, a retrieval-based speech playback framework. The method leverages internal hidden states as retrieval queries to precisely invoke pre-recorded audio segments for multi-turn conversations via semantic matching. By streamlining critical network layers, replacing generative modules with a lightweight retrieval head, and integrating an efficient turn-taking mechanism with end-to-end speech modeling, PersonaPlex effectively balances response speed and content controllability. Experimental results demonstrate that the system achieves a median latency of merely 383ms, yielding a 3–7× speedup over cascaded systems of comparable quality. Furthermore, user preference evaluations indicate that PersonaPlex significantly outperforms fast cascaded baselines while approaching the performance of high-quality, slower systems.

0 citationsRead paper

Interactive Generative Motion Editing via Scheduled Inpainting

Jul 31, 2026

Existing motion editing methods struggle to simultaneously achieve large-scale structural modifications and faithful preservation of original motion, while generative models often lack interactive editing capabilities. This work proposes scheduled inpainting—a novel approach that dynamically controls the balance between preserving and generating motion in specific spatiotemporal regions during inference of a generative model. For the first time, this method unifies generative motion synthesis with interactive editing. It enables flexible operations such as extension, stitching, and composition while maintaining natural motion quality, and achieves high-precision editing through fine-grained spatiotemporal control. Experiments demonstrate that the proposed method outperforms four baselines across multiple tasks, and both ablation studies and user evaluations confirm its effectiveness and practical utility.

0 citationsRead paper

BFMTrack: Latent Sequence Optimization for Physics-Based Motion Tracking with Behavioral Foundation Models

Jun 23, 2026

Existing behavioral foundation models (BFMs) struggle to accurately track time-varying targets—such as complex motion sequences—due to the absence of temporal dynamics in their latent spaces. This work proposes Latent Sequence Optimization (LSO), a method that directly optimizes temporally coherent latent trajectories within the BFM latent space. By integrating physics-based simulation rollouts with policy gradient updates, LSO achieves precise motion tracking without requiring handcrafted reward functions. The approach innovatively incorporates temporally correlated noise modeling, substantially enhancing trajectory smoothness and detail fidelity, thereby overcoming the limitation of BFMs to time-invariant tasks. Experiments demonstrate that LSO enables high-fidelity, highly generalizable motion reproduction across diverse scenarios, including dense trajectory tracking, sparse keyframe control, and real-world deployment on humanoid robots.

0 citationsRead paper