Institution profile

Wuhan AI Research

Research institutionasia · cn
Research library7linked papers
Opportunities0open roles
Selected work

Representative Papers

CT-OPD: Counterfactual Trace On-Policy Distillation for Diffusion Vision-Language Models

Sep 26, 2026

This study addresses the misalignment between preset masks and model decisions during the post-training of diffusion vision-language models, as well as the sparse feedback inherent in reinforcement learning. To overcome these challenges, we propose a counterfactual trajectory online distillation method. Specifically, this approach extracts masks from the student's current trajectory to reconstruct the teacher's endpoint states, leveraging the teacher's complete response to provide coherent token-level supervision that precisely aligns training signals with the model's revealed decisions. Experimental results demonstrate that our method yields an average improvement of 9.80 points across nine benchmarks. It significantly enhances multimodal understanding and reasoning capabilities while simultaneously improving image generation quality within a unified architecture.

0 citationsRead paper

HarnessWAM: Bridging Prediction and Deliberation in World Action Models

Aug 10, 2026

This work addresses the limitations of World Action Models (WAMs) in complex embodied tasks—specifically, their inadequate planning, state maintenance, and failure recovery stemming from a disconnect between prediction and reasoning. To overcome this, we propose an agent-based framework that employs a vision-language model–driven task manager to maintain a structured scene belief and task graph. High-level semantic plans are projected into sequences of atomic skills that respect both task dependencies and robot capability constraints. An event-driven dual-timescale feedback mechanism, coupled with a lightweight progress estimator, enables a verifiable and recoverable execution loop. Our approach introduces, for the first time, external structured state maintenance and closed-loop decision-making, substantially enhancing WAMs’ global planning and local fault tolerance. Experiments show our method achieves a 59.6% end-to-end success rate (69.9% on subtasks) on RoboMemArena and 23.7% on RoboCerebra Ideal, significantly outperforming existing approaches.

0 citationsRead paper

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects

Aug 08, 2026

This work addresses the challenges of data scarcity and misalignment between semantic representations and fixed acoustic supervision in end-to-end spoken dialogue modeling for low-resource Chinese dialects. To overcome these issues, the authors construct a scalable data pipeline for dialectal spoken dialogue synthesis and propose a two-stage post-training strategy incorporating a self-aligned speech supervision mechanism that dynamically aligns acoustic targets with the model’s evolving semantic representations. This approach achieves the first successful end-to-end spoken dialogue modeling for low-resource Chinese dialects, significantly outperforming existing baselines across multiple dialects. Substantial improvements are observed in dialect consistency, response quality, and speech intelligibility. The complete framework is open-sourced to facilitate future research in this underexplored domain.

0 citationsRead paper
Recent publications

Latest Papers

CT-OPD: Counterfactual Trace On-Policy Distillation for Diffusion Vision-Language Models

Sep 26, 2026

This study addresses the misalignment between preset masks and model decisions during the post-training of diffusion vision-language models, as well as the sparse feedback inherent in reinforcement learning. To overcome these challenges, we propose a counterfactual trajectory online distillation method. Specifically, this approach extracts masks from the student's current trajectory to reconstruct the teacher's endpoint states, leveraging the teacher's complete response to provide coherent token-level supervision that precisely aligns training signals with the model's revealed decisions. Experimental results demonstrate that our method yields an average improvement of 9.80 points across nine benchmarks. It significantly enhances multimodal understanding and reasoning capabilities while simultaneously improving image generation quality within a unified architecture.

0 citationsRead paper

HarnessWAM: Bridging Prediction and Deliberation in World Action Models

Aug 10, 2026

This work addresses the limitations of World Action Models (WAMs) in complex embodied tasks—specifically, their inadequate planning, state maintenance, and failure recovery stemming from a disconnect between prediction and reasoning. To overcome this, we propose an agent-based framework that employs a vision-language model–driven task manager to maintain a structured scene belief and task graph. High-level semantic plans are projected into sequences of atomic skills that respect both task dependencies and robot capability constraints. An event-driven dual-timescale feedback mechanism, coupled with a lightweight progress estimator, enables a verifiable and recoverable execution loop. Our approach introduces, for the first time, external structured state maintenance and closed-loop decision-making, substantially enhancing WAMs’ global planning and local fault tolerance. Experiments show our method achieves a 59.6% end-to-end success rate (69.9% on subtasks) on RoboMemArena and 23.7% on RoboCerebra Ideal, significantly outperforming existing approaches.

0 citationsRead paper

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects

Aug 08, 2026

This work addresses the challenges of data scarcity and misalignment between semantic representations and fixed acoustic supervision in end-to-end spoken dialogue modeling for low-resource Chinese dialects. To overcome these issues, the authors construct a scalable data pipeline for dialectal spoken dialogue synthesis and propose a two-stage post-training strategy incorporating a self-aligned speech supervision mechanism that dynamically aligns acoustic targets with the model’s evolving semantic representations. This approach achieves the first successful end-to-end spoken dialogue modeling for low-resource Chinese dialects, significantly outperforming existing baselines across multiple dialects. Substantial improvements are observed in dialect consistency, response quality, and speech intelligibility. The complete framework is open-sourced to facilitate future research in this underexplored domain.

0 citationsRead paper