Institution profile

Joy Future Academy

Academic institution
Research library27linked papers
Opportunities0open roles
Selected work

Representative Papers

ReCast: Attribution-Oriented Step Representation Learning for LLM-Based Agent Systems

Oct 08, 2026

This study addresses the challenges of tracing early-step errors in LLM agents and the limited attribution capability of existing hidden-state representations by proposing ReCast. This method extracts complementary pattern and bias features through layer selection and feature engineering, and trains an encoder via contrastive learning with a ranking objective to transform frozen LLM hidden states into step-level representations tailored for root cause localization. The main contributions include the ReCast framework and the release of the ReCast-2K dataset. Experimental results demonstrate that the proposed approach achieves state-of-the-art performance on the Hit@1 metric across four benchmarks, outperforming the strongest baseline by 5.65 and 9.19 percentage points, respectively.

0 citationsRead paper

SGF+: Decoupling Gradient Flows for Autoregressive Video Generation

Oct 07, 2026

This study addresses the gradient conflict in autoregressive video generation arising from shared parameters between denoising and context writing, which constrains generation quality. To overcome this limitation, we propose SGF+, a method that decouples gradient flows by assigning independent parameters to these distinct functional roles. The modules are interconnected via a causal attention mechanism, enabling joint optimization under the original objective without requiring auxiliary losses. Our approach significantly enhances long-horizon temporal consistency. Notably, when trained on merely five seconds of data, the model is capable of continuously generating high-quality videos spanning up to 24 hours, demonstrating substantial improvements in both efficiency and output fidelity for extended video synthesis.

0 citationsRead paper

JoyAI-Voice 2.0: A Full-Continuous Autoregressive Speech Generation Model with Semantic-Acoustic Joint Representation

Sep 29, 2026

This study addresses the challenges of decoupling and continuously modeling semantic and acoustic features in high-fidelity human-like speech generation by proposing a fully continuous dual-encoder architecture. The method decomposes speech into joint semantic-acoustic representations, which are planned by a causal autoregressive Transformer and subsequently synthesized into 48kHz audio via a local diffusion model. Generation quality is further optimized through the synergistic integration of flow matching, supervised fine-tuning, and DiffusionNFT reinforcement learning. Experimental results demonstrate that Seed-TTS reduces the word error rate to 2.51%, achieving a 14.9% relative improvement over the baseline, while attaining state-of-the-art performance on both the InstructTTSEval and MDVD-Eval benchmarks.

0 citationsRead paper

WorldLine: Action-Driven Visual Simulation for Robotic Manipulation

Sep 29, 2026

This study addresses the high costs of robot learning, imprecise action following in video models, and the difficulty of sharing data across embodiments by proposing a visual simulator that decouples dynamics learning from action grounding. Methodologically, it introduces an image-space action representation to unify control interfaces, combining relational regularization with few-step distillation for efficient causal reasoning. The model learns dynamics from 10,000 hours of unlabeled videos and achieves grounding using 1,000 hours of trajectory data, further incorporating multi-view training and failure augmentation. Experimental results demonstrate that the Intersection over Union (IoU) for failure trajectories improves by 0.16, prediction accuracy reaches 74%, and zero-shot transfer increases task success rates by up to 21.4%.

0 citationsRead paper

Where and When to Force: Routed Forcing for Streaming Avatars

Sep 25, 2026

This study addresses the dynamics and diversity collapse caused by distribution matching distillation in audio-driven streaming talking head generation. To overcome this, we propose Routed Forcing, which introduces a novel routing mechanism based on regional heterogeneity to integrate data-forcing and distribution-matching distillation. Specifically, our method adaptively routes differentiated supervision strategies according to semantic regions—person, mouth, and background—and noise stages, effectively balancing dynamic expressiveness, lip synchronization, and background stability. Experimental results demonstrate that, compared with the Self-Forcing baseline, Routed Forcing improves dynamics by 45% and diversity by 7–25%, while preserving high video quality and precise lip synchronization.

0 citationsRead paper
Recent publications

Latest Papers

ReCast: Attribution-Oriented Step Representation Learning for LLM-Based Agent Systems

Oct 08, 2026

This study addresses the challenges of tracing early-step errors in LLM agents and the limited attribution capability of existing hidden-state representations by proposing ReCast. This method extracts complementary pattern and bias features through layer selection and feature engineering, and trains an encoder via contrastive learning with a ranking objective to transform frozen LLM hidden states into step-level representations tailored for root cause localization. The main contributions include the ReCast framework and the release of the ReCast-2K dataset. Experimental results demonstrate that the proposed approach achieves state-of-the-art performance on the Hit@1 metric across four benchmarks, outperforming the strongest baseline by 5.65 and 9.19 percentage points, respectively.

0 citationsRead paper

SGF+: Decoupling Gradient Flows for Autoregressive Video Generation

Oct 07, 2026

This study addresses the gradient conflict in autoregressive video generation arising from shared parameters between denoising and context writing, which constrains generation quality. To overcome this limitation, we propose SGF+, a method that decouples gradient flows by assigning independent parameters to these distinct functional roles. The modules are interconnected via a causal attention mechanism, enabling joint optimization under the original objective without requiring auxiliary losses. Our approach significantly enhances long-horizon temporal consistency. Notably, when trained on merely five seconds of data, the model is capable of continuously generating high-quality videos spanning up to 24 hours, demonstrating substantial improvements in both efficiency and output fidelity for extended video synthesis.

0 citationsRead paper

JoyAI-Voice 2.0: A Full-Continuous Autoregressive Speech Generation Model with Semantic-Acoustic Joint Representation

Sep 29, 2026

This study addresses the challenges of decoupling and continuously modeling semantic and acoustic features in high-fidelity human-like speech generation by proposing a fully continuous dual-encoder architecture. The method decomposes speech into joint semantic-acoustic representations, which are planned by a causal autoregressive Transformer and subsequently synthesized into 48kHz audio via a local diffusion model. Generation quality is further optimized through the synergistic integration of flow matching, supervised fine-tuning, and DiffusionNFT reinforcement learning. Experimental results demonstrate that Seed-TTS reduces the word error rate to 2.51%, achieving a 14.9% relative improvement over the baseline, while attaining state-of-the-art performance on both the InstructTTSEval and MDVD-Eval benchmarks.

0 citationsRead paper

WorldLine: Action-Driven Visual Simulation for Robotic Manipulation

Sep 29, 2026

This study addresses the high costs of robot learning, imprecise action following in video models, and the difficulty of sharing data across embodiments by proposing a visual simulator that decouples dynamics learning from action grounding. Methodologically, it introduces an image-space action representation to unify control interfaces, combining relational regularization with few-step distillation for efficient causal reasoning. The model learns dynamics from 10,000 hours of unlabeled videos and achieves grounding using 1,000 hours of trajectory data, further incorporating multi-view training and failure augmentation. Experimental results demonstrate that the Intersection over Union (IoU) for failure trajectories improves by 0.16, prediction accuracy reaches 74%, and zero-shot transfer increases task success rates by up to 21.4%.

0 citationsRead paper

Where and When to Force: Routed Forcing for Streaming Avatars

Sep 25, 2026

This study addresses the dynamics and diversity collapse caused by distribution matching distillation in audio-driven streaming talking head generation. To overcome this, we propose Routed Forcing, which introduces a novel routing mechanism based on regional heterogeneity to integrate data-forcing and distribution-matching distillation. Specifically, our method adaptively routes differentiated supervision strategies according to semantic regions—person, mouth, and background—and noise stages, effectively balancing dynamic expressiveness, lip synchronization, and background stability. Experimental results demonstrate that, compared with the Self-Forcing baseline, Routed Forcing improves dynamics by 45% and diversity by 7–25%, while preserving high video quality and precise lip synchronization.

0 citationsRead paper