Institution profile

China Academy of Information and Communications Technology

Academic institutionasia · cn
Official website
Research library12linked papers
Opportunities0open roles
Selected work

Representative Papers

KineWorld: Action-Induced Transport Fields for Embodied World Modeling

Oct 05, 2026

This study addresses the misalignment between visual generation objectives and interaction requirements in existing embodied world models, as well as their neglect of critical sparse spatial variations. To this end, we propose KineWorld, a framework that incorporates robot kinematics into the spatial allocation of generative supervision. Specifically, it introduces Kinematic Transport Lifting (KTL) to construct camera-aligned transport fields and proposes Transport-Aware World Diffusion (TAWD) for reweighted flow matching. Combined with video latent grid calibration and a hybrid distribution normalization mechanism, KineWorld enables precise prediction of action consequences. Evaluated on RoboTwin 2.0, the framework achieves a single-view EWMScore-P of 68.95 and a multi-view TWB-Score of 54.82, validating a paradigm shift from appearance fitting toward action consequence modeling.

0 citationsRead paper

LoopMoEVR: Loop-Based Degradation-Aware Mixture-of-Experts for Unified UHD Video Restoration

Oct 04, 2026

This work addresses the parameter redundancy and limited generalization gains of existing unified ultra-high-definition (UHD) video restoration models by proposing a recurrent mixture-of-experts architecture. Specifically, we design a degradation-aware low-rank embedding module and an iterative spatiotemporal adaptive normalization module, while introducing an input-conditioned dynamic depth prediction mechanism. These components collectively establish a lightweight recurrent learning paradigm capable of efficiently handling cross-domain restoration tasks. Requiring only approximately 0.884M trainable parameters, the proposed model achieves state-of-the-art performance on both benchmark datasets and real-world scenarios for tasks such as dehazing and deraining. Ultimately, this approach enables the unified and efficient processing of diverse UHD video restoration tasks with minimal computational overhead.

0 citationsRead paper

Understanding Trajectory Heterogeneity in Federated World Model Learning

Oct 02, 2026

This study addresses the challenge of temporal context fragmentation and long-horizon prediction in federated world models, where trajectory ownership boundaries inherently limit information continuity. Using the MIMIC-IV dataset, we construct clients partitioned by illness severity and conduct hourly clinical forecasting experiments employing ten federated algorithms, including FedAvg and FedProx, integrated with local history rules and paired rolling evaluation techniques. Our contributions reveal how participation rates and data ownership constrain long-window coverage, demonstrating that 32-step window coverage reaches merely 7.55%–21.36%. Furthermore, we quantify the significant error escalation in FedAvg induced by fine-grained client partitioning. Ultimately, this work establishes a multidimensional evaluation framework for federated learning in clinical time-series forecasting, providing critical insights into the interplay between federation dynamics and predictive horizon limitations.

0 citationsRead paper

A Resource-Aware Behavior Reconstruction and Hierarchical Semantic Learning Framework for Host Intrusion Detection

Oct 01, 2026

This study addresses the challenge of extracting attack patterns from interleaved system call sequences by proposing ReSHID, a novel detection framework. The method introduces a resource-aware sequence reconstruction technique based on bipartite matching to recover continuous behavioral semantics. By integrating file descriptor lifecycle tracking, it constructs subject graphs and employs GATv2 graph neural networks with hierarchical semantic learning to capture multi-subject collaborative attack features, ultimately achieving efficient discrimination via a lightweight linear classifier. Experimental results demonstrate that ReSHID attains an F1 score of 98.64% and a ROC-AUC of 99.80%, while reducing n-gram features by 75.2%. Its overall performance significantly surpasses that of existing methods.

0 citationsRead paper

CodeMimicry: Exploiting Safety Generalization Lag in Large Language Models via Structured Code Completion

Sep 30, 2026

This study addresses the generalization lag of safety alignment in large language models within the code domain, which enables malicious intents to circumvent defenses via legitimate code structures. This work formally defines this failure mode for the first time and proposes a fully automated black-box jailbreak framework based on structured object-oriented code generation. Furthermore, it integrates latent space projection with activation steering to elucidate the underlying mechanisms of this vulnerability. Extensive experiments demonstrate that the proposed approach achieves an attack success rate of 96.25% across eight mainstream commercial large language models, requiring only 1.51 queries on average. These results significantly outperform existing baselines, highlighting both the efficacy of the method and the critical security gaps in current alignment strategies for code-generating models.

0 citationsRead paper
Recent publications

Latest Papers

KineWorld: Action-Induced Transport Fields for Embodied World Modeling

Oct 05, 2026

This study addresses the misalignment between visual generation objectives and interaction requirements in existing embodied world models, as well as their neglect of critical sparse spatial variations. To this end, we propose KineWorld, a framework that incorporates robot kinematics into the spatial allocation of generative supervision. Specifically, it introduces Kinematic Transport Lifting (KTL) to construct camera-aligned transport fields and proposes Transport-Aware World Diffusion (TAWD) for reweighted flow matching. Combined with video latent grid calibration and a hybrid distribution normalization mechanism, KineWorld enables precise prediction of action consequences. Evaluated on RoboTwin 2.0, the framework achieves a single-view EWMScore-P of 68.95 and a multi-view TWB-Score of 54.82, validating a paradigm shift from appearance fitting toward action consequence modeling.

0 citationsRead paper

LoopMoEVR: Loop-Based Degradation-Aware Mixture-of-Experts for Unified UHD Video Restoration

Oct 04, 2026

This work addresses the parameter redundancy and limited generalization gains of existing unified ultra-high-definition (UHD) video restoration models by proposing a recurrent mixture-of-experts architecture. Specifically, we design a degradation-aware low-rank embedding module and an iterative spatiotemporal adaptive normalization module, while introducing an input-conditioned dynamic depth prediction mechanism. These components collectively establish a lightweight recurrent learning paradigm capable of efficiently handling cross-domain restoration tasks. Requiring only approximately 0.884M trainable parameters, the proposed model achieves state-of-the-art performance on both benchmark datasets and real-world scenarios for tasks such as dehazing and deraining. Ultimately, this approach enables the unified and efficient processing of diverse UHD video restoration tasks with minimal computational overhead.

0 citationsRead paper

Understanding Trajectory Heterogeneity in Federated World Model Learning

Oct 02, 2026

This study addresses the challenge of temporal context fragmentation and long-horizon prediction in federated world models, where trajectory ownership boundaries inherently limit information continuity. Using the MIMIC-IV dataset, we construct clients partitioned by illness severity and conduct hourly clinical forecasting experiments employing ten federated algorithms, including FedAvg and FedProx, integrated with local history rules and paired rolling evaluation techniques. Our contributions reveal how participation rates and data ownership constrain long-window coverage, demonstrating that 32-step window coverage reaches merely 7.55%–21.36%. Furthermore, we quantify the significant error escalation in FedAvg induced by fine-grained client partitioning. Ultimately, this work establishes a multidimensional evaluation framework for federated learning in clinical time-series forecasting, providing critical insights into the interplay between federation dynamics and predictive horizon limitations.

0 citationsRead paper

A Resource-Aware Behavior Reconstruction and Hierarchical Semantic Learning Framework for Host Intrusion Detection

Oct 01, 2026

This study addresses the challenge of extracting attack patterns from interleaved system call sequences by proposing ReSHID, a novel detection framework. The method introduces a resource-aware sequence reconstruction technique based on bipartite matching to recover continuous behavioral semantics. By integrating file descriptor lifecycle tracking, it constructs subject graphs and employs GATv2 graph neural networks with hierarchical semantic learning to capture multi-subject collaborative attack features, ultimately achieving efficient discrimination via a lightweight linear classifier. Experimental results demonstrate that ReSHID attains an F1 score of 98.64% and a ROC-AUC of 99.80%, while reducing n-gram features by 75.2%. Its overall performance significantly surpasses that of existing methods.

0 citationsRead paper

CodeMimicry: Exploiting Safety Generalization Lag in Large Language Models via Structured Code Completion

Sep 30, 2026

This study addresses the generalization lag of safety alignment in large language models within the code domain, which enables malicious intents to circumvent defenses via legitimate code structures. This work formally defines this failure mode for the first time and proposes a fully automated black-box jailbreak framework based on structured object-oriented code generation. Furthermore, it integrates latent space projection with activation steering to elucidate the underlying mechanisms of this vulnerability. Extensive experiments demonstrate that the proposed approach achieves an attack success rate of 96.25% across eight mainstream commercial large language models, requiring only 1.51 queries on average. These results significantly outperform existing baselines, highlighting both the efficacy of the method and the critical security gaps in current alignment strategies for code-generating models.

0 citationsRead paper