Institution profile

Hangzhou Normal University

Academic institutionasia · cn
Official website
Research library52linked papers
Opportunities0open roles
Selected work

Representative Papers

iCATS: Fast Video Generation via Interaction-Aware Sparse Attention and Timestep-Adaptive Sparsity

Oct 08, 2026

This study addresses the quality degradation and efficiency bottlenecks in diffusion Transformer-based video generation caused by approximation errors in sparse attention. To this end, we propose a training-free acceleration framework that introduces a novel quadratic importance estimation based on query-key dot-product interactions, coupled with a signal-to-noise ratio (SNR)-guided dynamic sparsity scheduling strategy. Furthermore, an irregular cluster-tail merging mechanism is designed to optimize GPU kernel utilization. Evaluated on HunyuanVideo and Wan2.1, the proposed framework achieves 2.03× and 1.55× speedups, respectively, while preserving high-fidelity generation quality with PSNR scores of 31.017 dB and 29.301 dB. These results demonstrate that our approach effectively realizes the synergistic optimization of inference efficiency and generation quality.

0 citationsRead paper

Dijkstra Is NOT Greedy: A Global-to-Local Proof of Correctness

Oct 06, 2026

This paper challenges the conventional classification of Dijkstra’s algorithm as a greedy method, rectifying the widespread misconception that local choices inherently lead to global optimality. It proposes a three-step elimination method and a deterministic path-space contraction strategy, which are formally proven in conjunction with optimal substructure and boundary state reduction theory. These analyses reveal that the algorithm fundamentally constitutes a dynamic programming process deriving local solutions from the global shortest path. By establishing this paradigm-shifting "global-to-local" perspective, the study constructs a more concise correctness proof and clarifies the algorithm's essential nature. Ultimately, this work provides a novel theoretical interpretive framework for shortest-path problems in graph theory.

0 citationsRead paper

One-Step Generative Modeling via Training Dynamics Action

Sep 30, 2026

This study addresses the disconnect between theoretical objectives and practical deployment in existing one-step generative models, where training often overlooks the implementation costs of parameter sharing. To bridge this gap, we propose TDAction, a framework that formulates training dynamics as an optimal control problem for the first time, achieving cost-aware optimal transport through local parameter sharing. Specifically, we derive a closed-form Batch Tangent Action-to-Go value function and introduce soft terminal control, stochastic tangent probing, and low-rank approximation techniques to enhance optimization efficiency. By preserving the distributional evolution process while significantly reducing implementation overhead, our method achieves state-of-the-art performance on ImageNet 256×256, yielding an FID below 1.1 without requiring distillation.

0 citationsRead paper

Is Personalized Modality Weighting Actually Personalized? A Controlled Audit of Per-User Weighting Claims in Multimodal Recommenders

Aug 06, 2026

This work investigates whether user-specific modality weighting mechanisms widely adopted in multimodal recommendation systems genuinely capture individual user preferences. To this end, the authors propose an auditing framework comprising two metrics—real-GM and real-shuf—that evaluate personalization efficacy by comparing six personalized weighting methods against global weights and shuffled user-weight assignments, all under a unified collaborative filtering backbone. Experimental results across three short-video and one cross-domain e-commerce dataset reveal that performance gains from most methods stem primarily from increased model capacity rather than authentic user signals, with gating mechanisms often inducing spurious personalization due to their reliance on shared embeddings. Notably, global modality weights already achieve nearly all attainable gains, while personalized weighting shows no consistent improvement; the proposed audit framework effectively identifies architectures that truly encode user-specific patterns.

0 citationsRead paper

Task-Oriented Wave Processing with Stacked Intelligent Metasurfaces: Framework, Fusion, and Challenges

Jul 20, 2026

This work addresses the performance conflicts and resource contention arising from conventional task-agnostic channels in 6G scenarios characterized by deep convergence of heterogeneous services. To overcome these challenges, the paper proposes a novel physical-layer computing paradigm based on stacked intelligent metasurfaces (SIMs). This paradigm reconfigures the wireless environment into a programmable signal processor, establishing a unified mapping framework that directly links service requirements to wave-domain synthesis. By leveraging the deep computational architecture of SIMs, the approach enables intrinsic co-design of sensing, communication, and computing. Notably, it pioneers the use of SIMs for task-oriented wave manipulation, thereby transitioning multi-service systems from mere coexistence to true symbiosis. Numerical results demonstrate that the proposed method effectively mitigates resource conflicts and significantly enhances overall system performance, offering a foundational enabler for service-native 6G architectures.

0 citationsRead paper
Recent publications

Latest Papers

iCATS: Fast Video Generation via Interaction-Aware Sparse Attention and Timestep-Adaptive Sparsity

Oct 08, 2026

This study addresses the quality degradation and efficiency bottlenecks in diffusion Transformer-based video generation caused by approximation errors in sparse attention. To this end, we propose a training-free acceleration framework that introduces a novel quadratic importance estimation based on query-key dot-product interactions, coupled with a signal-to-noise ratio (SNR)-guided dynamic sparsity scheduling strategy. Furthermore, an irregular cluster-tail merging mechanism is designed to optimize GPU kernel utilization. Evaluated on HunyuanVideo and Wan2.1, the proposed framework achieves 2.03× and 1.55× speedups, respectively, while preserving high-fidelity generation quality with PSNR scores of 31.017 dB and 29.301 dB. These results demonstrate that our approach effectively realizes the synergistic optimization of inference efficiency and generation quality.

0 citationsRead paper

Dijkstra Is NOT Greedy: A Global-to-Local Proof of Correctness

Oct 06, 2026

This paper challenges the conventional classification of Dijkstra’s algorithm as a greedy method, rectifying the widespread misconception that local choices inherently lead to global optimality. It proposes a three-step elimination method and a deterministic path-space contraction strategy, which are formally proven in conjunction with optimal substructure and boundary state reduction theory. These analyses reveal that the algorithm fundamentally constitutes a dynamic programming process deriving local solutions from the global shortest path. By establishing this paradigm-shifting "global-to-local" perspective, the study constructs a more concise correctness proof and clarifies the algorithm's essential nature. Ultimately, this work provides a novel theoretical interpretive framework for shortest-path problems in graph theory.

0 citationsRead paper

One-Step Generative Modeling via Training Dynamics Action

Sep 30, 2026

This study addresses the disconnect between theoretical objectives and practical deployment in existing one-step generative models, where training often overlooks the implementation costs of parameter sharing. To bridge this gap, we propose TDAction, a framework that formulates training dynamics as an optimal control problem for the first time, achieving cost-aware optimal transport through local parameter sharing. Specifically, we derive a closed-form Batch Tangent Action-to-Go value function and introduce soft terminal control, stochastic tangent probing, and low-rank approximation techniques to enhance optimization efficiency. By preserving the distributional evolution process while significantly reducing implementation overhead, our method achieves state-of-the-art performance on ImageNet 256×256, yielding an FID below 1.1 without requiring distillation.

0 citationsRead paper

Is Personalized Modality Weighting Actually Personalized? A Controlled Audit of Per-User Weighting Claims in Multimodal Recommenders

Aug 06, 2026

This work investigates whether user-specific modality weighting mechanisms widely adopted in multimodal recommendation systems genuinely capture individual user preferences. To this end, the authors propose an auditing framework comprising two metrics—real-GM and real-shuf—that evaluate personalization efficacy by comparing six personalized weighting methods against global weights and shuffled user-weight assignments, all under a unified collaborative filtering backbone. Experimental results across three short-video and one cross-domain e-commerce dataset reveal that performance gains from most methods stem primarily from increased model capacity rather than authentic user signals, with gating mechanisms often inducing spurious personalization due to their reliance on shared embeddings. Notably, global modality weights already achieve nearly all attainable gains, while personalized weighting shows no consistent improvement; the proposed audit framework effectively identifies architectures that truly encode user-specific patterns.

0 citationsRead paper

Task-Oriented Wave Processing with Stacked Intelligent Metasurfaces: Framework, Fusion, and Challenges

Jul 20, 2026

This work addresses the performance conflicts and resource contention arising from conventional task-agnostic channels in 6G scenarios characterized by deep convergence of heterogeneous services. To overcome these challenges, the paper proposes a novel physical-layer computing paradigm based on stacked intelligent metasurfaces (SIMs). This paradigm reconfigures the wireless environment into a programmable signal processor, establishing a unified mapping framework that directly links service requirements to wave-domain synthesis. By leveraging the deep computational architecture of SIMs, the approach enables intrinsic co-design of sensing, communication, and computing. Notably, it pioneers the use of SIMs for task-oriented wave manipulation, thereby transitioning multi-service systems from mere coexistence to true symbiosis. Numerical results demonstrate that the proposed method effectively mitigates resource conflicts and significantly enhances overall system performance, offering a foundational enabler for service-native 6G architectures.

0 citationsRead paper