Institution profile

Canva

Industry researchaustralasia · au
Official website
Research library16linked papers
Opportunities0open roles
Selected work

Representative Papers

FlashGaze: Training-Free Multi-Scale Patch Pruning For Efficient Video Understanding

Oct 03, 2026

This study addresses the prohibitive computational overhead incurred by spatiotemporal redundancy when multimodal large language models process long videos. We propose a training-free video understanding acceleration method that leverages pixel-space differences as a proxy metric, eliminating the need for auxiliary networks. Prior to Vision Transformer (ViT) encoding, our approach employs quadtree dynamic programming to jointly optimize the dropping, merging, and retention strategies for multi-scale patches, thereby efficiently pruning redundant visual tokens. Experimental results demonstrate that the proposed method preserves 98% of the original accuracy while achieving a 5.4× speedup in ViT encoding and a 17× acceleration during prefilling. Furthermore, GPU memory consumption is reduced by 1.8×. These improvements significantly enhance the efficiency of long video understanding without compromising model performance.

0 citationsRead paper

Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE

Sep 29, 2026

This study addresses the semantic routing fragmentation in traditional Mixture-of-Experts (MoE) for video diffusion models caused by uniformity constraints. We propose SplitMoE, a novel architecture that introduces a role-disentangled sparsity mechanism to explicitly separate semantic and general experts, thereby breaking the uniformity trap. Furthermore, we incorporate prototype-guided routing with push-pull regularization to effectively overcome routing imbalances induced by spatiotemporal redundancy in visual data, achieving precise semantic clustering. Experimental results demonstrate that, under an equivalent parameter budget, our method significantly improves convergence speed, routing coherence, and video generation quality. This work provides a new pathway toward modality-aware efficient scaling for video diffusion models.

0 citationsRead paper
Recent publications

Latest Papers

FlashGaze: Training-Free Multi-Scale Patch Pruning For Efficient Video Understanding

Oct 03, 2026

This study addresses the prohibitive computational overhead incurred by spatiotemporal redundancy when multimodal large language models process long videos. We propose a training-free video understanding acceleration method that leverages pixel-space differences as a proxy metric, eliminating the need for auxiliary networks. Prior to Vision Transformer (ViT) encoding, our approach employs quadtree dynamic programming to jointly optimize the dropping, merging, and retention strategies for multi-scale patches, thereby efficiently pruning redundant visual tokens. Experimental results demonstrate that the proposed method preserves 98% of the original accuracy while achieving a 5.4× speedup in ViT encoding and a 17× acceleration during prefilling. Furthermore, GPU memory consumption is reduced by 1.8×. These improvements significantly enhance the efficiency of long video understanding without compromising model performance.

0 citationsRead paper

Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE

Sep 29, 2026

This study addresses the semantic routing fragmentation in traditional Mixture-of-Experts (MoE) for video diffusion models caused by uniformity constraints. We propose SplitMoE, a novel architecture that introduces a role-disentangled sparsity mechanism to explicitly separate semantic and general experts, thereby breaking the uniformity trap. Furthermore, we incorporate prototype-guided routing with push-pull regularization to effectively overcome routing imbalances induced by spatiotemporal redundancy in visual data, achieving precise semantic clustering. Experimental results demonstrate that, under an equivalent parameter budget, our method significantly improves convergence speed, routing coherence, and video generation quality. This work provides a new pathway toward modality-aware efficient scaling for video diffusion models.

0 citationsRead paper