Institution profile

SenseTime

Industry researchasia · cn
Official website
Research library65linked papers
Opportunities0open roles
Selected work

Representative Papers

PageWeaver: KV-Guided Query Unions for Sparse Attention

Oct 08, 2026

This study addresses the low GPU utilization in dynamic sparse attention caused by small query support sets by proposing a page-affinity-based query assembly mechanism. It enables shared KV page loading through bounded search and ID-aware kernels, while an execution-level query union strategy eliminates tensor rearrangement and cross-page partial outputs to optimize non-local reuse. Computational efficiency is further enhanced by integrating FP8 KV caching, Tensor Core tile padding, and dual-CTA kernels. Evaluated on NVIDIA H200 GPUs, the proposed approach achieves a 1.7× speedup over the FlashInfer baseline and improves prefill throughput by 7.88%–14.36%.

0 citationsRead paper

Phase-aware video generation for physics-grounded dynamics and interactions

Oct 08, 2026

This study addresses the significant challenge of preserving phase-specific motion characteristics while modeling physically consistent interactions in solid-gas coupled dynamics video generation. To this end, we construct a large-scale trajectory corpus comprising 700,000 samples and propose a phase-aware dual-branch architecture that explicitly disentangles and collaboratively models the distinct dynamics of solid and gas phases. Furthermore, a spatiotemporal cross-attention mechanism is designed to precisely capture cross-phase physical coupling relationships. Extensive evaluations demonstrate that our approach significantly outperforms existing methods in terms of motion adherence, physical plausibility, and visual quality, thereby achieving high-fidelity, physics-compliant generation of solid-gas interaction videos.

0 citationsRead paper

Self-correction Optimization for Interleaved Multimodal Generation

Oct 07, 2026

This study addresses the challenges of visual subject inconsistency, poor temporal coherence, and insufficient physical plausibility in multimodal interleaved image-text generation. To this end, we propose a training-free self-correcting optimization framework. Building upon classifier-free guidance (CFG), our method introduces a complementary constraint mechanism that jointly governs novel event incorporation and state preservation, achieving self-correction by imposing minimal modifications to the guided updates without requiring additional training or computational overhead. Experimental results demonstrate that the proposed approach significantly enhances both the temporal coherence and physical realism of generated content across complex benchmarks. Furthermore, it successfully generalizes to long-horizon video generation tasks, such as robotic manipulation, highlighting its broad applicability and effectiveness.

0 citationsRead paper

RealtimeWAM: One-Step Asynchronous World Action Models

Oct 05, 2026

This work addresses the inference bottlenecks in world action models caused by multi-step denoising and serial expert execution. To overcome these limitations, it proposes a single-step generation and asynchronous inference framework. Methodologically, Teacher-Anchored Consistency Distillation (TACD) is introduced to enable single-step action generation while eliminating iterative error accumulation. Additionally, a Cross-Expert Wavefront Pipeline (CEWP) is designed to facilitate asynchronous inference by overlapping video and action module computations through block-level KV cache sharing. This hybrid architecture incurs less than 1% performance degradation on benchmarks such as LIBERO while achieving an approximately 25× inference speedup on H100 GPUs, substantially enhancing real-time control efficiency.

0 citationsRead paper

Think Before You Score: Thinking Reward Model for Visual Generation

Sep 29, 2026

This study addresses the limitation of existing visual reward models, which predominantly rely on implicit evaluation and lack adaptive criteria. We pioneer a "think-then-score" paradigm by introducing the Thinking Reward Model, which leverages the reasoning capabilities of large language models to first formulate instance-specific scoring rubrics before generating fine-grained, point-level rewards for guiding visual generation optimization. Furthermore, we propose the PD-GRPO algorithm to mitigate score polarization, effectively reconciling pairwise preference supervision with fine-grained scoring. By integrating reinforcement learning with fine-grained reward modeling, our approach achieves state-of-the-art performance among open-source models and substantially enhances the reinforcement training of diverse visual generation models.

0 citationsRead paper
Recent publications

Latest Papers

PageWeaver: KV-Guided Query Unions for Sparse Attention

Oct 08, 2026

This study addresses the low GPU utilization in dynamic sparse attention caused by small query support sets by proposing a page-affinity-based query assembly mechanism. It enables shared KV page loading through bounded search and ID-aware kernels, while an execution-level query union strategy eliminates tensor rearrangement and cross-page partial outputs to optimize non-local reuse. Computational efficiency is further enhanced by integrating FP8 KV caching, Tensor Core tile padding, and dual-CTA kernels. Evaluated on NVIDIA H200 GPUs, the proposed approach achieves a 1.7× speedup over the FlashInfer baseline and improves prefill throughput by 7.88%–14.36%.

0 citationsRead paper

Phase-aware video generation for physics-grounded dynamics and interactions

Oct 08, 2026

This study addresses the significant challenge of preserving phase-specific motion characteristics while modeling physically consistent interactions in solid-gas coupled dynamics video generation. To this end, we construct a large-scale trajectory corpus comprising 700,000 samples and propose a phase-aware dual-branch architecture that explicitly disentangles and collaboratively models the distinct dynamics of solid and gas phases. Furthermore, a spatiotemporal cross-attention mechanism is designed to precisely capture cross-phase physical coupling relationships. Extensive evaluations demonstrate that our approach significantly outperforms existing methods in terms of motion adherence, physical plausibility, and visual quality, thereby achieving high-fidelity, physics-compliant generation of solid-gas interaction videos.

0 citationsRead paper

Self-correction Optimization for Interleaved Multimodal Generation

Oct 07, 2026

This study addresses the challenges of visual subject inconsistency, poor temporal coherence, and insufficient physical plausibility in multimodal interleaved image-text generation. To this end, we propose a training-free self-correcting optimization framework. Building upon classifier-free guidance (CFG), our method introduces a complementary constraint mechanism that jointly governs novel event incorporation and state preservation, achieving self-correction by imposing minimal modifications to the guided updates without requiring additional training or computational overhead. Experimental results demonstrate that the proposed approach significantly enhances both the temporal coherence and physical realism of generated content across complex benchmarks. Furthermore, it successfully generalizes to long-horizon video generation tasks, such as robotic manipulation, highlighting its broad applicability and effectiveness.

0 citationsRead paper

RealtimeWAM: One-Step Asynchronous World Action Models

Oct 05, 2026

This work addresses the inference bottlenecks in world action models caused by multi-step denoising and serial expert execution. To overcome these limitations, it proposes a single-step generation and asynchronous inference framework. Methodologically, Teacher-Anchored Consistency Distillation (TACD) is introduced to enable single-step action generation while eliminating iterative error accumulation. Additionally, a Cross-Expert Wavefront Pipeline (CEWP) is designed to facilitate asynchronous inference by overlapping video and action module computations through block-level KV cache sharing. This hybrid architecture incurs less than 1% performance degradation on benchmarks such as LIBERO while achieving an approximately 25× inference speedup on H100 GPUs, substantially enhancing real-time control efficiency.

0 citationsRead paper

Think Before You Score: Thinking Reward Model for Visual Generation

Sep 29, 2026

This study addresses the limitation of existing visual reward models, which predominantly rely on implicit evaluation and lack adaptive criteria. We pioneer a "think-then-score" paradigm by introducing the Thinking Reward Model, which leverages the reasoning capabilities of large language models to first formulate instance-specific scoring rubrics before generating fine-grained, point-level rewards for guiding visual generation optimization. Furthermore, we propose the PD-GRPO algorithm to mitigate score polarization, effectively reconciling pairwise preference supervision with fine-grained scoring. By integrating reinforcement learning with fine-grained reward modeling, our approach achieves state-of-the-art performance among open-source models and substantially enhances the reinforcement training of diverse visual generation models.

0 citationsRead paper