Institution profile

Shanghai University of Finance and Economics

Academic institutionasia · cn
Official website
Research library346linked papers
Opportunities0open roles
Selected work

Representative Papers

MeViS: A Multi-Modal Dataset for Referring Motion Expression Video Segmentation

Aug 19, 2025IEEE Transactions on Pattern Analysis and Machine Intelligence

Existing video segmentation datasets emphasize static attribute descriptions, neglecting the critical role of motion in video understanding. To address this, we introduce MeViS—the first multimodal video segmentation dataset explicitly guided by motion expression—comprising 33K human-annotated text and audio motion descriptions across 2,006 complex scenes and 8,171 objects, supporting four tasks: Referring Video Object Segmentation (RVOS), Audio-Visual Object Segmentation (AVOS), Referring Multi-Object Tracking (RMOT), and Referring Motion Expression Grounding (RMEG). MeViS pioneers motion semantics as the core referential cue, breaking the static-dominant paradigm. We further propose LMPM++, a model integrating multimodal aligned annotation, motion-aware modeling, and joint audio-visual-linguistic representation, achieving new state-of-the-art performance on RVOS, AVOS, and RMOT. Comprehensive evaluation of 15 mainstream methods reveals systematic motion reasoning bottlenecks; leveraging MeViS significantly improves segmentation and tracking accuracy, advancing motion-centric video understanding.

22 citationsRead paper

AlignDiff: Exploiting Model-Intrinsic Information for Better Preference Data Selection

Sep 05, 2026

Aligning large language models with human preferences remains a challenge, primarily due to the critical role of preference data quality in effective alignment. Existing datasets are frequently plagued by inherent noise and distribution shifts, which inherently limit model performance. To bridge this gap, we propose AlignDiff, a preference data filtering framework driven by intrinsic model signals. AlignDiff first identifies samples with clear preferences using both positive and inverse signals, then prioritizes the more challenging samples based on the average negative log-likelihood gap, encouraging the model to learn richer information from them. AlignDiff is evaluated on two widely used model families (LLaMA and Qwen) and three benchmarks widely adopted in the alignment community (AlpacaEval 2.0, Arena-Hard, and MT-Bench). Across all settings, it consistently outperforms seven strong baselines. We conduct comprehensive ablation studies to validate the effectiveness of AlignDiff, and further show that difficulty-based curriculum learning improves model performance.

1 citationsRead paper
Recent publications

Latest Papers