Score
Designs algorithms and processing pipelines that convert episode-level instructions or sparse temporal annotations into dense per-frame supervisory signals by decomposing episodes into sub-instructions, aligning each video frame with matched action targets, and producing per-frame training targets such as action labels, trajectories, or flow-field supervision. Builds and evaluates framewise alignment and local supervision-conversion modules that generate the dense, frame-level supervision used to train or analyze temporally detailed models.
This work addresses the scarcity of high-quality procedural annotations in long videos caused by ASR noise and audio-visual temporal misalignment. It proposes the first training-free automated pipeline that leverages shot segmentation and cross-modal alignment filtering, followed by structured, temporally aligned fine-grained procedural step generation using multimodal large language models such as Qwen2.5-VL and DeepSeek-R1. This approach enables large-scale, training-free dense procedural video annotation, resulting in DenseStep2M—a novel dataset comprising 100,000 videos and 2 million steps. The method significantly advances performance on dense captioning, step localization, and cross-modal retrieval tasks, while demonstrating strong zero-shot generalization capabilities.
为了解决视频语言建模中长时间动态表示的问题,本文通过引入具有细粒度时间标注的Kairos数据集来支持更精细和长范围的视频内容理解。
To address the global, ambiguous, and spatially imprecise nature of text-based prompting for image-to-video generation, this work introduces *In-Video Instruction*—a novel paradigm that embeds structured visual signals (e.g., overlaid text, arrows, motion trajectories) directly into input frames, enabling pixel-level spatial alignment and unambiguous, fine-grained control. Built upon state-of-the-art video diffusion models—including Veo 3.1, Kling 2.5, and Wan 2.2—we design a lightweight visual instruction encoder and conditional injection mechanism, allowing models to interpret and execute spatially grounded, multi-object, multi-action instructions without fine-tuning. Experiments demonstrate substantial improvements in instruction adherence and action localization accuracy, particularly in complex multi-object scenarios. This work provides the first systematic validation that off-the-shelf video generation models can reliably parse embedded visual instructions, establishing a scalable, high-precision pathway for controllable video synthesis.
To address the high annotation cost of fully supervised video action detection, this paper proposes an action-agnostic frame-level weak supervision paradigm (AAPL), requiring only sparse, unsupervised keyframe annotations—without exhaustive video scanning or instance-level localization—thereby drastically reducing labeling effort. Methodologically, we design an end-to-end temporal detection model coupled with a weakly supervised learning strategy to accurately localize action segments from sparse, non-instance-aligned frame labels. A contrastive pseudo-label distillation mechanism is further introduced to enhance the temporal convolutional network’s capacity for modeling action boundaries. Evaluated on five standard benchmarks—including THUMOS’14 and ActivityNet 1.3—our approach achieves state-of-the-art or competitive performance using significantly fewer annotations than video-level or point-level supervision, establishing, for the first time, a new paradigm that simultaneously delivers high detection accuracy and low annotation cost.
This work addresses the challenges of temporally inconsistent predictions—such as flickering—in human-centric dense video tasks under motion, occlusion, and illumination changes, compounded by the scarcity of multi-task paired video supervision. To this end, we propose a scalable, photorealistic synthetic human video generation method that, for the first time, provides both frame-level and sequence-level pixel-wise annotations, including depth, surface normals, and masks. Leveraging this data, we develop a unified Vision Transformer (ViT)-based dense prediction architecture that integrates CSE human geometric priors with a lightweight channel reweighting module. Our approach employs a two-stage training strategy—static pretraining followed by dynamic sequence fine-tuning—to jointly optimize spatial and temporal consistency. The method achieves state-of-the-art performance on THuman2.1 and Hi4D benchmarks and demonstrates strong generalization to in-the-wild real-world videos.
为解决视频生成中的运动不一致问题,提出MotionSpec方法,通过频谱轨迹一致性(STC)和局部流一致性(LFC)增强视频的运动连贯性和真实性。
Existing knowledge distillation methods assume a constant supervisory value for each sample, overlooking the dynamic evolution of student capabilities and thereby inducing data redundancy and training inefficiency. To address this limitation, this work proposes a student-curriculum coupled framework that innovatively decouples supervision credibility from necessity. By integrating online policy distillation, anchor-frontier curricula, and closed-loop feedback control, the framework enables the student to dynamically activate or suspend specific supervision subsets on demand according to its current task-specific deficiencies. Experimental results demonstrate that the proposed method achieves an average recall improvement of 5.1% across three benchmarks while reducing the number of training samples by 60% and decreasing computational time by 50.4%.
研究通过构建Video-IFBench基准,评估多模态大语言模型在视频理解场景中遵循指令的能力,使用半自动数据构建管道生成1.5K样本,揭示了当前模型在处理复杂指令上的挑战。
本文提出VideoX-Qwen框架,通过构建大规模数据集和统一训练模型解决基于指令的视频编辑问题,提高了编辑质量和内容保真度。
This work addresses a critical limitation in existing chart-to-code generation methods, which rely on reference code containing unobservable latent variables for supervision, often leading to model hallucination and over-specification. The study systematically identifies this issue and introduces an observation-aligned supervision framework that restricts training objectives to quantities directly inferable from chart images—such as boxplot statistics, pie chart proportions, and histogram bin weights—ensuring alignment between supervision signals and visual observations. By integrating chart understanding from vision-language models with supervised fine-tuning and data rewriting techniques, the proposed approach significantly improves both the accuracy of observable attribute recovery and code executability on ChartMimic and ChartX benchmarks, demonstrating the pivotal role of observation-aligned supervision in enhancing model performance.