Score
Designs, builds, and evaluates algorithms, models, and systems that automatically extract, represent, and interpret information from images and video to produce actionable outputs. This includes end-to-end pipelines for image/video acquisition and preprocessing, feature extraction and representation learning, object detection and recognition, semantic/instance segmentation, motion estimation and tracking, depth and 3D reconstruction, pose estimation, and assessment of performance and robustness.
To address the high cost and low efficiency of manual annotation in video object detection and segmentation, this paper proposes a lightweight, end-to-end automated annotation framework. Methodologically, it introduces the first deep integration of YOLOv8-based video tracking and SAM-based interactive segmentation, augmented by an adaptive frame-sampling strategy and a Gradio-powered visual interface, resulting in a modular, open-source, and extensible prototype system. Key contributions include: (i) a lightweight co-design of tracking and segmentation models; (ii) efficient, interactive annotation support for multi-object, long-duration videos; and (iii) rapid generation of annotated datasets enabling a closed-loop annotation–training pipeline. Experiments on multiple public video benchmarks demonstrate over 10× higher annotation throughput compared to manual labeling, strong inter-annotator consistency, and substantial performance gains for downstream detection and segmentation models trained on the generated data.
To address performance limitations in object detection and semantic segmentation under complex scenarios—including occlusion, small objects, and cross-domain generalization—this paper proposes a novel multimodal detection paradigm synergizing large language models (LLMs). Methodologically, it systematically integrates CNNs, YOLOv5/v8, and DETR architectures into an LLM-augmented inference framework, augmented by scalable data pipelines, model pruning, and quantization, and evaluated via a multi-dimensional metric system based on mAP and mIoU. Key contributions include: (1) bridging the gap between traditional feature engineering and end-to-end deep learning; (2) introducing a dynamic context enhancement mechanism tailored for challenging environments; and (3) achieving state-of-the-art accuracy-efficiency trade-offs on COCO and ADE20K. The fully open-sourced, reproducible framework significantly improves model generalizability and robustness across diverse real-world conditions.
This work addresses motion-induced artifacts in unsupervised 3D reconstruction of dynamic scenes caused by moving foreground objects. To tackle this, we propose a novel method that jointly models unsupervised 2D object decomposition and 3D geometric consistency. Our approach introduces FlowCapsules—a flow-guided network for unsupervised foreground segmentation—and a transient-object-mask-driven robust optimization kernel that detects and excludes dynamic objects under multi-view geometric consistency constraints. This kernel subsequently guides weighted bundle adjustment and NeRF training. Crucially, our method is the first to jointly learn object-centric representations and scene-level geometry without any annotations or controlled capture conditions. Experiments demonstrate significant improvements in both SfM and NeRF reconstruction accuracy on casually captured dynamic scenes, effectively eliminating motion artifacts. The framework advances explicit object-aware 3D understanding in open-world vision applications.
Existing video understanding research overlooks how structural characteristics of datasets—such as motion complexity, temporal span, hierarchical composition, and multimodal richness—guide the evolution of model architectures. Method: We propose a dataset-centric analytical framework that systematically interprets mainstream architectures—including two-stream networks, 3D CNNs, RNNs, Transformers, and multimodal foundation models—as responses to dataset-imposed inductive biases. Our approach integrates literature review with architecture–bias–task alignment analysis, unifying inductive bias theory and multimodal learning paradigms. Contribution/Results: We establish, for the first time, a unified “dataset → inductive bias → model design” framework, revealing the intrinsic logic underlying architectural evolution. The framework yields principled, generalizable design guidelines for video understanding models that balance scalability and task adaptability, advancing both theoretical understanding and practical model development.
To address the inefficiency, time consumption, and high resource demands of manual mask generation in visual effects (VFX) production, this paper proposes a text-prompt-driven automated video segmentation pipeline. The method integrates text-guided instance segmentation for flexible object localization, fine-grained frame-wise semantic segmentation, and temporally consistent video object tracking, deployed via lightweight containerization to ensure seamless integration with artist workflows. Key contributions include: (i) the first end-to-end incorporation of text prompting into VFX-grade video segmentation; (ii) a multi-stage collaborative architecture ensuring inter-frame segmentation stability; and (iii) structured, editable outputs—including mask sequences, confidence maps, and trajectory metadata. Experiments demonstrate that the system reduces pre-compositing preparation time by 72% on average and decreases manual intervention by 89%, significantly enhancing automation and productivity across VFX pipelines.
This work addresses the challenge of achieving real-time performance, high accuracy, and energy efficiency in embedded vision systems operating on resource-constrained hardware. The authors propose an algorithm-hardware co-design methodology tailored for DSP/FPGA platforms, optimizing edge, corner, and blob detection operators through hardware-aware algorithmic refinements and quantization techniques. To further enhance throughput without compromising image quality, the approach incorporates inter-frame redundancy elimination and adaptive frame averaging strategies. Experimental results demonstrate that, compared to conventional solutions, the proposed method delivers significantly improved processing speed and energy efficiency, enabling scalable and highly effective real-time embedded vision across diverse applications such as automotive systems, surveillance, and robotics.
This work proposes SuperCam, a novel camera architecture designed to address the inefficiency of conventional cameras in resource-constrained settings, where they generate excessive redundant data that hinders downstream visual tasks. SuperCam uniquely integrates adaptive superpixel segmentation directly into the hardware pipeline, enabling online, lightweight data compression and preservation of critical visual information at the point of capture. By synergizing this design with edge computing, the system substantially reduces memory footprint and bandwidth requirements. Experimental results demonstrate that SuperCam consistently outperforms traditional approaches across multiple vision tasks—including semantic segmentation, object detection, and monocular depth estimation—thereby validating its feasibility and superiority for efficient perception under stringent resource limitations.
This study addresses the challenge of automated AfroBeats dance motion analysis by proposing a marker-free, equipment-agnostic end-to-end video analytics framework. Methodologically, it integrates YOLOv8/v11 with the Segment Anything Model (SAM) to achieve high-precision dancer detection and pixel-level instance segmentation—overcoming the limitations of bounding-box-based approaches. Motion quantification is performed via trajectory tracking and inter-frame displacement analysis, enabling step counting, spatial coverage measurement, and rhythm consistency assessment. Evaluated on a 49-second real-world AfroBeats video, the system achieves 94% detection precision, 89% recall, and an 83% IoU for SAM-based segmentation. Quantitative analysis reveals statistically significant disparities between lead and supporting dancers in step count (+23%), motion intensity (+37%), and spatial occupancy (+42%). To our knowledge, this is the first work to apply SAM for fine-grained African dance motion analysis, establishing a scalable, high-fidelity visual computing paradigm for markerless dance evaluation.
This study addresses the limited generalization capability of existing motion-based AI-generated video detection methods, which rely heavily on inter-frame motion biases present in datasets rather than genuine forgery artifacts. The authors systematically evaluate four state-of-the-art motion-based detectors and, for the first time, explicitly demonstrate that these models exploit motion shortcuts through biased training and testing data. Through data rebalancing, spatial augmentations, and cross-dataset evaluations, they reveal the fragility of such approaches. Experiments show that when motion bias is eliminated in a newly curated dataset, all motion-based detectors degrade to random-chance performance, whereas frequency-domain detectors maintain consistently high accuracy, underscoring the superior robustness of frequency-based methods.