computer vision

Designs, builds, and evaluates algorithms, models, and systems that automatically extract, represent, and interpret information from images and video to produce actionable outputs. This includes end-to-end pipelines for image/video acquisition and preprocessing, feature extraction and representation learning, object detection and recognition, semantic/instance segmentation, motion estimation and tracking, depth and 3D reconstruction, pose estimation, and assessment of performance and robustness.

computervision

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.4
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$219K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

To address the high cost and low efficiency of manual annotation in video object detection and segmentation, this paper proposes a lightweight, end-to-end automated annotation framework. Methodologically, it introduces the first deep integration of YOLOv8-based video tracking and SAM-based interactive segmentation, augmented by an adaptive frame-sampling strategy and a Gradio-powered visual interface, resulting in a modular, open-source, and extensible prototype system. Key contributions include: (i) a lightweight co-design of tracking and segmentation models; (ii) efficient, interactive annotation support for multi-object, long-duration videos; and (iii) rapid generation of annotated datasets enabling a closed-loop annotation–training pipeline. Experiments on multiple public video benchmarks demonstrate over 10× higher annotation throughput compared to manual labeling, strong inter-annotator consistency, and substantial performance gains for downstream detection and segmentation models trained on the generated data.

Automate video annotation for computer vision modelsDevelop efficient video tracking and segmentation toolReduce time and resources for labeled data generation

Deep Learning and Machine Learning - Object Detection and Semantic Segmentation: From Theory to Applications

Oct 21, 2024
JR
Jintao Ren
🏛️ Aarhus University | Indiana University | Kyoto University | AppCubic | Rutgers University | University of Wisconsin-Madison | Georgia Institute of Technology | National Taiwan Normal University | University of Hawaii | Xi'an Jiaotong-Liverpool University | Zhejiang University | Purdue University

To address performance limitations in object detection and semantic segmentation under complex scenarios—including occlusion, small objects, and cross-domain generalization—this paper proposes a novel multimodal detection paradigm synergizing large language models (LLMs). Methodologically, it systematically integrates CNNs, YOLOv5/v8, and DETR architectures into an LLM-augmented inference framework, augmented by scalable data pipelines, model pruning, and quantization, and evaluated via a multi-dimensional metric system based on mAP and mIoU. Key contributions include: (1) bridging the gap between traditional feature engineering and end-to-end deep learning; (2) introducing a dynamic context enhancement mechanism tailored for challenging environments; and (3) achieving state-of-the-art accuracy-efficiency trade-offs on COCO and ADE20K. The fully open-sourced, reproducible framework significantly improves model generalizability and robustness across diverse real-world conditions.

Bridging traditional methods with modern AI for large-scale detection tasksExploring object detection and semantic segmentation from theory to applicationsReviewing state-of-the-art deep learning architectures for computer vision

Object Learning and Robust 3D Reconstruction

Apr 22, 2025
SS
Sara Sabour
🏛️ University of Toronto

This work addresses motion-induced artifacts in unsupervised 3D reconstruction of dynamic scenes caused by moving foreground objects. To tackle this, we propose a novel method that jointly models unsupervised 2D object decomposition and 3D geometric consistency. Our approach introduces FlowCapsules—a flow-guided network for unsupervised foreground segmentation—and a transient-object-mask-driven robust optimization kernel that detects and excludes dynamic objects under multi-view geometric consistency constraints. This kernel subsequently guides weighted bundle adjustment and NeRF training. Crucially, our method is the first to jointly learn object-centric representations and scene-level geometry without any annotations or controlled capture conditions. Experiments demonstrate significant improvements in both SfM and NeRF reconstruction accuracy on casually captured dynamic scenes, effectively eliminating motion artifacts. The framework advances explicit object-aware 3D understanding in open-world vision applications.

3D dynamic object detection via geometric consistencyRobust 3D modeling with transient object masksUnsupervised 2D object segmentation using motion cues

Video Understanding by Design: How Datasets Shape Architectures and Insights

Sep 11, 2025
LW
Lei Wang
🏛️ Griffith University | Data61/CSIRO | Australian National University | University of New South Wales

Existing video understanding research overlooks how structural characteristics of datasets—such as motion complexity, temporal span, hierarchical composition, and multimodal richness—guide the evolution of model architectures. Method: We propose a dataset-centric analytical framework that systematically interprets mainstream architectures—including two-stream networks, 3D CNNs, RNNs, Transformers, and multimodal foundation models—as responses to dataset-imposed inductive biases. Our approach integrates literature review with architecture–bias–task alignment analysis, unifying inductive bias theory and multimodal learning paradigms. Contribution/Results: We establish, for the first time, a unified “dataset → inductive bias → model design” framework, revealing the intrinsic logic underlying architectural evolution. The framework yields principled, generalizable design guidelines for video understanding models that balance scalability and task adaptability, advancing both theoretical understanding and practical model development.

Analyzing how video datasets impose structural biases on model architecturesProviding guidance for aligning designs with dataset invariances and scalabilityReinterpreting model evolution as responses to dataset-driven inductive pressures

Automated Video Segmentation Machine Learning Pipeline

Jul 09, 2025
JM
Johannes Merz
🏛️ Image Engine Design Inc.

To address the inefficiency, time consumption, and high resource demands of manual mask generation in visual effects (VFX) production, this paper proposes a text-prompt-driven automated video segmentation pipeline. The method integrates text-guided instance segmentation for flexible object localization, fine-grained frame-wise semantic segmentation, and temporally consistent video object tracking, deployed via lightweight containerization to ensure seamless integration with artist workflows. Key contributions include: (i) the first end-to-end incorporation of text prompting into VFX-grade video segmentation; (ii) a multi-stage collaborative architecture ensuring inter-frame segmentation stability; and (iii) structured, editable outputs—including mask sequences, confidence maps, and trajectory metadata. Experiments demonstrate that the system reduces pre-compositing preparation time by 72% on average and decreases manual intervention by 89%, significantly enhancing automation and productivity across VFX pipelines.

Automates slow mask generation in VFX productionEnsures temporally consistent video segmentation masksReduces manual effort for preliminary composites

Latest Papers

What's happening recently
View more

This work addresses the challenge of achieving real-time performance, high accuracy, and energy efficiency in embedded vision systems operating on resource-constrained hardware. The authors propose an algorithm-hardware co-design methodology tailored for DSP/FPGA platforms, optimizing edge, corner, and blob detection operators through hardware-aware algorithmic refinements and quantization techniques. To further enhance throughput without compromising image quality, the approach incorporates inter-frame redundancy elimination and adaptive frame averaging strategies. Experimental results demonstrate that, compared to conventional solutions, the proposed method delivers significantly improved processing speed and energy efficiency, enabling scalable and highly effective real-time embedded vision across diverse applications such as automotive systems, surveillance, and robotics.

edge detectionembedded systemslatency

This work proposes SuperCam, a novel camera architecture designed to address the inefficiency of conventional cameras in resource-constrained settings, where they generate excessive redundant data that hinders downstream visual tasks. SuperCam uniquely integrates adaptive superpixel segmentation directly into the hardware pipeline, enabling online, lightweight data compression and preservation of critical visual information at the point of capture. By synergizing this design with edge computing, the system substantially reduces memory footprint and bandwidth requirements. Experimental results demonstrate that SuperCam consistently outperforms traditional approaches across multiple vision tasks—including semantic segmentation, object detection, and monocular depth estimation—thereby validating its feasibility and superiority for efficient perception under stringent resource limitations.

computer visionedge devicesmemory-limited

AfroBeats Dance Movement Analysis Using Computer Vision: A Proof-of-Concept Framework Combining YOLO and Segment Anything Model

Dec 03, 2025
KO
Kwaku Opoku-Ware
🏛️ University of Idaho | Kwame Nkrumah University of Science and Technology

This study addresses the challenge of automated AfroBeats dance motion analysis by proposing a marker-free, equipment-agnostic end-to-end video analytics framework. Methodologically, it integrates YOLOv8/v11 with the Segment Anything Model (SAM) to achieve high-precision dancer detection and pixel-level instance segmentation—overcoming the limitations of bounding-box-based approaches. Motion quantification is performed via trajectory tracking and inter-frame displacement analysis, enabling step counting, spatial coverage measurement, and rhythm consistency assessment. Evaluated on a 49-second real-world AfroBeats video, the system achieves 94% detection precision, 89% recall, and an 83% IoU for SAM-based segmentation. Quantitative analysis reveals statistically significant disparities between lead and supporting dancers in step count (+23%), motion intensity (+37%), and spatial occupancy (+42%). To our knowledge, this is the first work to apply SAM for fine-grained African dance motion analysis, establishing a scalable, high-fidelity visual computing paradigm for markerless dance evaluation.

Automated analysis of AfroBeats dance movements using computer visionEvaluating technical feasibility of YOLO and SAM integration for motion analysisTracking and quantifying dancer steps, spatial coverage, and rhythm consistency

This study addresses the limited generalization capability of existing motion-based AI-generated video detection methods, which rely heavily on inter-frame motion biases present in datasets rather than genuine forgery artifacts. The authors systematically evaluate four state-of-the-art motion-based detectors and, for the first time, explicitly demonstrate that these models exploit motion shortcuts through biased training and testing data. Through data rebalancing, spatial augmentations, and cross-dataset evaluations, they reveal the fragility of such approaches. Experiments show that when motion bias is eliminated in a newly curated dataset, all motion-based detectors degrade to random-chance performance, whereas frequency-domain detectors maintain consistently high accuracy, underscoring the superior robustness of frequency-based methods.

AI-generated video detectiondataset biasgeneralization

Hot Scholars

RH

Richard Han

Professor of Computing, Macquarie University
mobile computingdrone systemswireless sensor networkssecurity
FS

Federica Sarro

Professor, University College London
AI EngineeringSBSEAutomated Software EngineeringEmpirical Software Engineering
GD

Giordano d'Aloisio

Postdoctoral Researcher, Università degli Studi dell'Aquila
Software FairnessSustainabilitySoftware EngineeringEmpirical Software Engineering
SK

Shehryar Khattak

NASA Jet Propulsion Lab
RoboticsPerceptionComputer VisionSLAM
MZ

Minghui Zheng

J. Mike Walker '66 Department of Mechanical Engineering, Texas A&M University
RoboticsPlanningControlRobotic Disassembly