build computer vision models

Designs, implements, and evaluates models, algorithms, and end-to-end pipelines that process visual data (images and video) to detect and localize objects, perform semantic and instance segmentation, and track multiple objects over time. Builds and analyzes methods for image processing, depth and pose estimation, visual grounding, segmentation modeling/analysis, scene understanding, and related computer vision techniques and algorithms.

buildcomputervisionmodels

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.11
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$220K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Deep Learning and Machine Learning - Object Detection and Semantic Segmentation: From Theory to Applications

Oct 21, 2024
JR
Jintao Ren
🏛️ Aarhus University | Indiana University | Kyoto University | AppCubic | Rutgers University | University of Wisconsin-Madison | Georgia Institute of Technology | National Taiwan Normal University | University of Hawaii | Xi'an Jiaotong-Liverpool University | Zhejiang University | Purdue University

To address performance limitations in object detection and semantic segmentation under complex scenarios—including occlusion, small objects, and cross-domain generalization—this paper proposes a novel multimodal detection paradigm synergizing large language models (LLMs). Methodologically, it systematically integrates CNNs, YOLOv5/v8, and DETR architectures into an LLM-augmented inference framework, augmented by scalable data pipelines, model pruning, and quantization, and evaluated via a multi-dimensional metric system based on mAP and mIoU. Key contributions include: (1) bridging the gap between traditional feature engineering and end-to-end deep learning; (2) introducing a dynamic context enhancement mechanism tailored for challenging environments; and (3) achieving state-of-the-art accuracy-efficiency trade-offs on COCO and ADE20K. The fully open-sourced, reproducible framework significantly improves model generalizability and robustness across diverse real-world conditions.

Bridging traditional methods with modern AI for large-scale detection tasksExploring object detection and semantic segmentation from theory to applicationsReviewing state-of-the-art deep learning architectures for computer vision

This work addresses the challenge of achieving real-time performance, high accuracy, and energy efficiency in embedded vision systems operating on resource-constrained hardware. The authors propose an algorithm-hardware co-design methodology tailored for DSP/FPGA platforms, optimizing edge, corner, and blob detection operators through hardware-aware algorithmic refinements and quantization techniques. To further enhance throughput without compromising image quality, the approach incorporates inter-frame redundancy elimination and adaptive frame averaging strategies. Experimental results demonstrate that, compared to conventional solutions, the proposed method delivers significantly improved processing speed and energy efficiency, enabling scalable and highly effective real-time embedded vision across diverse applications such as automotive systems, surveillance, and robotics.

edge detectionembedded systemslatency

WS$^2$: Weakly Supervised Segmentation using Before-After Supervision in Waste Sorting

Sep 08, 2025
AM
Andrea Marelli
🏛️ Politecnico di Milano | EURECOM

To address the challenges of heavy reliance on dense pixel-level annotations and poor adaptability to diverse industrial conditions in scrap sorting—particularly for foreign object detection and segmentation—this paper proposes a novel weakly supervised learning paradigm termed “before-after supervision,” leveraging image differences before and after human intervention. Methodologically, we design an end-to-end semantic segmentation framework integrating differential feature modeling and multi-view consistency constraints, enabling unified evaluation of diverse weak supervision strategies. Key contributions include: (1) the first formulation of operator removal actions as natural weak supervision signals; (2) the release of WS², the first multi-view, high-resolution weakly supervised dataset specifically designed for industrial scrap sorting; and (3) empirical validation on WS² demonstrating that multiple weakly supervised methods achieve over 90% of fully supervised performance, substantially reducing annotation costs while maintaining practical deployability.

Automating visual recognition of unwanted items in waste sortingLeveraging before-after image differences for operator removal actionsReducing labeling efforts with weakly supervised segmentation

To address insufficient geometric-semantic joint understanding of robots in unstructured environments, this paper proposes an end-to-end modular RGB-D understanding pipeline. Our method introduces a novel hybrid mask generation mechanism integrating SAM2 with a lightweight semantic classifier for pixel-level semantic segmentation and instance awareness; incorporates ReID-enhanced cross-frame human tracking and semantic-weighted TSDF point cloud fusion to ensure geometric fidelity and semantic consistency; and outputs structured scene representations in USD format. Evaluated on ADE20K, our approach achieves 47.0% mIoU—surpassing SegFormer and OneFormer in boundary accuracy—and attains a reconstruction error of only 25.3 mm. It runs 1.81× faster than prior methods and has been validated for deployment feasibility on real-world Kinect data.

Enables efficient human tracking and point cloud fusionImproves semantic segmentation accuracy using hybrid SAM2 modelIntegrates semantic segmentation with geometric reconstruction for robotics

Object Learning and Robust 3D Reconstruction

Apr 22, 2025
SS
Sara Sabour
🏛️ University of Toronto

This work addresses motion-induced artifacts in unsupervised 3D reconstruction of dynamic scenes caused by moving foreground objects. To tackle this, we propose a novel method that jointly models unsupervised 2D object decomposition and 3D geometric consistency. Our approach introduces FlowCapsules—a flow-guided network for unsupervised foreground segmentation—and a transient-object-mask-driven robust optimization kernel that detects and excludes dynamic objects under multi-view geometric consistency constraints. This kernel subsequently guides weighted bundle adjustment and NeRF training. Crucially, our method is the first to jointly learn object-centric representations and scene-level geometry without any annotations or controlled capture conditions. Experiments demonstrate significant improvements in both SfM and NeRF reconstruction accuracy on casually captured dynamic scenes, effectively eliminating motion artifacts. The framework advances explicit object-aware 3D understanding in open-world vision applications.

3D dynamic object detection via geometric consistencyRobust 3D modeling with transient object masksUnsupervised 2D object segmentation using motion cues

Latest Papers

What's happening recently
View more

This work proposes a training-free object detection method tailored for scenarios involving minor data variations where model training and annotation are impractical, such as GUI automation testing. By leveraging a segmentation foundation model—e.g., SAM—to generate image segments and integrating classical feature engineering for object classification, the approach rapidly adapts to new targets or interface changes without any training or labeled data. Evaluated on an in-vehicle navigation icon detection task, the method achieves performance comparable to learning-based detectors like YOLO, while entirely eliminating the need for model training. This significantly reduces deployment time and cost, demonstrating strong practical utility through its efficiency and adaptability.

foundation modelGUI testingicon detection

This work proposes SuperCam, a novel camera architecture designed to address the inefficiency of conventional cameras in resource-constrained settings, where they generate excessive redundant data that hinders downstream visual tasks. SuperCam uniquely integrates adaptive superpixel segmentation directly into the hardware pipeline, enabling online, lightweight data compression and preservation of critical visual information at the point of capture. By synergizing this design with edge computing, the system substantially reduces memory footprint and bandwidth requirements. Experimental results demonstrate that SuperCam consistently outperforms traditional approaches across multiple vision tasks—including semantic segmentation, object detection, and monocular depth estimation—thereby validating its feasibility and superiority for efficient perception under stringent resource limitations.

computer visionedge devicesmemory-limited

This work addresses the challenges of high early-stage uncertainty in manufacturing monitoring system development—leading to redundant modeling and substantial training costs—and the limited transferability of filtering pipelines in cross-domain image segmentation tasks. To tackle these issues, the authors propose a problem-centric design paradigm that constructs an abstract system model to continuously accumulate and retrieve historical segmentation tasks along with their associated filtering pipelines, enabling solution reuse and incremental optimization. The approach integrates similarity-based problem retrieval, abstract modeling, pipeline reuse, and a retrieval-augmented evolutionary learning mechanism. Experimental results demonstrate that the method significantly reduces training costs and late-stage revision risks, provides the first systematic validation of filtering pipeline transferability across similar segmentation tasks, and achieves a favorable balance among complexity, technical requirements, and reliability under lightweight model constraints.

cross-domain transferabilityfilter pipelinesmonitoring systems

This work addresses the high cost of deploying traditional warehouse vision systems, which typically require repeated data collection, annotation, and retraining for each new environment. To overcome this limitation, the authors propose a generalizable deployment framework that eliminates the need for on-site retraining. By optimizing camera placement, implementing strategic image triggering mechanisms, selecting robust foundation models, and employing effective ensemble strategies, the system is trained once in the lab using only laboratory-collected data. The study demonstrates, for the first time, that such a model can successfully generalize across diverse real-world warehouse settings. Specifically, in the task of detecting fork anomalies in vertical material handling systems, deployment is reduced to simply installing cameras, capturing images, and directly applying the pre-trained model—bypassing costly annotation and retraining cycles and significantly lowering real-world deployment costs.

anomaly detectioncomputer visionmodel generalization

This study addresses the challenge of accurately segmenting irregular and densely packed components in electronic waste by presenting the first systematic comparison between the general-purpose foundation model SAM2 and the lightweight task-specific model YOLOv8. Leveraging a newly curated dataset of 1,456 high-quality annotated images and diverse data augmentation strategies, experiments demonstrate that YOLOv8 significantly outperforms SAM2, achieving an mAP50 of 98.8% and an mAP50–95 of 85%, with notably more precise boundary delineation. Although SAM2 offers greater structural flexibility, it suffers from mask overlap and contour inconsistency issues. The findings underscore that off-the-shelf vision foundation models require targeted adaptation to suit industrial recycling applications. The authors publicly release the dataset and benchmark framework to support further research in this domain.

e-waste disassemblyirregular component segmentationmaterial recovery

Hot Scholars

GZ

Guangtao Zhai

Professor, IEEE Fellow, Shanghai Jiao Tong University
Multimedia Signal ProcessingVisual Quality AssessmentQoEAI Evaluation
YW

Yunchao Wei

Professor, Beijing Jiaotong University, UTS, UIUC, NUS
Computer VisionMachine Learning
GL

Guosheng Lin

Nanyang Technological University
Computer VisionMachine Learning
HH

Haoyang Huang

JD Explore Academy (present) | StepFun | Microsoft Research
Multimodal & Multilingual Foundation Model
WL

Wenbo Li

The Chinese University of Hong Kong
Computer VisionDeep Learning