Score
Designs, implements, and analyzes methods for detecting, simulating, augmenting, and reasoning about occlusions in visual data (images and video), including generating occlusion augmentations, simulating occlusion events, detecting occluded regions and masks, projecting scene representations to camera views, and propagating labels under partial visibility. Builds and evaluates recovery and compositing techniques — mask-guided inpainting, type-aware occlusion recovery, pixel-replacement or filtering fallbacks, video compositing and vandalism repair — and performs occlusion robustness testing and strategy selection to preserve downstream perception performance.
This work addresses the insufficient robustness of object recognition models under local visual occlusion. We propose a dual-path occlusion-resilient method based on frozen diffusion models: occlusion is formulated as an image completion task, and context-aware diffusion features are extracted from intermediate layers of Stable Diffusion to (i) inpaint occluded regions at the input level and (ii) augment discriminative representations via feature-level fusion. Key contributions include: (1) the first integration of pretrained diffusion model intermediate features for occlusion-robust classification; (2) a novel dual-path enhancement paradigm synergizing input restoration and feature fusion; and (3) the construction of the first benchmark dataset targeting realistic occlusion scenarios. Experiments demonstrate substantial improvements—e.g., significant Top-1 accuracy gains for ResNet and ViT under synthetic ImageNet occlusions—and an average 12.3% performance boost over baselines on real-world occlusion benchmarks.
Severe occlusion leads to significant loss and corruption of target information, substantially degrading classification performance. This work proposes an occlusion-agnostic and severity-adaptive classification method that enhances robustness during training through multi-level random masking and, at test time, dynamically masks disruptive regions based on visual anomaly detection to estimate occlusion severity and adaptively select the optimal model. To the best of our knowledge, this is the first approach to jointly achieve occlusion-type invariance and severity-aware optimization. On occluded images, the proposed method improves AUC_occ by 18.5% over standard training and by 23.7% compared to fine-tuning without occlusion.
Existing 3D layout-guided image generation methods struggle to accurately model occlusion relationships among objects, often resulting in inconsistent geometry and scale for occluded entities in synthesized scenes. To address this limitation, this work proposes SeeThrough3D, the first approach to explicitly model occlusion relations in text-to-image generation. By introducing an Occlusion-aware 3D Scene Representation (OSCR), objects are represented as semi-transparent 3D bounding boxes, enabling explicit reasoning about occluded regions. The method integrates a flow-based diffusion model with a mask-based self-attention mechanism to achieve precise 3D layout control while preserving occlusion consistency. SeeThrough3D supports multi-object binding and viewpoint-consistent occlusion synthesis, generalizes well to unseen object categories, and significantly enhances both the realism and layout fidelity of complex scene generation.
To address the performance bottleneck of interactive occlusion boundary estimation (IOBE) caused by the scarcity of real-world occlusion boundary (OB) annotations, this paper proposes DNMMSI—the first interactive IOBE method supporting multi-stroke interventions—and introduces two novel benchmarks: OB-FUTURE, a geometrically precise synthetic benchmark built upon our self-developed differentiable tool Mesh2OB, and OB-LabName, a high-fidelity real-world benchmark comprising 120 images with pixel-level ground truth. Key contributions include: (i) establishing the first formal interactive OB estimation paradigm; (ii) introducing the first end-to-end training framework leveraging geometry-differentiable synthetic data, eliminating the need for domain adaptation; and (iii) releasing the first high-fidelity real OB benchmark alongside a fully open-sourced toolchain. Experiments demonstrate that DNMMSI, trained solely on synthetic data, significantly outperforms state-of-the-art fully automatic methods, while OB-LabName achieves superior annotation accuracy compared to existing benchmarks.
This work addresses the heavy reliance on manual intervention—such as hand-crafted mask resizing and occlusion inpainting—in existing product catalog image synthesis methods. The authors propose a model-agnostic, end-to-end automated framework that generates high-quality composites from only a product image and a background, leveraging a novel dimension-aware masking algorithm and an occlusion-aware hybrid inpainting mechanism. Key contributions include the dimension-aware masking strategy, the occlusion-aware inpainting approach, the CatalogStitch-Eval benchmark comprising 58 complex real-world scenes, and an accompanying visualization toolkit. Experiments demonstrate consistent and significant improvements in both synthesis quality and efficiency across three representative models—ObjectStitch, OmniPaint, and InsertAnything—without requiring any post-processing.
Real-time hardware-accelerated ray tracing suffers from high noise and computational cost in ambient occlusion (AO) and soft shadow estimation due to limited samples per pixel. This work proposes an occluder point reuse framework that unifies AO and area-light shadow estimation as integrals over the occluder domain. By reusing ray samples within neighborhoods of first-hit occluder points and combining unbiased and biased estimators through multiple importance sampling, the method achieves efficient sample reuse guided by occluder consistency rather than traditional visibility consistency, better aligning with scene geometry. Experiments demonstrate that, at comparable computational cost, the proposed approach outperforms non-reuse baselines in both AO and soft shadow quality, and further surpasses existing light-sample reuse techniques in shadow fidelity.
Existing video object removal methods struggle to preserve scene consistency after eliminating objects involved in complex physical interactions—such as collisions—often resulting in implausible or distorted outputs. This work addresses this limitation by introducing high-order causal reasoning and explicit physical consistency into the task for the first time. The authors construct a synthetic dataset featuring counterfactual object removal using Kubric and HUMOTO, and leverage a vision-language model to identify regions affected by the removal. These regions then guide a video diffusion model to generate edits that adhere to physical laws. Experimental results demonstrate that the proposed approach significantly outperforms existing methods on both synthetic and real-world data, effectively maintaining dynamic scene coherence after object removal.