Score
Design and implement methods that detect and localize anomalies in three-dimensional volumetric data without task-specific labeled training (zero-shot / ZSAD), producing lightweight, efficient models or pipelines that can run batch-wise across samples; such work includes adapting or leveraging frozen two-dimensional foundation models and techniques to transfer 2D priors into 3D anomaly localization.
Existing zero-shot anomaly detection methods for 3D brain MRI are limited by their reliance on slice-level features, which neglect the intrinsic three-dimensional spatial structure among voxels. This work proposes a training-free, prompt-free, and supervision-free zero-shot anomaly detection framework that leverages multi-axis 2D foundation models to extract features from orthogonal slices, aggregates them into 3D local voxel tokens enriched with cubic spatial context, and directly applies distance-based batch anomaly scoring. To the best of our knowledge, this is the first approach to achieve fully training-free zero-shot anomaly detection on 3D brain MRI volumes. The method efficiently produces compact 3D representations on standard GPUs, successfully extending the zero-shot capabilities of 2D foundation models to full 3D volumes while maintaining both simplicity and robustness.
This work addresses the scarcity of annotated data in 3D medical imaging and the limited applicability of existing 2D zero-shot anomaly detection methods by proposing CS3F, a training-free 3D zero-shot anomaly detection framework. CS3F leverages a frozen 2D vision transformer and employs multi-axis slice decomposition, neighboring slice feature pooling, and a coarse-to-fine voxel-level tokenization strategy to preserve lesion signals while enabling high-resolution anomaly localization. It introduces a novel cross-subject similarity metric to compute anomaly scores. Experiments demonstrate that CS3F achieves effective performance on brain MRI (metastases, gliomas, stroke) and lung CT datasets, validating that frozen 2D foundation models can be successfully adapted for 3D anomaly detection, with fine-grained tokenization efficacy influenced by lesion contrast and imaging modality.
This work proposes a training-free, zero-shot visual anomaly localization method that operates without access to normal samples of the target category or any textual prompts. By leveraging a pre-trained DDIM model, the approach performs diffusion inversion followed by partial denoising at intermediate timesteps to reconstruct the input image; anomalies are localized through discrepancies between the original and reconstructed images. As the first method to achieve high-precision spatial anomaly localization under purely visual zero-shot conditions, it eliminates reliance on language prompts or auxiliary modalities. The proposed technique attains state-of-the-art performance on the VISA dataset, significantly advancing the practicality and accuracy of zero-shot anomaly detection.
This work addresses the challenge of 3D point cloud anomaly detection, where the scarcity and diversity of anomalous samples typically restrict training to normal data alone, thereby limiting model generalization. To overcome this, the authors propose a modular framework that enhances unsupervised training by synthesizing diverse pseudo-anomalies. Specifically, they construct a parametric deformation model based on local PCA coordinate systems, enabling anisotropic, direction-gated, and normal/tangential displacement fields to generate a rich variety of geometric defects. The approach is highly versatile, compatible with both reconstruction- and offset-prediction-based detection paradigms. Experiments on AnomalyShapeNet and Real3D-AD demonstrate significant improvements in both object-level and point-level anomaly detection and localization performance, while ablation studies confirm the effectiveness of individual components and robustness to noise.
Zero-shot anomaly detection (ZSAD) aims to identify anomalies across domains without access to target-domain training samples; however, its generalizability is severely hindered by substantial discrepancies in foreground objects, anomaly appearances, and background distributions. To address this, we propose a CLIP-based universal ZSAD framework. Our method introduces the first object-agnostic text prompt learning mechanism, decoupling foreground semantics from normal/abnormal discrimination modeling. We further design learnable, domain-agnostic “normal” and “abnormal” text prompts and leverage vision-language feature alignment to enable zero-shot anomaly scoring and pixel-level localization. Critically, our approach requires no target-domain annotations or fine-tuning. Extensive experiments across 17 industrial defect and medical imaging datasets demonstrate significant improvements over existing ZSAD methods, achieving both strong cross-domain generalization and high localization accuracy.
This work addresses the limitations of existing zero-shot 3D anomaly detection methods, which often render point clouds into 2D images and consequently lose critical geometric details, leading to insufficient sensitivity to local anomalies. To overcome this, the authors propose the BTP framework, which pioneers the use of pretrained point cloud–language models for this task. BTP aligns multi-granularity point cloud patches with textual embeddings, integrates geometric descriptors, and leverages auxiliary point cloud data for joint representation learning, thereby significantly enhancing the model’s ability to perceive and localize structural anomalies. Extensive experiments on the Real3D-AD and Anomaly-ShapeNet benchmarks demonstrate that BTP substantially outperforms current state-of-the-art methods, achieving the best-reported performance in zero-shot 3D anomaly detection.
This work identifies and formally characterizes the “consistent anomaly” problem in zero-shot anomaly classification and segmentation (AC/AS)—a systematic bias in distance-based methods caused by recurrent, visually similar anomalies—rooted in similarity scaling imbalance and k-nearest-neighbor exhaustion. To mitigate bias propagation, we propose CoDeGraph, a graph-based framework integrating multi-stage graph construction, community-aware structural optimization, and pseudo-mask-guided vision-language supervision. Furthermore, we introduce a training-free 3D volumetric patching strategy enabling truly zero-shot, voxel-level MRI anomaly segmentation. Our method requires neither anomaly annotations nor 3D training data. Evaluated on multiple benchmarks, it achieves significant improvements in localization accuracy—particularly under high-anomaly-similarity conditions—and enhances robustness. This advances the practicality of text-driven anomaly detection in medical imaging.
This work addresses zero-shot anomaly detection in scenarios where normal data from the target domain is unavailable, a setting in which existing CLIP-based methods suffer from spatial misalignment and insufficient sensitivity to fine-grained anomalies. To overcome these limitations, the authors propose a decoupled prompting mechanism built upon a spatially aware TIPS vision-language model: a fixed prompt facilitates image-level anomaly detection, while a learnable prompt enables pixel-level localization. Global anomaly scores are further refined through local evidence fusion, effectively bridging the distribution gap between global and local features without requiring complex auxiliary modules. Evaluated on seven industrial datasets, the method achieves consistent improvements—1.1–3.9% in image-level detection and 1.5–6.9% in pixel-level localization—significantly outperforming current state-of-the-art approaches.
This work addresses the limitations of existing zero-shot anomaly detection methods, which rely solely on spatial features and struggle to capture subtle textural and structural deviations—particularly in the absence of target-domain data. To overcome this, we propose the first frequency-aware zero-shot anomaly detection framework, introducing a novel local spatial-frequency discrepancy modeling approach. Our method employs a Local Frequency Compensation Module (LFCM), a Frequency Discrepancy Anchor Projector (FDAP), and Asymmetric Anchor Supervision (AAS) to construct normal and anomalous frequency-domain anchors for relative similarity discrimination. Notably, it achieves language-agnostic detection without requiring textual prompts. Extensive experiments demonstrate that our approach sets new state-of-the-art average performance across 13 industrial and medical benchmarks, excelling in both image-level recognition and pixel-level localization tasks.
This work addresses the challenge of zero-shot anomaly detection and localization in the absence of labeled anomalous samples by introducing the first purely vision-based framework that entirely dispenses with text encoders and cross-modal alignment. Built upon the Vision Transformer architecture, the method incorporates learnable semantic tokens representing normal and anomalous concepts, coupled with a spatially aware cross-attention (SCA) mechanism and a self-alignment fusion (SAF) strategy to enable robust and efficient anomaly discrimination. The proposed framework achieves state-of-the-art performance across 13 industrial and medical benchmark datasets and seamlessly integrates with various pretrained visual backbones—such as CLIP and DINOv2—significantly enhancing its generalizability and practical applicability.