🤖 AI Summary
Existing zero-shot anomaly detection methods for 3D brain MRI are limited by their reliance on slice-level features, which neglect the intrinsic three-dimensional spatial structure among voxels. This work proposes a training-free, prompt-free, and supervision-free zero-shot anomaly detection framework that leverages multi-axis 2D foundation models to extract features from orthogonal slices, aggregates them into 3D local voxel tokens enriched with cubic spatial context, and directly applies distance-based batch anomaly scoring. To the best of our knowledge, this is the first approach to achieve fully training-free zero-shot anomaly detection on 3D brain MRI volumes. The method efficiently produces compact 3D representations on standard GPUs, successfully extending the zero-shot capabilities of 2D foundation models to full 3D volumes while maintaining both simplicity and robustness.
📝 Abstract
Zero-shot anomaly detection (ZSAD) has gained increasing attention in medical imaging as a way to identify abnormalities without task-specific supervision, but most advances remain limited to 2D datasets. Extending ZSAD to 3D medical images has proven challenging, with existing methods relying on slice-wise features and vision-language models, which fail to capture volumetric structure. In this paper, we introduce a fully training-free framework for ZSAD in 3D brain MRI that constructs localized volumetric tokens by aggregating multi-axis slices processed by 2D foundation models. These 3D patch tokens restore cubic spatial context and integrate directly with distance-based, batch-level anomaly detection pipelines. The framework provides compact 3D representations that are practical to compute on standard GPUs and require no fine-tuning, prompts, or supervision. Our results show that training-free, batch-based ZSAD can be effectively extended from 2D encoders to full 3D MRI volumes, offering a simple and robust approach for volumetric anomaly detection.