🤖 AI Summary
This study addresses the bottleneck in few-shot anomaly detection where discrete text struggles to capture fine-grained visual discrepancies, and proposes the VD-DeepStack framework. This method innovatively conditions explicit visual differences between query and reference images within language reasoning. By integrating DINO features with the hierarchical structure of large vision-language models, it designs multi-depth difference evidence injection paths alongside an auxiliary visual context path, injecting dense discrepancy evidence into the decoder to enhance the synergy between visual comparison and linguistic reasoning. Evaluated across four industrial and two medical benchmarks, the proposed framework significantly outperforms existing text-based comparative reasoning methods, effectively advancing few-shot anomaly detection performance.
📝 Abstract
Few-shot visual anomaly detection is fundamentally a visual comparison task, requiring fine-grained inspection of a query against normal references. Many recent methods based on large vision-language models (LVLMs) emphasize comparative reasoning through language chain-of-thought. Yet discrete, abstract descriptions may underrepresent dense, fine-grained visual differences, leaving a gap between visual comparison and its expression in language. To address this gap, we propose Visual Difference DeepStack (VD-DeepStack), which explicitly conditions language reasoning on query-reference visual differences. Specifically, we fuse DINO features with the LVLM visual hierarchy to strengthen fine-grained representations, then construct dense difference evidence from residuals between query features and softly matched reference features. The difference-evidence path injects spatially weighted difference vectors into query-image states at multiple decoder depths, while an auxiliary visual-context path provides fine-grained appearance information to support their interpretation. Experiments on 4 industrial and 2 medical anomaly benchmarks demonstrate substantial improvements in few-shot anomaly detection over baselines relying on textual comparative reasoning. These results support mitigating the visual comparison-reasoning gap through the joint design of comparison representations and their integration into the decoder. Code will be released upon acceptance.