🤖 AI Summary
This study addresses the challenge that vision-language models (VLMs) frequently suffer from classification failures in domains deviating from their pretraining distributions due to missing features. To overcome this, we propose a training-free inductive visual logic framework that exploits the asymmetry between the descriptive and discriminative capabilities of VLMs to transcend feature space limitations. Specifically, our method constructs a visual feature dictionary by extracting representations from few-shot images via dual-modal prompting, and employs a hierarchical filtering mechanism to enable inference-time classification. Extensive evaluations across multiple benchmarks demonstrate that the proposed approach achieves state-of-the-art aggregated accuracy while generating interpretable and feature-traceable predictions, effectively resolving the problem of cross-domain few-shot classification.
📝 Abstract
Generative vision-language models (VLMs) such as Qwen-VL and LLaVA achieve strong zero-shot performance on tasks overlapping with their pretraining distribution, yet fail on specialized domains where the required discriminative features were never learned, a regime we term distant out-of-distribution (OOD). Standard adaptation methods cannot overcome this representational absence because they operate within the encoder's existing feature space. However, VLMs retain a robust descriptive capacity even when discrimination collapses: a model that cannot classify a medical scan can still articulate its visual patterns. Exploiting this asymmetry, we introduce Inductive Visual Logic (IVL), a training-free framework that constructs classification knowledge from the model's surviving descriptive ability. IVL extracts visual traits from few-shot support images through dual-mode prompting, combining semantic descriptions with primitive visual observations, and organizes them into per-class trait dictionaries. At inference, hierarchical filtering identifies spatially grounded trait evidence for classification. Across multiple distant-OOD benchmarks, IVL achieves the highest aggregate accuracy under two VLM backbones while producing interpretable, trait-traceable predictions.