Inductive Visual Logic for Few-Shot Out-Of-Distribution Adaptation in VLMs

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that vision-language models (VLMs) frequently suffer from classification failures in domains deviating from their pretraining distributions due to missing features. To overcome this, we propose a training-free inductive visual logic framework that exploits the asymmetry between the descriptive and discriminative capabilities of VLMs to transcend feature space limitations. Specifically, our method constructs a visual feature dictionary by extracting representations from few-shot images via dual-modal prompting, and employs a hierarchical filtering mechanism to enable inference-time classification. Extensive evaluations across multiple benchmarks demonstrate that the proposed approach achieves state-of-the-art aggregated accuracy while generating interpretable and feature-traceable predictions, effectively resolving the problem of cross-domain few-shot classification.
📝 Abstract
Generative vision-language models (VLMs) such as Qwen-VL and LLaVA achieve strong zero-shot performance on tasks overlapping with their pretraining distribution, yet fail on specialized domains where the required discriminative features were never learned, a regime we term distant out-of-distribution (OOD). Standard adaptation methods cannot overcome this representational absence because they operate within the encoder's existing feature space. However, VLMs retain a robust descriptive capacity even when discrimination collapses: a model that cannot classify a medical scan can still articulate its visual patterns. Exploiting this asymmetry, we introduce Inductive Visual Logic (IVL), a training-free framework that constructs classification knowledge from the model's surviving descriptive ability. IVL extracts visual traits from few-shot support images through dual-mode prompting, combining semantic descriptions with primitive visual observations, and organizes them into per-class trait dictionaries. At inference, hierarchical filtering identifies spatially grounded trait evidence for classification. Across multiple distant-OOD benchmarks, IVL achieves the highest aggregate accuracy under two VLM backbones while producing interpretable, trait-traceable predictions.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Out-of-Distribution
Few-Shot Adaptation
Generative VLMs
Innovation

Methods, ideas, or system contributions that make the work stand out.

Inductive Visual Logic
Out-of-Distribution Adaptation
Vision-Language Models
Few-Shot Learning
Training-Free