DVLA-RL++: Dual-Level Vision-Language Alignment with Reinforcement Learning Gating for Few-Shot Learning

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the instability of image-text prototypes caused by contextual contamination in few-shot learning. To this end, it proposes a complementary semantic purification framework coupled with a counterfactual reinforcement learning gating mechanism. The method dynamically optimizes vision-language fusion intensity by decoupling intrinsic attributes from distracting contexts. Furthermore, it introduces an ambiguity-dependent rejection margin and intrinsic semantic anchors to achieve sparse evidence allocation, while constructing unbiased gradient estimates based on paired reference trajectories, supported by theoretical stability analysis. Extensive experiments demonstrate that the proposed approach achieves state-of-the-art accuracy across standard, fine-grained, and cross-domain benchmarks, yielding an average improvement of 1.4% over existing baselines.
📝 Abstract
Few-shot learning aims to recognize novel categories from limited labeled examples. Recent studies incorporate textual semantics to compensate for limited visual observations and improve class representations. However, high image-text agreement may reflect both intrinsic object properties and incidental context, making support prototypes susceptible to contextual contamination. To address this problem, we propose DVLA-RL++, which extends DVLA-RL with complementary semantic purification (CSP) and counterfactual reinforcement-learning gating (CRG). Specifically, CSP generates intrinsic and nuisance descriptions from labeled supports and compares their agreement with each support token. An ambiguity-dependent rejection margin guides sparse evidence allocation, while an intrinsic semantic anchor fills the unassigned mass to provide a fallback when visual evidence is unreliable. CRG learns layer-wise semantic fusion strengths using a reward that balances recognition performance and nuisance exposure. An independently executed reference trajectory on the same episode provides a paired learning signal. Theoretical analysis relates retained evidence and anchor quality to prototype stability and establishes conditions for unbiased on-policy gradient estimation. Experiments on standard, fine-grained, and cross-domain benchmarks show state-of-the-art accuracy, with an average gain of 1.4% over DVLA-RL. The project page is available at https://peacelwh.github.io/TPAMI27-DVLA-RLpp/.
Problem

Research questions and friction points this paper is trying to address.

Few-Shot Learning
Vision-Language Alignment
Contextual Contamination
Semantic Purification
Innovation

Methods, ideas, or system contributions that make the work stand out.

Few-Shot Learning
Vision-Language Alignment
Reinforcement Learning
Semantic Purification
Counterfactual Gating