Retrieval-Augmented Visual Prompting: Guiding Foundation Models in Two-Photon Imaging

📅 2026-08-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决双光子成像中模型适应性问题,提出一种基于检索增强的视觉提示方法RAVP,通过在输入中注入外部视觉记忆来指导基础模型。
📝 Abstract
Two-photon calcium imaging presents a challenging setting for foundation models: image appearance varies substantially across recordings and experimental conditions, annotations are scarce, and rapid adaptation is often needed. Rather than adapting model weights through fine-tuning, we ask whether a foundation model can be guided at inference time by injecting external visual memory directly into its input. We implement this idea with SAM 3 and introduce Retrieval-Augmented Visual Prompting (RAVP), a framework in which each target tile is augmented with a retrieved annotated exemplar whose bounding box is used as a concept prompt. RAVP turns retrieval into a form of visual prompting and enables adaptation through input design alone. We study multiple exemplar selection strategies, including fluorescence-guided heuristics and a lightweight recall predictor trained to estimate which exemplar is most informative for a target tile. Experiments on the Allen Brain Observatory show that exemplar-augmented inference consistently strengthens zero-shot neuron detection and instance segmentation. Ablation studies further show that a single carefully selected exemplar is more effective than prompting with multiple retrieved examples. These results position inference-time visual memory injection as a simple and effective alternative to parameter adaptation for foundation models in specialized biomedical imaging.
Problem

Research questions and friction points this paper is trying to address.

Two-Photon Imaging
Foundation Models
Visual Memory
Inference Time
Adaptation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Retrieval-Augmented Visual Prompting
external visual memory
inference-time adaptation
visual prompting
exemplar selection
🔎 Similar Papers
No similar papers found.
S
Salvatore Calcagno
PeRCeiVe Lab, Department of Electrical, Electronic and Computer Engineering, University of Catania, Catania, Italy
M
Marco Finocchiaro
PeRCeiVe Lab, Department of Electrical, Electronic and Computer Engineering, University of Catania, Catania, Italy
G
Giovanni Bellitto
PeRCeiVe Lab, Department of Electrical, Electronic and Computer Engineering, University of Catania, Catania, Italy
Daniela Giordano
Daniela Giordano
Professor of Artificial Intelligence at the University of Catania
Medical informaticsHuman-computer InteractionCognitive Systems
Concetto Spampinato
Concetto Spampinato
University of Catania
Deep LearningArtificial IntelligenceComputer VisionMedical Image Analysis
F
Federica Proietto Salanitri
PeRCeiVe Lab, Department of Electrical, Electronic and Computer Engineering, University of Catania, Catania, Italy