RVLM: Recursive Vision-Language Models with Adaptive Depth

๐Ÿ“… 2026-03-25
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the limitations of conventional vision-language models in medical AIโ€”namely, their lack of interpretability, low computational efficiency in iterative systems, and fixed reasoning depthโ€”by proposing a recursive generate-and-execute framework. The approach enables auditable clinical reasoning through code-based invocation of visual sub-agents that manipulate images and accumulate evidence, while a lightweight adaptive routing controller, termed RRouter, dynamically adjusts reasoning depth. This method uniquely integrates auditable, code-driven reasoning with adaptive depth control. Evaluated on BraTS 2023 meningioma MRI and MIMIC-CXR chest X-ray datasets, it effectively identifies cross-modal inconsistencies, generates structured reports, and accurately detects view-specific artifacts, substantially enhancing both transparency and computational efficiency in medical AI systems.

Technology Category

Computer Vision: Multi-modal VisionKnowledge Representation and Reasoning: Diagnosis and Abductive ReasoningPlanning, Routing, and Scheduling: Planning with Language Models

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGResponsible Web: Machine-in-the-loop, human agency and autonomyGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
๐Ÿ“ Abstract
Medical AI systems face two fundamental limitations. First, conventional vision-language models (VLMs) perform single-pass inference, yielding black-box predictions that cannot be audited or explained in clinical terms. Second, iterative reasoning systems that expose intermediate steps rely on fixed iteration budgets wasting compute on simple cases while providing insufficient depth for complex ones. We address both limitations with a unified framework. RVLM replaces single-pass inference with an iterative generate-execute loop: at each step, the model writes Python code, invokes vision sub-agents, manipulates images, and accumulates evidence. Every diagnostic claim is grounded in executable code, satisfying auditability requirements of clinical AI governance frameworks. RRouter makes iteration depth adaptive: a lightweight controller predicts the optimal budget from task-complexity features, then monitors progress and terminates early when reasoning stalls. We evaluate on BraTS 2023 Meningioma (brain MRI) and MIMIC-CXR (chest X-ray) using Gemini 2.5 Flash without fine-tuning. Across repeated runs, RVLM shows high consistency on salient findings (e.g., mass presence and enhancement) and can detect cross-modal discrepancies between Fluid-Attenuated Inversion Recovery (FLAIR) signal characteristics and segmentation boundaries. On MIMIC-CXR, it generates structured reports and correctly recognises view-specific artefacts. Code: https://github.com/nican2018/rvlm.
Problem

Research questions and friction points this paper is trying to address.

vision-language models
iterative reasoning
auditability
adaptive depth
medical AI
Innovation

Methods, ideas, or system contributions that make the work stand out.

Recursive Vision-Language Model
Adaptive Iteration Depth
Executable Reasoning
Clinical Auditability
RRouter
๐Ÿ’ผ Related Jobs
No related jobs found.
N
Nicanor Mayumu
Department of Computer Science, University of Wollongong in Dubai, Dubai Knowledge Park, Dubai, UAE
Z
Zeenath Khan
Department of Computer Science, University of Wollongong in Dubai, Dubai Knowledge Park, Dubai, UAE
M
Melodena Stephens
Department of Academic Affairs, Mohammed Bin Rashid School of Government, City Walk, Building B02, Dubai, UAE
P
Patrick Mukala
Department of Computer Science, University of Wollongong in Dubai, Dubai Knowledge Park, Dubai, UAE
Farhad Oroumchian
Farhad Oroumchian
Professor of Computer Science, University of Wollongong in Dubai
Information RetrievalNatural language ProcessingArtificial IntelligenceData Mining