BIRD: Distilling Decision Boundaries into Rationales for MLLM Adaptation

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the disconnection between visual cues and decision logic during domain adaptation of multimodal large language models by proposing a boundary-aware rationale distillation framework. The method identifies decision boundaries through model confusion analysis, retrieves and verifies evidence via retrieval-augmented generation, and distills these into single-shot reasoning rationales to guide supervised fine-tuning, thereby enabling self-improving domain adaptation. Experimental results demonstrate that the proposed approach significantly outperforms existing baselines on medical and chart visual question answering tasks, effectively enhancing answer discriminability and the reliability of multimodal reasoning.
📝 Abstract
Adapting general-purpose multimodal large language models (MLLMs) to specialized domains requires learning domain-specific decision criteria, which often hinge on subtle visual distinctions between otherwise plausible answers. Rationale augmentation aims to expose such evidence through additional observations or inter-sample comparisons, yet a visually valid cue is not necessarily decision-relevant: it may describe how samples differ without changing the model's relative preference between competing answers. We therefore introduce BIRD, a self-improving Boundary-Informed Rationale Distillation framework that uses model-specific confusions to locate unresolved local decision boundaries and distills the evidence that resolves these confusions into rationales. For each sample, BIRD retrieves candidate neighbors from the target MLLM's own representation space and selects the most confusable one according to its answer preferences. It then generates answer-blind candidate evidence from their visual differences and functionally verifies which evidence most effectively strengthens the model's preference for the correct answer while avoiding inappropriate transfer across the pair. The verified evidence is then distilled into a single-sample rationale for standard supervised fine-tuning. Experiments on medical and chart VQA show that BIRD outperforms competing rationale-augmentation methods across two target MLLMs, while further analyses demonstrate clearer separation of confusable answers and stronger gains from model-matched supervision.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Large Language Models
Domain Adaptation
Decision Boundaries
Rationale Augmentation
Visual Question Answering
Innovation

Methods, ideas, or system contributions that make the work stand out.

Decision Boundary Distillation
Rationale Augmentation
Multimodal Large Language Models
Self-improving Framework
Confusion-aware Retrieval
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Anglin Liu
HKUST(GZ)
Y
Yanlin Wu
HKUST(GZ)
R
Ruichao Chen
HKUST
Yuting Zhang
Yuting Zhang
HKUST(GZ)
rPPGComputer Vision
Qingyuan Zeng
Qingyuan Zeng
Xiamen University
computer vision
P
Pengxiang Cai
HKUST(GZ)
Z
Ziqi Gong
HKUST(GZ)
M
Muchen Li
HKUST(GZ)
Jintai Chen
Jintai Chen
Assistant Professor@HKUST(GZ)
AI for HealthcareMultimodal LearningDeep Tabular Learning