SKstars at SHROOM: Visions Agreement-Guided Ensembling of Zero-Shot and LoRA-Adapted Vision--Language Models

📅 2026-09-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过结合零样本和LoRA调整模型的方法,解决了细粒度幻觉检测问题,并在SHROOM-Visions 2026竞赛中验证了其有效性。
📝 Abstract
This paper describes the SKstars submission to SHROOM-Visions 2026, a shared task on fine-grained hallucination detection in large vision-language model outputs. The task requires systems to identify hallucinated character spans, assign hallucination categories, and provide confidence estimates for their predictions. Our approach combines zero-shot predictions from Qwen2.5-VL-72B-Instruct with those of a LoRA-adapted Qwen2.5-VL-7B-Instruct model. The outputs of the two models are integrated through a lightweight ensemble procedure, followed by span refinement and confidence adjustment. We evaluate the main system components on a small internal development subset and report the performance of the submitted system on the official English test set. SKstars achieved a Cor+Lbl score of 0.2902, ranking 15th among 29 teams, and obtained Cor and IoU scores of 0.3642 and 0.3151, respectively, ranking 18th on both metrics. The results show that combining a large zero-shot model with a smaller adapted model provides a practical framework for multilingual and fine-grained hallucination localization, while also highlighting the difficulty of transferring development-set improvements to hidden test data. Code and predictions: https://github.com/aliathar1401/SK-Stars-shroom-visions-2026
Problem

Research questions and friction points this paper is trying to address.

hallucination detection
vision-language models
confidence estimates
fine-grained
Innovation

Methods, ideas, or system contributions that make the work stand out.

Zero-Shot
LoRA-Adapted Models
Ensembling
Fine-Grained Hallucination Detection
Vision-Language Models
🔎 Similar Papers
Ali Athar
Ali Athar
Amazon
Computer VisionVideo SegmentationObject Tracking
I
Imran Ahsan
Department of Smart City, Chung-Ang University
J
Joon-Yong Jung
Department of Radiology, Seoul St. Mary’s Hospital, The Catholic University of Korea