CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition

πŸ“… 2026-07-28
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the lack of systematic evaluation of multimodal in-context learning (ICL) capabilities, which hinders the identification of bottlenecks in vision-language models’ joint reasoning and knowledge acquisition. To bridge this gap, we propose CLBench-V, the first benchmark that formally defines multimodal ICL and introduces a three-dimensional evaluation framework encompassing context localization, application of new information, and acquisition of novel knowledge. We develop an automated pipeline to curate and generate high-quality data spanning diverse domains, including scientific reasoning, finance, long-document understanding, spatial reasoning, and web-based visual question answering. Evaluating six state-of-the-art multimodal models on 3,443 instances reveals that overall performance remains limited, with the best model achieving only a score of 0.2847. Among them, InternVL3.5-30B-A3B excels in context localization and knowledge learning, while Qwen3.5-Plus demonstrates superior performance in applying newly provided information.
πŸ“ Abstract
Real-world tasks often require models to learn from task-specific context rather than relying only on pre-trained knowledge. While recent work has highlighted this capability as context learning, existing evaluations mainly focus on textual contexts. In many practical settings, however, the context to be learned from is multimodal: scientific findings are conveyed through figures and tables, financial indicators are scattered across converted reports, and spatial decisions depend on maps, scenes, or web pages. We introduce CLBench-V, a benchmark for multimodal context learning that addresses the difficulty of localizing where context use breaks down by organizing tasks around three dimensions: context grounding, new information application, and new knowledge learning. CLBench-V combines converted public benchmarks with newly constructed datasets spanning domains such as science, finance, long-document understanding, spatial reasoning, and web-based visual question answering. To reduce the cost of constructing domain-specific context-learning tasks, we further use automated construction and filtering procedures for our newly built datasets. Across 3,443 instances and six recent multimodal models, the best overall score is only 0.2847, indicating that multimodal context learning remains far from saturated. Moreover, InternVL3.5-30B-A3B performs best on context grounding and new knowledge learning, while Qwen3.5-Plus performs best on new information application. We further analyze judge reliability, context length, image count, and representative failure cases. Code is available at https://github.com/IamLihua/CLBench-V.
Problem

Research questions and friction points this paper is trying to address.

multimodal context learning
context grounding
knowledge acquisition
visual question answering
benchmark evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

multimodal context learning
context grounding
benchmark
automated dataset construction
knowledge acquisition
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
L
Lai Wei
School of Computer Science, Shanghai Jiao Tong University; Zhongguancun Academy
C
Chengqi Li
School of Computer Science, Shanghai Jiao Tong University; Shanghai Innovation Institute
J
Jiapeng Li
School of Computer Science, Shanghai Jiao Tong University; Shanghai Innovation Institute
R
Ruina Hu
Zhongguancun Academy
Y
Yue Wang
Zhongguancun Academy
W
Weiran Huang
School of Computer Science, Shanghai Jiao Tong University; Shanghai Innovation Institute