PRISM-VLM: A Multi-Axis Discriminative Benchmark for Compact Vision-Language Models

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
针对紧凑视觉-语言模型的评估问题,提出PRISM-VLM多轴判别基准,通过七个维度综合评分,更可靠地区分模型性能。
📝 Abstract
Compact vision-language models (VLMs) now power a growing share of multimodal applications. The benchmarks used to compare them, however, inherit a frontier-centric design: each model is reduced to a single accuracy number, narrowing the inter-model gap on saturated suites and pressing models into low-score bands on harder ones. We introduce PRISM-VLM, a multi-axis discriminative benchmark that scores every item along seven axes covering the recurring failure modes (task quality, behavioral robustness, and capability bottlenecks) and combines them into a single PScore, with items recycled from fifteen public benchmarks. Across compact VLMs from the past two years, PScore separates model pairs more reliably than prior single-axis benchmarks under an item-level paired bootstrap, and surfaces behavioral differences these benchmarks average away. Even models with statistically indistinguishable PScores diverge sharply along the per-axis profile, particularly on sycophancy, which is nearly orthogonal to single-prompt accuracy. We will release the full pipeline, prompts, and per-item annotations.
Problem

Research questions and friction points this paper is trying to address.

Compact Vision-Language Models
Benchmark Design
Single-Axis Evaluation
Inter-Model Gap
Behavioral Differences
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-axis discriminative benchmark
compact vision-language models
PScore
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.