Beyond Balanced Accuracy: A Resolution and Parity-Controlled Benchmark for Vision-Language and Vision-Only Defect Assessment in UAV Power-Line Inspection

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究通过ElecVQA-Bench测试了视觉-语言模型与纯视觉模型在无人机电力线缺陷评估中的表现差异,调整输入分辨率等参数后发现两者差距缩小,挑战了视觉-语言模型普遍优越的说法。
📝 Abstract
Vision-language models (VLMs) are often reported to outperform task-specific vision backbones for unmanned aerial vehicle (UAV) power-line defect assessment. We test that claim on ElecVQA-Bench, a 56,972-item benchmark derived from the public InsPLAD dataset, across six evaluation choices: partition, evaluated item set, label space, replication, input resolution, and side information. On a matched partition, a Swin Transformer and the strongest adapted VLM differ by only 0.03 points at binary screening. At seven-way defect typing, increasing the vision backbones from 224 px to the measured pixel budget of the VLM preprocessor narrows the gap against InternVL3.5-8B from +20.53 to -0.57 points for ResNet-50 and from +23.67 to +4.70 points for Swin-T. A pixel-budget audit shifts Qwen3-VL-8B macro recall by 10.78 points, yet a source-pixel-matched InternVL control still leaves Qwen ahead by 7.43 to 13.61 points while using 56% fewer visual tokens, so neither source pixels nor token budget explains the difference between the two VLMs. A two-seed global replication changes Qwen binary accuracy and seven-way macro recall by 0.86 and 1.02 points. After split-specific retraining, Qwen does not lead at crop or image level, and a 14-tower, three-seed replication reverses the sign across seeds, giving mean common-six macro recall of 0.9085 for Qwen against 0.9509 for ResNet-50. No split regime yields a family-level advantage that survives multiple-comparison correction. The study supports a benchmark-audit contribution rather than a general claim of VLM superiority.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
UAV Power-Line Inspection
Defect Assessment
Benchmarking
Performance Comparison
Innovation

Methods, ideas, or system contributions that make the work stand out.

vision-language models
resolution control
benchmark audit
UAV power-line inspection
visual backbones
🔎 Similar Papers
No similar papers found.
L
Linghao Zhang
State Grid Sichuan Electric Power Research Institute, Chengdu 610041, China; Power System Security and Operation Key Laboratory of Sichuan Province
S
Siyu Xiang
State Grid Sichuan Electric Power Research Institute, Chengdu 610041, China; Power System Security and Operation Key Laboratory of Sichuan Province
Junwei Kuang
Junwei Kuang
Beijing Institute of Technology
Recommender SystemData MiningHealthcare
P
Peiyu Yi
State Grid Sichuan Electric Power Research Institute, Chengdu 610041, China; Power System Security and Operation Key Laboratory of Sichuan Province