🤖 AI Summary
This study addresses the lack of perceptual reliability evaluation for large vision-language models (LVLMs) under physically constrained imaging. We propose the first physics-constrained benchmark based on the DORI standard, generating 54,000 question-answer pairs through a controllable synthesis pipeline. By integrating a discriminability annotation framework with mask-conditioned statistical feature propagation, this work systematically evaluates model robustness across varying distances, illumination levels, and viewing angles. Our findings reveal that pixel density, rather than model scale, predominantly drives perceptual failures, and that compact open-source models outperform commercial baselines under long-range, low-light conditions. Ultimately, this research establishes a new paradigm for multimodal evaluation under real-world physical constraints.
📝 Abstract
Large vision-language models (LVLMs) have demonstrated remarkable performance on multimodal reasoning benchmarks, yet their perceptual reliability under physically constrained imaging conditions remains poorly understood. Existing evaluations predominantly assume ideal visual inputs and therefore fail to characterize how camera distance, illumination, viewpoint, and pixel density fundamentally affect semantic recoverability. We introduce SynDORBench, the first physically grounded benchmark for evaluating LVLM perceptual robustness under DORI-calibrated conditions aligned with human visual capability standards. SynDORBench comprises over 54k question--answer pairs generated through a controllable synthetic pipeline that systematically varies viewing distance, lighting, camera geometry, and action pose according to physically interpretable pixel-density regimes. To support scalable low-visibility supervision, we further propose a discernibility annotation framework that propagates human perceptual labels using mask-conditioned statistical features and ensemble learning. We evaluate 16 open-source LVLMs, a commercial LVLM baseline, and YOLO11x across human-presence classification and action recognition tasks under progressively degraded visibility conditions. Our results reveal that perceptual failure in LVLMs is strongly governed by pixel density and physical imaging constraints rather than model scale alone. Surprisingly, several compact open-source LVLMs outperform larger commercial baselines and substantially exceed YOLO11x robustness under long-range and low-light conditions. SynDORBench establishes a new benchmark paradigm for physically grounded multimodal evaluation, enabling systematic analysis of LVLM reliability under real-world perceptual constraints and direct comparison against human visibility thresholds.