SynDORBench: Evaluating LVLM Perceptual Robustness Under Physically Constrained Visibility Conditions

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of perceptual reliability evaluation for large vision-language models (LVLMs) under physically constrained imaging. We propose the first physics-constrained benchmark based on the DORI standard, generating 54,000 question-answer pairs through a controllable synthesis pipeline. By integrating a discriminability annotation framework with mask-conditioned statistical feature propagation, this work systematically evaluates model robustness across varying distances, illumination levels, and viewing angles. Our findings reveal that pixel density, rather than model scale, predominantly drives perceptual failures, and that compact open-source models outperform commercial baselines under long-range, low-light conditions. Ultimately, this research establishes a new paradigm for multimodal evaluation under real-world physical constraints.
📝 Abstract
Large vision-language models (LVLMs) have demonstrated remarkable performance on multimodal reasoning benchmarks, yet their perceptual reliability under physically constrained imaging conditions remains poorly understood. Existing evaluations predominantly assume ideal visual inputs and therefore fail to characterize how camera distance, illumination, viewpoint, and pixel density fundamentally affect semantic recoverability. We introduce SynDORBench, the first physically grounded benchmark for evaluating LVLM perceptual robustness under DORI-calibrated conditions aligned with human visual capability standards. SynDORBench comprises over 54k question--answer pairs generated through a controllable synthetic pipeline that systematically varies viewing distance, lighting, camera geometry, and action pose according to physically interpretable pixel-density regimes. To support scalable low-visibility supervision, we further propose a discernibility annotation framework that propagates human perceptual labels using mask-conditioned statistical features and ensemble learning. We evaluate 16 open-source LVLMs, a commercial LVLM baseline, and YOLO11x across human-presence classification and action recognition tasks under progressively degraded visibility conditions. Our results reveal that perceptual failure in LVLMs is strongly governed by pixel density and physical imaging constraints rather than model scale alone. Surprisingly, several compact open-source LVLMs outperform larger commercial baselines and substantially exceed YOLO11x robustness under long-range and low-light conditions. SynDORBench establishes a new benchmark paradigm for physically grounded multimodal evaluation, enabling systematic analysis of LVLM reliability under real-world perceptual constraints and direct comparison against human visibility thresholds.
Problem

Research questions and friction points this paper is trying to address.

Large Vision-Language Models
Perceptual Robustness
Physical Imaging Constraints
Visibility Conditions
Benchmark Evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Physically grounded benchmark
Perceptual robustness
Controllable synthetic pipeline
Discernibility annotation framework
Large vision-language models
J
Jeremy Stephen Gabriel Yee
Singapore Institute of Technology
Z
Zhengkui Wang
Singapore Institute of Technology
Zhiyuan Zhang
Zhiyuan Zhang
Singapore Management University
3D Point CloudDeep LearningPattern Recognition3D Biometrics
A
Avinash Anand
Singapore Institute of Technology
T
Timothy Liu
NVIDIA AI Technology Center
B
Benedict Chan
Singapore Institute of Technology
A
Aik Beng Ng
NVIDIA AI Technology Center
Simon See
Simon See
nvidia
applied mathematicsAImachine learningHigh Performance ComputingSimulation