Score
Designing and running human-subject evaluations and perceptual metrics to measure visual quality, physics-based reasoning, novel-view fidelity, and temporal coherence of models, and comparing model performance against baselines while accounting for cost and efficiency.
Accurately measuring visual fidelity and predicting human perception remains challenging, particularly in distinguishing perceptual differences across object categories and simplification types. Method: We propose a systematic evaluation framework integrating polygon mesh simplification (using two distinct algorithms) with psychophysical experiments—namely naming time, subjective rating, and pairwise preference judgments—to rigorously compare perceptual responses to animals versus man-made objects under controlled simplification levels. Contribution/Results: This work introduces the first three-factor controlled analysis jointly varying simplification type, degree, and object category. We find naming time and preference tasks exhibit higher sensitivity to fidelity degradation than subjective ratings; conversely, state-of-the-art image- and mesh-based automatic metrics reliably predict only subjective ratings, failing to model higher-order perceptual decisions. These findings expose fundamental cognitive limitations of prevailing fidelity metrics and provide empirical grounding and a methodological paradigm for developing human-centered graphics fidelity assessment frameworks.
This work proposes a human-centered evaluation paradigm for visual processing systems that moves beyond reliance on single image quality metrics, which often fail to capture human perception and user preferences. By integrating objective image quality assessment (IQA), human perceptual experiments, and fine-grained modeling of user preferences, the study establishes a context-aware, comprehensive evaluation framework. The research uncovers significant discrepancies between widely used image quality metrics and actual human judgments, thereby offering both theoretical insights and methodological support for developing more application-aligned evaluation protocols for visual models.
This study investigates whether vision-language models (VLMs) can effectively emulate human judgments in perceptual image quality assessment, potentially replacing costly psychophysical experiments. For the first time, we systematically evaluate six VLMs—four closed-source and two open-source—against human judgments across three dimensions: contrast, color saturation, and overall preference, using established psychophysical data as a benchmark. Our analysis integrates attribute-weighted evaluation and intra-model consistency metrics. Results show that VLMs achieve up to 0.93 correlation with human judgments on color saturation but perform notably weaker on contrast. Most models align with human behavior by prioritizing color saturation in overall preference. A key contribution is revealing the trade-off between model self-consistency and human alignment, and demonstrating that enhancing perceptual separability improves human-model agreement.
This study investigates whether mainstream image/video quality metrics—such as SSIM, LPIPS, and VMAF—faithfully model human low-level visual mechanisms, including contrast sensitivity, contrast masking, and contrast matching. To this end, we introduce the first interpretability benchmark framework specifically designed for low-level vision properties, grounded in psychophysical principles: contrast sensitivity function testing, masking threshold estimation, and contrast-matching discrimination tasks. We systematically evaluate 33 full-reference metrics under this framework. Results show that LPIPS and MS-SSIM effectively capture contrast masking, whereas VMAF exhibits significant deficiencies; SSIM over-responds to high-frequency distortions, while MS-SSIM substantially mitigates this bias. Our analysis uncovers structural shortcomings in current metrics’ perceptual modeling, revealing misalignments with neurobiologically grounded vision principles. This work establishes a new paradigm for interpretable, neurophysiologically plausible quality assessment and provides empirical foundations for next-generation perceptual metrics.
Existing video-language model (VLM) evaluation benchmarks are vulnerable to spurious visual or textual shortcuts, yielding inflated scores and failing to reliably assess spatiotemporal and physical reasoning capabilities. To address this, we introduce Minimal Video Pairs (MVP), a rigorous benchmark comprising 55K high-quality video-question-answer samples. MVP pioneers the “minimal-difference video pair” paradigm: for each question, two semantically similar videos yield opposite correct answers, compelling models to perform deep physical reasoning rather than exploit superficial cues. The benchmark integrates first- and third-person videos, robot interaction sequences, and cognitive-science-inspired intuitive physics data, and employs dual-sample joint evaluation with strict paired annotation. Human accuracy is 92.9%, while the strongest open-source VLM achieves only 40.2%—significantly above the 25% random baseline—demonstrating MVP’s unprecedented bias-resilience and discriminative power for evaluating physical understanding.
This work addresses the lack of human-level perceptual reasoning and judgment consistency in blind image quality assessment (BIQA) models. To this end, we propose a perception–reasoning cascaded framework that explicitly models the human cognitive chain—“sensory input → implicit reasoning → quality judgment”—as a learnable, self-consistent reasoning path. We further introduce a reinforcement learning reward mechanism grounded in self-generated quality descriptions, balancing alignment with human preferences and internal logical consistency. Our method integrates human annotations, natural language generation, and ROUGE-1-based interpretability evaluation to achieve end-to-end interpretable BIQA. Experiments demonstrate state-of-the-art performance: highest Pearson and Spearman correlation coefficients among existing methods; ROUGE-1 score of 0.512—significantly surpassing the baseline (0.443)—validating high fidelity to human reasoning chains.
Existing large vision-language model benchmarks are often confined to single tasks and lack comprehensive evaluation of fine-grained, multi-dimensional perception and reasoning capabilities in human-centric scenarios. To address this gap, this work proposes the MHPR benchmark, which encompasses three key dimensions: individual humans, multi-person interactions, and human-object interactions. It introduces a four-level data hierarchy and an automated high-quality annotation pipeline, ACVG, enabling joint assessment of fine-grained attributes and high-level semantics. Efficient automatic annotation is achieved through category-level attribute decomposition, attribute-specific rewriting, and multi-model voting. Combined with supervised fine-tuning and a reinforcement learning strategy based on failure-case analysis, this approach significantly enhances instruction following, robustness, and complex scene reasoning on Qwen2.5-VL-7B, achieving performance comparable to much larger models and demonstrating the effectiveness and practicality of MHPR.
This work addresses the lack of fine-grained evaluation of physical plausibility in current video generation models, which hinders the identification of specific causes behind violations of physical laws during dynamic processes. To this end, we introduce a large-scale benchmark grounded in expert human reasoning, featuring fine-grained reasoning trajectories that include temporal localization, structured failure categories, and natural language explanations across 22 physical phenomena. The benchmark integrates real reference videos, expert annotations, and a physics-based taxonomy to form a high-quality human-evaluated dataset. Experiments reveal that among videos generated by state-of-the-art models in physics-critical scenarios, 83.3% (third-person) and 93.5% (first-person) contain at least one human-identifiable physical inconsistency, underscoring the urgent need for standardized evaluation protocols and highlighting the diagnostic value of our benchmark.
This work addresses the discrepancy between high benchmark scores and fragile real-world perceptual capabilities of multimodal models by proposing a fine-grained evaluation framework that bridges the gap between automated metrics and human judgment. Built upon 1,038 high-information-density images and over 12,000 instance-level scoring rules, the framework introduces dual-stream criteria—Must-Right and Easy-Wrong—employs circular peer review to construct gold-standard annotations, and implements a “fail-as-penalty” gated scoring mechanism. Experiments demonstrate that this approach significantly improves alignment with human assessments, exposes critical reliability gaps in dense visual scenes, reveals an 8% perception performance disparity between open- and closed-source models, and validates that the gated metric better reflects human perception than conventional benchmarks.