Humanity's Sixth Sense: Benchmarking Intuitive Visual Reasoning in Multimodal Models

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the oversight of intuitive visual reasoning in existing multimodal benchmarks by introducing HSS, which establishes intuitive visual reasoning as a quantifiable evaluation dimension for the first time. Encompassing spatiotemporal, social, and abstract implicit information, this benchmark bridges the gap between low-level perception and high-level cognitive analysis through a structured taxonomy, manual prompt engineering, and agent-based dynamic visual manipulation techniques. Experimental results demonstrate that the best-performing model achieves an accuracy of only 53.6%, substantially lagging behind the human baseline of 93.1%. This significant performance disparity reveals critical deficiencies in the intuitive reasoning capabilities of current multimodal large language models, highlighting the necessity for further advancement in this domain.
📝 Abstract
Humans perceive far more in a scene than what is explicitly depicted: a single glance captures past causes and future trajectories; a quick peek determines if a vehicle can fit between two parked cars; a few seconds of video reveals who holds authority in a room; and a fleeting clip highlights subtle abstract patterns like unwritten rules or hidden labels. This capacity reflects a form of humanity's sixth sense: an intuitive reasoning mechanism that recovers implicit information beyond raw sensory perception. Crucially, this rapid, zero-shot visual intuition underpins everyday navigation and social interaction, making it a vital capability for Multimodal Large Language Models (MLLMs) deployed alongside people. Existing visual benchmarks, however, target either deliberate expert-level analysis in academic and mathematical domains or low-level perception, leaving the intuitive reasoning that people perform largely untested. To bridge this gap, we introduce Humanity's Sixth Sense (HSS), a benchmark for intuitive visual reasoning. HSS spans diverse image and video inputs, organizes items under a structured taxonomy, and pairs each with human-written prompts probing the implicit temporal, spatial, social, and abstract structure that people infer at a glance. Frontier MLLMs fall short of human performance: participants reach 93.1% accuracy, while the strongest model, GPT-6-astra, reaches only 53.6% even at maximum reasoning effort. Despite excelling in many complex tasks that require advanced perception and knowledge, current models still struggle significantly on these visual tasks that are intuitive for humans. We further explore agentic setup that apply dynamic visual manipulation to HSS, which narrows but does not close the gap. HSS establishes intuitive visual reasoning as a measurable axis and directs attention to a capability that scaling on current benchmarks has so far left behind.
Problem

Research questions and friction points this paper is trying to address.

Intuitive Visual Reasoning
Multimodal Large Language Models
Benchmark
Implicit Information
Sixth Sense
Innovation

Methods, ideas, or system contributions that make the work stand out.

Intuitive Visual Reasoning
Multimodal Benchmark
Humanity's Sixth Sense
Agentic Setup
Implicit Information Inference
Xingang Guo
Xingang Guo
University of Illinois at Urbana-Champaign
OptimizationDynamics and ControlMachine Learning
J
Jing Gu
Elorian
B
Brian Jang
Scale AI
R
Renxiong Wang
Scale AI
Utkarsh Tyagi
Utkarsh Tyagi
University of Maryland, College Park
AIMachine LearningNLPMultimodal
D
Daniel Quigley
Scale AI
S
Steven Li
Scale AI
D
David Yan
Elorian
Daniel Yue Zhang
Daniel Yue Zhang
Amazon AGI, University of Notre Dame
Natural Language UnderstandingMisinformation Detection Edge Computing Human-Cyber-Physical Systems
D
Darvin Yi
Scale AI
F
Forrest Huang
Elorian
H
HiJae Kim
Scale AI
T
Tianyi Zhang
Elorian
J
Jared Lichtarge
Elorian
J
Jihua Huang
Elorian
L
Le Xue
Elorian
Manan Tomar
Manan Tomar
Postdoctoral Researcher at Microsoft Research NYC
Reinforcement LearningDeep LearningMachine Learning
Q
Qiuyi Richard Zhang
Elorian
R
Ruofei Yu
Elorian
Seth Neel
Seth Neel
Google
computer sciencemachine learningprivacyfairness
Y
Yaning Hu
Elorian
M
Marcella Valentine
Elorian
X
Xinzhe Jiang
Elorian
D
Daniel Evans
Scale AI
Chenguang Wang
Chenguang Wang
UC Santa Cruz