🤖 AI Summary
Existing benchmarks inadequately assess the ability of multimodal large language models to integrate visual business evidence in executive decision-making. To address this gap, this work introduces C-SUITEBENCH, a multimodal benchmark that systematically evaluates nine state-of-the-art models across 50 real-world business scenarios under both textual and multimodal conditions, focusing on CEO-level strategic decisions. Through controlled experiments, ablation studies, and visual channel analyses, the study uncovers a “multimodal integration paradox”: while visual information substantially enhances risk prediction and argumentation capabilities, it consistently impairs constrained decisions such as resource allocation, primarily due to signal crowding effects. These findings reveal that perception and action constitute distinct bottlenecks in multimodal agents, highlighting a critical divergence between interpretive and operational intelligence in complex business contexts.
📝 Abstract
Large language models are increasingly applied as autonomous decision-making agents. However, in executive business decisions, existing benchmarks are limited to textonly settings. This makes it unclear whether models can perceive visual business evidence and effectively integrate it to improve decision quality. We introduce C-SUITEBENCH, a controlled multimodal benchmark that includes five decision tasks under paired text-only and multimodal conditions across 50 scenarios. We place nine frontier models in the role of a chief executive officer and evaluate their decision-making ability. Multimodal inputs consistently improve evidence-centric reasoning, with the largest and most reliable gains appearing in risk forecasting and board-facing justification. However, we uncover a multimodal integration paradox: adding visual business information degrades constrained resource allocation for all nine models, even as visual grounding itself improves. Ablation experiments reveal that this failure emerges from signal crowding, although each visual channel helps individually, their combination disrupts constraint satisfaction during decoding. These findings demonstrate that visual perception and constrained action are separable bottlenecks in multimodal agents, and that indiscriminate visual augmentation can harm high-stakes decision making, motivating selective grounding strategies for future executive AI systems.