🤖 AI Summary
This study addresses the limitation of existing financial benchmarks, which are typically confined to single modalities or tasks and thus fail to evaluate the complex reasoning capabilities of multimodal large language models. To this end, we construct an S&P 500 earnings dataset spanning text, audio, and visual modalities, accompanied by a twelve-task evaluation framework. Methodologically, we introduce a novel tri-modal complementary and non-redundant design that simulates analyst decision-making processes, systematically comparing image-text, audio-text, and arbitrary modality translation architectures. Experimental results demonstrate that smaller omni-modal models outperform larger dual-modal proprietary counterparts, validating the effectiveness and complementarity of tri-modal data. This work bridges a critical gap in the evaluation of multimodal financial forecasting.
📝 Abstract
Financial forecasting from earnings conference calls requires models to reason over complex corporate disclosures, market expectations, and subtle communication signals. However, existing financial benchmarks are often limited to unimodal inputs or single-task settings, making it difficult to evaluate whether multimodal large language models (LLMs) can support real-world financial analysis. In this paper, we introduce MM-FinEval, a novel benchmark designed to evaluate multimodal LLMs across multiple financial tasks. MM-FinEval spans a diverse timeline from 2019 to 2022. The entire proposed dataset contains 2,045 S\&P 500 conference earning calls as inputs and 12 financial task labels as outputs. Each input contains three modalities: a word-to-word text transcript of the earning call, the corresponding presentation slides used during the call, and the entire audio recording. To establish a rigorous evaluation framework, we analyze 19 baseline models across three distinct model categories: Image-Text, Audio-Text, and Any-to-Any configurations. We observe that small-size Any-to-Any models processing all three modalities achieve strong performance, even when compared against larger proprietary models restricted to two-modality inputs. This indicates that our tri-modal dataset design introduces useful, non-redundant information. These results validate that text, audio, and visual data serve as important, complementary signals that mimic the decision-making process of expert human analysts.