FinEval-KR: A Financial Domain Evaluation Framework for Large Language Models' Knowledge and Reasoning

📅 2026-04-11
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing LLM evaluation in finance struggles to disentangle knowledge acquisition from reasoning capability and lacks fine-grained attribution analysis. Method: We propose the first fine-grained evaluation framework for Chinese financial LLMs, featuring a dual-metric system (“knowledge score” and “reasoning score”), cognitively interpretable scoring grounded in Bloom’s Taxonomy, and an open-source financial reasoning dataset covering 22 subdomains—constructed via integration of a Chinese financial knowledge graph, domain-adaptive prompting, and multi-dimensional error attribution. Contributions/Results: First, we identify a critical bottleneck: finance-specialized models underperform general-purpose LLMs in knowledge transfer applications. Second, empirical analysis confirms that higher-order reasoning ability and cognitive hierarchy are primary determinants of model performance. Third, we demonstrate that state-of-the-art finance-specialized LLMs still lag behind general-purpose counterparts overall.

Technology Category

Knowledge Representation and Reasoning: Knowledge AcquisitionCognitive Modeling & Cognitive Systems: Conceptual Inference and ReasoningMachine Learning: Large Multimodal Models (LMMs)

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systemsGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
Large Language Models (LLMs) demonstrate significant potential but face challenges in complex financial reasoning tasks requiring both domain knowledge and sophisticated reasoning. Current evaluation benchmarks often fall short by not decoupling these capabilities indicators from single task performance and lack root cause analysis for task failure. To address this, we introduce FinEval-KR, a novel evaluation framework for decoupling and quantifying LLMs' knowledge and reasoning abilities independently, proposing distinct knowledge score and reasoning score metrics. Inspired by cognitive science, we further propose a cognitive score based on Bloom's taxonomy to analyze capabilities in reasoning tasks across different cognitive levels. We also release a new open-source Chinese financial reasoning dataset covering 22 subfields to support reproducible research and further advancements in financial reasoning. Our experimental results reveal that LLM reasoning ability and higher-order cognitive ability are the core factors influencing reasoning accuracy. We also specifically find that even top models still face a bottleneck with knowledge application. Furthermore, our analysis shows that specialized financial LLMs generally lag behind the top general large models across multiple metrics.
Problem

Research questions and friction points this paper is trying to address.

Evaluating LLMs' financial knowledge and reasoning separately
Addressing lack of root cause analysis in financial tasks
Assessing cognitive levels in financial reasoning using Bloom's taxonomy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Decouples knowledge and reasoning abilities evaluation
Introduces cognitive score based on Bloom's taxonomy
Releases open-source Chinese financial reasoning dataset
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Shaoyu Dou
Shaoyu Dou
Ant Group
AIOps (AI for Operation)Anomaly detectionTime series analysis
Y
Yutian Shen
Ant Group, Shanghai University of Finance and Economics
M
Mofan Chen
Ant Group, Shanghai University of Finance and Economics
Z
Zixuan Wang
Shanghai University of Finance and Economics
J
Jiajie Xu
Shanghai University of Finance and Economics
Q
Qi Guo
Ant Group
K
Kailai Shao
Ant Group
C
Chao Chen
Ant Group
H
Haixiang Hu
Ant Group
H
Haibo Shi
Ant Group
M
Min Min
Shanghai University of Finance and Economics
L
Liwen Zhang
Shanghai University of Finance and Economics