Unlocking Data Value in Finance: A Study on Distillation and Difficulty-Aware Training

📅 2026-03-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenges faced by large language models in the financial domain—namely, dense domain-specific terminology, complex numerical reasoning, and low tolerance for factual inaccuracies. The authors propose a post-training data construction paradigm grounded in data quality and a difficulty-verifiability distribution. They generate high-quality chain-of-thought supervision data through multi-stage knowledge distillation and introduce a reinforcement learning strategy that explicitly accounts for both task difficulty and answer verifiability. The resulting model, ODA-Fin-RL-8B, consistently outperforms same-scale open-source state-of-the-art models across nine financial benchmarks. Furthermore, the study releases high-quality supervised fine-tuning and reinforcement learning datasets alongside model weights, offering the first systematic evidence of how post-training data characteristics critically determine the performance of financial large language models.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: (Large) Language ModelsData Mining & Knowledge Management: Linked Open Data, Knowledge Graphs & KB Completion

Application Category

Semantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsEconomics, Online Markets and Human Computation: Trust and reliance of crowd workers and data experts on GenAISearch and Retrieval-Augmented AI: Large language models for search
📝 Abstract
Large Language Models (LLMs) have demonstrated strong general capabilities, yet their deployment in finance remains challenging due to dense domain-specific terminology, stringent numerical reasoning requirements, and low tolerance for factual errors. We conduct a controlled empirical study showing that in specialized vertical domains, performance is largely determined by the quality and difficulty/verifiability profile of post-training data. We introduce \textbf{ODA-Fin-SFT-318k}, constructed via multi-stage distillation and verification to produce high-quality Chain-of-Thought supervision, and \textbf{ODA-Fin-RL-12k}, curated for hard-but-verifiable tasks that balance reward precision and task diversity. Using standard SFT and RL pipelines, we show that high-quality CoT distillation establishes a robust foundation during SFT, while difficulty- and verifiability-aware sampling improves RL generalization. Evaluated on nine benchmarks spanning general financial tasks, sentiment analysis, and numerical reasoning, our ODA-Fin-RL-8B consistently surpasses open-source state-of-the-art (SOTA) financial LLMs of comparable size. We release our ODA-Fin-SFT-318k and ODA-Fin-RL-12k datasets, along with trained models to advance data-centric financial AI research.
Problem

Research questions and friction points this paper is trying to address.

financial LLMs
domain-specific terminology
numerical reasoning
factual errors
data quality
Innovation

Methods, ideas, or system contributions that make the work stand out.

distillation
difficulty-aware training
verifiability-aware sampling
Chain-of-Thought
financial LLMs
🔎 Similar Papers
No similar papers found.