Constructing Synthetic Instruction Datasets for Improving Reasoning in Domain-Specific LLMs: A Case Study in the Japanese Financial Domain

📅 2026-03-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the ongoing challenge of balancing domain-specific expertise with robust reasoning capabilities in large language models. We propose a general-purpose methodology that systematically and automatically transforms domain vocabulary into high-quality synthetic instruction data enriched with Chain-of-Thought (CoT) reasoning trajectories. Applying this approach to the Japanese financial domain, we construct a large-scale instruction dataset comprising approximately 9.5 billion tokens. Our experiments demonstrate that CoT length significantly influences model performance while also revealing inherent limitations. After large-scale instruction tuning, the resulting model substantially outperforms baseline models on financial-domain benchmarks. Both the dataset and the fine-tuned model are publicly released on Hugging Face to support further research.

Technology Category

Natural Language Processing: (Large) Language ModelsMachine Learning: Large Multimodal Models (LMMs)Application Domains: Humanities & Computational Social Science

Application Category

Economics, Online Markets and Human Computation: Humans versus LLMs for data annotation and labelingWeb Mining and Content Analysis: Large pretrained models with web dataSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactions
📝 Abstract
In adapting LLMs to specific domains, achieving both domain expertise and reasoning ability remains an urgent challenge. This study proposes a general method for constructing high-quality synthetic instruction data for any domain, starting from domain-specific vocabulary. As a demonstration, we applied this method to the financial domain and constructed a large-scale instruction dataset totaling approximately 9.5 billion tokens with Chain-of-Thought reasoning traces. Evaluation results confirmed performance improvements over baseline models on financial benchmarks, demonstrating the effectiveness of our approach. We also report findings on the impact of reasoning trace length on performance and its limitations. Lastly, we open-source our models and datasets on https://huggingface.co/nri-ai .
Problem

Research questions and friction points this paper is trying to address.

domain-specific LLMs
reasoning ability
synthetic instruction datasets
financial domain
Chain-of-Thought
Innovation

Methods, ideas, or system contributions that make the work stand out.

synthetic instruction data
domain-specific LLMs
Chain-of-Thought reasoning
financial domain adaptation
instruction tuning
🔎 Similar Papers
No similar papers found.
Y
Yuma Okochi
Nomura Research Institute, Ltd.
F
Fabio Milentiansen Sim
NRI Indonesia
T
Tomoyasu Okada
Nomura Research Institute, Ltd.