🤖 AI Summary
This work addresses the ongoing challenge of balancing domain-specific expertise with robust reasoning capabilities in large language models. We propose a general-purpose methodology that systematically and automatically transforms domain vocabulary into high-quality synthetic instruction data enriched with Chain-of-Thought (CoT) reasoning trajectories. Applying this approach to the Japanese financial domain, we construct a large-scale instruction dataset comprising approximately 9.5 billion tokens. Our experiments demonstrate that CoT length significantly influences model performance while also revealing inherent limitations. After large-scale instruction tuning, the resulting model substantially outperforms baseline models on financial-domain benchmarks. Both the dataset and the fine-tuned model are publicly released on Hugging Face to support further research.
📝 Abstract
In adapting LLMs to specific domains, achieving both domain expertise and reasoning ability remains an urgent challenge. This study proposes a general method for constructing high-quality synthetic instruction data for any domain, starting from domain-specific vocabulary. As a demonstration, we applied this method to the financial domain and constructed a large-scale instruction dataset totaling approximately 9.5 billion tokens with Chain-of-Thought reasoning traces. Evaluation results confirmed performance improvements over baseline models on financial benchmarks, demonstrating the effectiveness of our approach. We also report findings on the impact of reasoning trace length on performance and its limitations. Lastly, we open-source our models and datasets on https://huggingface.co/nri-ai .