Institution profile

Nomura Research Institute

Academic institutionasia · jp
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Learning Dispute Structure for Settlement Prediction in Financial ADR: A Multi-Task and Cross-Institutional Approach

Jul 19, 2026

This study addresses the lack of a unified data and modeling framework in financial alternative dispute resolution (ADR), which hinders accurate prediction of settlement outcomes. To overcome this limitation, the authors integrate data from multiple Japanese ADR institutions and propose a functional-label-based annotation scheme to characterize dispute structures. Building upon this representation, they develop a multi-task learning model that jointly performs dispute classification and settlement prediction. The work introduces, for the first time, a cross-institutional shared representation of dispute structure and demonstrates its partial generalizability across diverse ADR domains. Experimental results show that incorporating structural information significantly enhances settlement prediction performance, and when combined with large language models, the approach achieves or surpasses state-of-the-art results across multiple domains.

0 citationsRead paper

Constructing Synthetic Instruction Datasets for Improving Reasoning in Domain-Specific LLMs: A Case Study in the Japanese Financial Domain

Mar 01, 2026

This work addresses the ongoing challenge of balancing domain-specific expertise with robust reasoning capabilities in large language models. We propose a general-purpose methodology that systematically and automatically transforms domain vocabulary into high-quality synthetic instruction data enriched with Chain-of-Thought (CoT) reasoning trajectories. Applying this approach to the Japanese financial domain, we construct a large-scale instruction dataset comprising approximately 9.5 billion tokens. Our experiments demonstrate that CoT length significantly influences model performance while also revealing inherent limitations. After large-scale instruction tuning, the resulting model substantially outperforms baseline models on financial-domain benchmarks. Both the dataset and the fine-tuned model are publicly released on Hugging Face to support further research.

0 citationsRead paper
Recent publications

Latest Papers

Learning Dispute Structure for Settlement Prediction in Financial ADR: A Multi-Task and Cross-Institutional Approach

Jul 19, 2026

This study addresses the lack of a unified data and modeling framework in financial alternative dispute resolution (ADR), which hinders accurate prediction of settlement outcomes. To overcome this limitation, the authors integrate data from multiple Japanese ADR institutions and propose a functional-label-based annotation scheme to characterize dispute structures. Building upon this representation, they develop a multi-task learning model that jointly performs dispute classification and settlement prediction. The work introduces, for the first time, a cross-institutional shared representation of dispute structure and demonstrates its partial generalizability across diverse ADR domains. Experimental results show that incorporating structural information significantly enhances settlement prediction performance, and when combined with large language models, the approach achieves or surpasses state-of-the-art results across multiple domains.

0 citationsRead paper

Constructing Synthetic Instruction Datasets for Improving Reasoning in Domain-Specific LLMs: A Case Study in the Japanese Financial Domain

Mar 01, 2026

This work addresses the ongoing challenge of balancing domain-specific expertise with robust reasoning capabilities in large language models. We propose a general-purpose methodology that systematically and automatically transforms domain vocabulary into high-quality synthetic instruction data enriched with Chain-of-Thought (CoT) reasoning trajectories. Applying this approach to the Japanese financial domain, we construct a large-scale instruction dataset comprising approximately 9.5 billion tokens. Our experiments demonstrate that CoT length significantly influences model performance while also revealing inherent limitations. After large-scale instruction tuning, the resulting model substantially outperforms baseline models on financial-domain benchmarks. Both the dataset and the fine-tuned model are publicly released on Hugging Face to support further research.

0 citationsRead paper