TULIP: Adapting Open-Source Large Language Models for Underrepresented Languages and Specialized Financial Tasks

📅 2025-08-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited applicability of open-source large language models (LLMs) to low-resource languages (e.g., Turkish) and vertical domains (e.g., finance), along with their weak capabilities in sensitive information handling and domain-knowledge integration. To this end, we propose a language-domain co-adaptation framework. Methodologically, we build a five-stage pipeline—comprising continual pretraining, controllable synthetic data generation, supervised fine-tuning, and customized benchmark evaluation—based on Llama 3.1 8B and Qwen 2.5 7B. Our key contribution is the first deep coupling of low-resource language adaptation with financial domain knowledge injection, enhanced via controlled synthetic data to improve modeling of sensitive information. Experimental results demonstrate that the adapted models significantly outperform baselines across Turkish financial NER, question answering, and compliance text analysis tasks, achieving an average +18.7% F1-score improvement—validating the effectiveness and generalizability of our joint optimization paradigm.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: (Large) Language ModelsApplication Domains: Humanities & Computational Social Science

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSearch and Retrieval-Augmented AI: Vertical and domain-specific search
📝 Abstract
Thanks to the growing popularity of large language models over the years, there is great potential for their applications in finance. Despite the exceptional performance of larger proprietary models, which are presented as black-box solutions through APIs, smaller models that can be hosted on-premise present opportunities for adaptability and privacy. Especially in cases where the management of sensitive information and application of domain knowledge is important, like finance, enhancing the capabilities of smaller models becomes crucial, notably for underrepresented languages. In this work, we introduce TULIP models, which adapt Llama 3.1 8B and Qwen 2.5 7B for domain and language adaptation, focusing on financial Turkish use cases. The five-stage development pipeline involves data collection, continual pre-training (CPT), benchmark design, synthetic data generation and supervised fine-tuning (SFT). The results show that the capabilities of the models can be enhanced to effectively accomplish targeted tasks in this specific domain and language.
Problem

Research questions and friction points this paper is trying to address.

Adapting open-source LLMs for underrepresented languages
Enhancing smaller models for specialized financial tasks
Addressing privacy and domain knowledge in finance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adapts Llama and Qwen models for finance
Uses five-stage pipeline including CPT and SFT
Enhances models for underrepresented languages like Turkish
🔎 Similar Papers
No similar papers found.
İ
İrem Demirtaş
Data Science & AI, Prometeia SPA
B
Burak Payzun
Data Science & AI, Prometeia SPA
S
Seçil Arslan
Data Science & AI, Prometeia SPA