Institution profile

Prometeia

Industry researcheurope · it
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Benchmarks Are Not Validation: A System-Level View of Financial LLM Applications

Jul 30, 2026

Current evaluations of financial large language models (LLMs) rely excessively on static benchmarks and lack comprehensive validation across the full system stack. This work proposes the first LLM full-stack verification framework tailored to financial scenarios, encompassing data, model, retrieval, generation, agent behavior, governance, and deployment layers, advocating for verification as an ongoing engineering practice. By introducing a multi-rater LLM-as-a-judge mechanism integrated with scoring rubrics, consistency checks, and auditability, the framework uncovers system failure modes that static benchmarks fail to capture. The study defines critical failure types and advances novel directions—including system-aware benchmarks, agent trajectory validation, rater alignment protocols, and lifecycle-oriented verification standards—thereby shifting the evaluation paradigm from score-driven metrics toward evidence-based readiness for real-world decision-making.

0 citationsRead paper

TULIP: Adapting Open-Source Large Language Models for Underrepresented Languages and Specialized Financial Tasks

Aug 22, 2025

This study addresses the limited applicability of open-source large language models (LLMs) to low-resource languages (e.g., Turkish) and vertical domains (e.g., finance), along with their weak capabilities in sensitive information handling and domain-knowledge integration. To this end, we propose a language-domain co-adaptation framework. Methodologically, we build a five-stage pipeline—comprising continual pretraining, controllable synthetic data generation, supervised fine-tuning, and customized benchmark evaluation—based on Llama 3.1 8B and Qwen 2.5 7B. Our key contribution is the first deep coupling of low-resource language adaptation with financial domain knowledge injection, enhanced via controlled synthetic data to improve modeling of sensitive information. Experimental results demonstrate that the adapted models significantly outperform baselines across Turkish financial NER, question answering, and compliance text analysis tasks, achieving an average +18.7% F1-score improvement—validating the effectiveness and generalizability of our joint optimization paradigm.

0 citationsRead paper
Recent publications

Latest Papers

Benchmarks Are Not Validation: A System-Level View of Financial LLM Applications

Jul 30, 2026

Current evaluations of financial large language models (LLMs) rely excessively on static benchmarks and lack comprehensive validation across the full system stack. This work proposes the first LLM full-stack verification framework tailored to financial scenarios, encompassing data, model, retrieval, generation, agent behavior, governance, and deployment layers, advocating for verification as an ongoing engineering practice. By introducing a multi-rater LLM-as-a-judge mechanism integrated with scoring rubrics, consistency checks, and auditability, the framework uncovers system failure modes that static benchmarks fail to capture. The study defines critical failure types and advances novel directions—including system-aware benchmarks, agent trajectory validation, rater alignment protocols, and lifecycle-oriented verification standards—thereby shifting the evaluation paradigm from score-driven metrics toward evidence-based readiness for real-world decision-making.

0 citationsRead paper

TULIP: Adapting Open-Source Large Language Models for Underrepresented Languages and Specialized Financial Tasks

Aug 22, 2025

This study addresses the limited applicability of open-source large language models (LLMs) to low-resource languages (e.g., Turkish) and vertical domains (e.g., finance), along with their weak capabilities in sensitive information handling and domain-knowledge integration. To this end, we propose a language-domain co-adaptation framework. Methodologically, we build a five-stage pipeline—comprising continual pretraining, controllable synthetic data generation, supervised fine-tuning, and customized benchmark evaluation—based on Llama 3.1 8B and Qwen 2.5 7B. Our key contribution is the first deep coupling of low-resource language adaptation with financial domain knowledge injection, enhanced via controlled synthetic data to improve modeling of sensitive information. Experimental results demonstrate that the adapted models significantly outperform baselines across Turkish financial NER, question answering, and compliance text analysis tasks, achieving an average +18.7% F1-score improvement—validating the effectiveness and generalizability of our joint optimization paradigm.

0 citationsRead paper