FinRegQA-EU: Corruption-Based Preference Data for Grounded EU Financial Regulatory Question Answering

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the deficiency of large language models in EU financial regulatory knowledge and the inability of existing evaluations to detect fabricated citations. We construct an end-to-end evaluation and optimization pipeline based on official EBA/ESMA corpora. Methodologically, we employ an LLM-as-a-Judge protocol with multi-model consensus filtering for data curation, and inject regulatory failure modes to generate preference pairs for model fine-tuning. Our analysis reveals a critical blind spot: while supervised fine-tuning achieves high overall scores, it yields the lowest citation F1 and exacerbates hallucinations. We demonstrate that DPO and GRPO achieve superior trade-offs between conciseness and accuracy. Finally, we release the FinRegQA-EU benchmark and evaluation stack, validating that DPO and GRPO significantly improve citation F1 while exposing the flaws of standard evaluations in rewarding spurious verbosity.
📝 Abstract
Large Language Models (LLMs) struggle with region-specific factual knowledge, particularly in financial regulation. While benchmarks such as CFinBench, provide broad coverage of financial knowledge in other regions, no comparable resource exists for European financial regulation. We close this gap by proposing an end-to-end pipeline for evaluating and improving LLMs on European financial regulatory question answering. Our dataset is grounded in the official Q&A corpora of the European Banking Authority (EBA) and the European Securities and Markets Authority (ESMA). We evaluate candidate answers with a pointwise LLM-as-a-Judge protocol using three judges from distinct model families, retain only unanimously judged pairs, and sharpen the rejected side through a taxonomy of regulatory failure modes : law swaps, article swaps, hallucinated citations, and hallucinated text. By fine-tuning on the resulting preference pairs exposes a divergence at the core of our findings: supervised fine-tuning attains the highest judge score of any method yet got the lowest rule-based Citation F1. SFT is producing longer answers with three times as many citations, most of them unsupported. DPO and GRPO offer the better trade-off with more concise answers and the highest citation F1 among fine-tuned models. Standard LLM-as-a-Judge evaluation rewards citation density as evidence of grounding and cannot detect when those citations are fabricated. Thus, we detect a blind spot that matters wherever answers must be verifiable. We release the benchmark and evaluation stack at https://github.com/auliakharis/FinRegQA-EU.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Financial Regulation
Question Answering
Hallucination
Evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Financial Regulatory QA
Preference Data Corruption
LLM-as-a-Judge
Hallucination Taxonomy
Citation F1
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Aulia Kharis Rakhmasari
ETH Zürich
F
Fan Yu
ETH Zürich
Alexander Hoyle
Alexander Hoyle
AI Center Postdoctoral Fellow, ETH Zürich
Natural Language ProcessingComputational Social ScienceMachine Learning
E
Elliot Ash
ETH Zürich