🤖 AI Summary
This study addresses cross-lingual causal relation extraction in English–Spanish financial texts through a question-answering framework to enable causal inference. It systematically evaluates the performance of multilingual large language models—including mBERT, mBART, Llama 3.1, and the GPT series—across encoder, encoder-decoder, and decoder-only architectures, offering the first comparative analysis of prompt engineering, few-shot learning, and supervised fine-tuning strategies for financial causal extraction. Experimental results demonstrate that supervised fine-tuning on combined English and Spanish data substantially enhances cross-lingual transfer performance. Notably, the fine-tuned GPT-4.1 Mini model achieves joint top rank on the English subtask (4.8140) and third place on the Spanish subtask (4.7753), underscoring the effectiveness of task-specific fine-tuning.
📝 Abstract
This paper describes team HSA_CORAL's submission to the FinCausal 2026 shared task on extracting cause-effect relations from financial narratives via extractive question answering in English and Spanish. We compare three modeling families: (i) encoder-only token tagging with multilingual BERT, (ii) encoder-decoder generation with multilingual BART, and (iii) decoder-only LLMs (Llama 3.1 and GPT variants) using prompt refinement, few-shot demonstrations, and supervised fine-tuning. Across settings, prompting and few-shot examples yield competitive performance, while supervised fine-tuning provides the largest gains. Our best system, GPT-4.1 Mini fine-tuned on combined English and Spanish training data, achieves a tied highest score on the English subtask (score 4.8140) and ranks third on Spanish (score 4.7753) under the shared task's LLM-as-a-judge metric. Overall, the results highlight the value of task-specific adaptation and multilingual fine-tuning for cross-lingual transfer in financial causality QA.