🤖 AI Summary
General-purpose sentence embedding models struggle with financial domain terminology, semantic drift over time, and bilingual lexical misalignment in low-resource languages like Korean. Method: We propose the NMIXX model family and introduce KorFinSTS—the first Korean–English financial cross-lingual semantic textual similarity benchmark. Building upon the mbedding contrastive learning framework, we perform fine-grained domain adaptation on multilingual BGE-M3 using 18.8K high-quality triplets (including domain-specific paraphrases, hard negatives, and exact translations), and uniquely incorporate semantic drift typology to model financial semantic dynamics. Contribution/Results: Our approach achieves +0.10 and +0.22 improvements in Spearman correlation over prior open-source models on FinSTS and KorFinSTS, respectively—demonstrating state-of-the-art performance. Both the NMIXX models and the KorFinSTS benchmark are publicly released to advance research in financial cross-lingual representation learning.
📝 Abstract
General-purpose sentence embedding models often struggle to capture specialized financial semantics, especially in low-resource languages like Korean, due to domain-specific jargon, temporal meaning shifts, and misaligned bilingual vocabularies. To address these gaps, we introduce NMIXX (Neural eMbeddings for Cross-lingual eXploration of Finance), a suite of cross-lingual embedding models fine-tuned with 18.8K high-confidence triplets that pair in-domain paraphrases, hard negatives derived from a semantic-shift typology, and exact Korean-English translations. Concurrently, we release KorFinSTS, a 1,921-pair Korean financial STS benchmark spanning news, disclosures, research reports, and regulations, designed to expose nuances that general benchmarks miss.
When evaluated against seven open-license baselines, NMIXX's multilingual bge-m3 variant achieves Spearman's rho gains of +0.10 on English FinSTS and +0.22 on KorFinSTS, outperforming its pre-adaptation checkpoint and surpassing other models by the largest margin, while revealing a modest trade-off in general STS performance. Our analysis further shows that models with richer Korean token coverage adapt more effectively, underscoring the importance of tokenizer design in low-resource, cross-lingual settings. By making both models and the benchmark publicly available, we provide the community with robust tools for domain-adapted, multilingual representation learning in finance.