🤖 AI Summary
This study addresses the significant limitations of large language models in understanding Vietnamese figurative expressions—such as idioms and proverbs—that are deeply rooted in cultural context. To this end, the authors introduce VIVID, the first systematic evaluation benchmark for Vietnamese figurative language, comprising 1,636 annotated expressions labeled for complexity and semantic themes. They propose an evaluation framework integrating generative and discriminative tasks, augmented with a human-validated, aspect-level LLM-as-a-Judge mechanism (Cohen’s κ = 0.792). Experiments on eight state-of-the-art models reveal that Vietnamese-specific models substantially underperform multilingual counterparts (e.g., VinaLLaMA-7B scores 0.13 versus GPT-4o’s 2.46 on a 5-point scale), with all models scoring below 50% of the maximum. Notably, few-shot prompting degrades GPT-4o’s performance due to stylistic overfitting, underscoring a systemic deficiency in cultural pragmatic comprehension.
📝 Abstract
We present VIVID (Vietnamese Idioms for Validation and Interpretation Depth), the first systematic benchmark for evaluating culturally grounded figurative language understanding in Vietnamese. VIVID comprises 1,636 idioms and proverbs annotated with five complexity traits (literal expressions, pragmatic nuances, Sino-Vietnamese terms, uncommon vocabulary, folk knowledge) and seven semantic themes. We establish an evaluation framework combining generative and discriminative tasks, proposing an LLM-as-a-Judge approach with aspect-based prompting validated against human judgment (Cohen's kappa = 0.792). Evaluating eight state-of-the-art models reveals critical gaps: Vietnamese-specialized models drastically underperform multilingual systems (VinaLLaMA-7B: 0.13 vs. GPT-4o: 2.46), and even top models achieve less than 50% of maximum scores. Notably, few-shot prompting does not universally improve performance, with GPT-4o exhibiting degradation due to stylistic overfitting. Our analysis exposes systematic failures including literal over-interpretation, lexical gaps, and pragmatic flattening, demonstrating that current models lack cultural competence for nuanced figurative interpretation. VIVID provides an essential tool for advancing figurative language understanding in culturally rich contexts.