CultureTalk-ID: A Multi-Task Dialogue Benchmark for Cultural Commonsense in Indonesian Local Languages

📅 2026-07-23
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the lack of contextualized evaluation for Indonesian cultural commonsense, which hinders the capture of nuanced cultural distinctions in real-world discourse. To bridge this gap, we introduce CultureTalk-ID, the first dialogue-based benchmark for Indonesian and eleven regional languages, encompassing 13 cultural themes and comprising 4,496 authentic dialogues collected from native speakers through a multi-stage human-in-the-loop process. We propose a dialogue-centric framework for cultural commonsense evaluation, featuring three complementary tasks—dialogue reasoning, culturally faithful translation, and language-guided generation—to comprehensively assess large language models’ capabilities in understanding, transferring, and generating culturally grounded language. Emphasizing cultural authenticity and contextual integrity, this benchmark establishes a new paradigm for evaluating and enhancing the cultural awareness of large models.
📝 Abstract
Culture is lived through conversation, yet existing Indonesian cultural commonsense benchmarks evaluate LLMs on short and isolated prompts, stripping away the dialogic context in which cultural nuances actually surface. We introduce CultureTalk-ID, the first dialogue-based benchmark for cultural commonsense in Indonesian and its local languages, comprising 4,496 culturally grounded dialogues across 11 languages and 13 culturally salient topics, curated through a multi-stage human pipeline with native speakers to ensure authenticity. CultureTalk-ID introduces three complementary tasks, namely dialogue-based multiple-choice cultural commonsense reasoning, culturally faithful machine translation, and language steering, which jointly probe whether LLMs can understand, transfer, and generate culturally grounded language.
Problem

Research questions and friction points this paper is trying to address.

cultural commonsense
dialogue benchmark
Indonesian local languages
large language models
cultural context
Innovation

Methods, ideas, or system contributions that make the work stand out.

cultural commonsense
dialogue-based benchmark
multi-task evaluation
Indonesian local languages
culturally faithful translation
🔎 Similar Papers
No similar papers found.