dialect-aware evaluation

Designs and implements evaluation methods, annotation guidelines, and analysis procedures that measure and interpret speech- and voice-based system performance while accounting for dialectal and culturally specific linguistic and paralinguistic variation. Builds dialect-aware test sets, scoring metrics, and interpretive frameworks to detect misrecognition, bias, or culturally significant vocal cues and to support culturally informed interpretation of model outputs.

dialect-awareevaluation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.13
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Culturally-Aware Conversations: A Framework & Benchmark for LLMs

Oct 13, 2025
SH
Shreya Havaldar
🏛️ University of Pennsylvania

Existing LLM cultural adaptation benchmarks lack ecological validity, failing to reflect authentic cross-cultural interaction scenarios. Method: We propose CulturaBench—the first evaluation framework grounded in sociocultural theory—centered on the dynamic evolution of linguistic style across situational, relational, and cultural contexts. It innovatively introduces three core dimensions for cross-cultural NLP assessment: conversational framing, stylistic sensitivity, and subjective correctness. A diverse, multicultural annotator cohort co-constructed a high-ecological-validity benchmark dataset, featuring multi-dimensional human annotations and contextualized dialogue test cases. Contribution/Results: Empirical evaluation reveals systematic deficiencies in mainstream LLMs regarding dynamic stylistic adaptation and implicit cultural norm comprehension. CulturaBench establishes a theory-driven, reproducible evaluation paradigm and benchmark resource for culturally aware dialogue systems.

Current LLMs struggle with cultural adaptation during conversationsExisting cultural benchmarks misalign with real LLM interaction challengesIntroducing first framework evaluating LLMs in multicultural conversational settings

This study systematically evaluates racial bias in four major commercial automatic speech recognition (ASR) systems. Using the Pacific Northwest English Corpus—featuring speakers from African American, White, Chicano, and Yakama communities—the authors quantify cross-ethnic transcription disparities via a novel phoneme error rate (PER) metric integrated with sociophonetic annotations. Results reveal significantly higher PER for African American speakers across all systems; critically, all models substantially underrepresent dialectal phenomena such as low-vowel mergers, confirming that inadequate acoustic modeling of sociophonetic variation constitutes the primary source of bias. The study introduces an analytical framework linking PER to fine-grained sociophonetic features, identifying vowel quality variation as a key determinant of performance disparity. These findings underscore the necessity of incorporating dialectal diversity into ASR training and evaluation to advance fairness and robustness in speech technology.

Analyzing how sociophonetic variation causes differential ASR performanceEvaluating racial bias in commercial ASR systems across ethnic groupsIdentifying phonetic variation as primary source of ASR bias

Existing cultural competence evaluations predominantly rely on decontextualized correctness judgments, failing to capture the depth of understanding and reasoning that large language models (LLMs) exhibit in authentic, multicultural contexts. To address this, we propose a “thick culture” evaluation paradigm—a context-aware, scenario-based framework for cultural assessment that emphasizes situationally grounded response generation. We introduce four fine-grained, low-variance metrics—coverage, specificity, semantic depth, and coherence—to quantify cultural reasoning rigorously. Integrating contextualized benchmarks with multidimensional automated evaluation, our experiments reveal that conventional “thin” evaluation significantly overestimates model capabilities and yields high result variance. In contrast, our framework stably discriminates subtle differences in cultural understanding across state-of-the-art models, delivering more reliable, interpretable, and actionable evaluation signals.

Addressing limitations of de-contextualized cultural correctness assessmentsDeveloping thick evaluation methods for cultural understanding and reasoningEvaluating cultural competence in LLMs deployed across diverse environments

This work addresses the limitations of traditional sociodemographic prompting (SDP) in evaluating cultural alignment of large language models, as SDP is susceptible to confounding factors such as prompt sensitivity, decoding parameters, and task complexity, making it difficult to disentangle model bias from flaws in task design. To overcome these issues, the authors propose Inverse Sociodemographic Prompting (ISDP), which reframes cultural alignment assessment from a generation task into a discrimination task by prompting models to identify user groups based on real or simulated user behaviors. Experiments on the Goodreads-CSI dataset with models including Aya-23, Gemma-2, GPT-4o, and LLaMA-3.1 show that while models generally perform better on real user behaviors, their individual-level discriminative performance converges, revealing a significant bottleneck in current large language models’ capacity for deep, personalized cultural understanding.

biascultural alignmentgeneration vs. discrimination

This work proposes a multilingual framework for automatic intelligibility assessment of dysarthric speech, addressing the limitations of existing approaches that are often confined to a single language and struggle to model language-specific factors. The framework integrates universal phoneme recognition with language-specific phonemic mappings derived from contrastive phonological features, and incorporates sequence alignment to generate multidimensional intelligibility metrics. Notably, it introduces Phoneme Coverage (PhonCov)—a novel, alignment-free metric—that, together with Phone Error Rate (PER) and Phone Feature Error Rate (PFER), forms a comprehensive evaluation suite. Experiments on English, Spanish, Italian, and Tamil demonstrate that the framework effectively captures clinically relevant patterns of intelligibility degradation, with individual metrics benefiting differentially from phonemic mapping, alignment, or their combination.

dysarthriaintelligibility assessmentlanguage-specific factors

Latest Papers

What's happening recently
View more

This study addresses the pervasive lack of cultural competence in contemporary multilingual NLP models, which often fail to accurately interpret expressions deeply rooted in specific cultural contexts despite broad language coverage. Synthesizing insights from over 50 studies published between 2020 and 2026, the work advocates a paradigm shift from isolated language processing toward modeling the “communicative ecology,” integrating institutional norms, cultural scripts, and community practices as essential contextual dimensions. Through culturally aware evaluation benchmarks (e.g., Global-MMLU, CulturalBench), multimodal grounding of local knowledge, community-coconstructed datasets, and cultural alignment techniques, the research demonstrates that insufficient training data coverage is not the sole bottleneck—language choice, tokenization strategies, and translation benchmark design are equally critical. The paper calls for layered cultural evaluation frameworks and participatory alignment approaches to advance fair, inclusive, and culturally grounded NLP systems.

community-grounded evaluationcross-lingual transfercultural competence

This study addresses the lack of objective and interpretable vocal biomarkers in mental health assessment by proposing a transparent, clinically interpretable speech analysis framework. It systematically integrates multidimensional perceptual features—including prosody, voice quality, semantic coherence, syntactic structure, and sarcasm—by jointly leveraging acoustic and linguistic information. Using an XGBoost model enhanced with SHAP and LIME for interpretability, the framework identifies key features such as jitter, shimmer, lexical-syntactic patterns, and affective intonation from real-world clinical data and multiple benchmark datasets. Experimental results demonstrate robust associations between vocal irregularities and symptom severity in depression, anxiety, and ADHD. Ablation studies further confirm the most discriminative feature subsets, offering reliable and explainable vocal biomarkers to support clinical evaluation.

clinical decision-supportmental health assessmentperceptual speech features

This study addresses a critical gap in the evaluation of automatic speech recognition (ASR) systems by moving beyond conventional error-rate metrics to examine the emotional burden and adaptive costs users incur when ASR fails. Through qualitative research involving field interviews and open-ended narrative analysis across four U.S. dialect communities, the work integrates sociolinguistic and human-computer interaction theories to systematically investigate how ASR failures affect users’ emotions, behaviors, and self-perception. Findings reveal widespread experiences of frustration and self-doubt, alongside active linguistic accommodation strategies, underscoring significant deficiencies in cultural alignment and affective fairness. This research pioneers an expanded fairness framework for ASR that incorporates emotional impact, cultural inclusivity, and subjective user experience, thereby challenging the dominant paradigm that prioritizes accuracy alone.

algorithmic fairnessASR biasdialect marginalization

This study addresses the underexplored issue of fairness in phoneme-based automatic speech recognition (ASR) systems across demographic dimensions such as race, age, gender, and accent. It presents the first systematic evaluation of group-level biases in two open-source IPA transcription models—WhisperIPA and ZIPA—using multilingual, demographically annotated corpora. To account for linguistically acceptable phonemic variation, the authors introduce a novel metric, Soft PER (Phoneme Error Rate), which relaxes strict phoneme matching. Experimental results demonstrate that even when accommodating such permissible variation, both models exhibit significant performance disparities across languages, genders, accents, ethnicities, and age groups. These findings underscore persistent fairness challenges in current IPA-based ASR systems and highlight the need for more equitable model development and evaluation practices.

automatic speech recognitionbiasdemographic disparity

This study addresses the lack of comprehensive evaluation frameworks for cross-lingual text-to-speech (TTS) systems in low-resource languages that jointly assess perceptual quality, speaker similarity, and acoustic fidelity. The authors propose a reproducible multi-metric benchmark integrating MUSHRA/ABX subjective listening tests, Resemblyzer-based speaker similarity scores, and objective measures such as Mel-cepstral distortion (MCD) and F0 RMSE. For the first time, four state-of-the-art TTS systems are rigorously evaluated across four distinct speech domains—formal, conversational, literary, and emotional—in a low-resource setting. Results reveal that emotional speech synthesis poses the greatest challenge (average MCD: 12.03 dB), while conversational speech achieves the highest acoustic fidelity, with significant performance variations observed across systems and domains. The complete evaluation toolkit and dataset are publicly released to advance standardized assessment in low-resource TTS research.

acoustic fidelitydomain-specific analysislow-resource languages

Hot Scholars

MD

Muhammad Dehan Al Kautsar

Mohamed bin Zayed University of Artificial Intelligence
Natural Language ProcessingMultilingualityHuman-Centered NLP
FK

Fajri Koto

Assistant Professor (tenure-track), MBZUAI
Computational LinguisticsNatural Language ProcessingMultilingual NLPHuman-centered NLP
RK

Ratna Kandala

Postdoc, University of Kansas
SyntaxNatural Language ProcessingComputational LinguisticsCognitive Science
PH

Pan Hui

Chair Professor, Nokia Chair in Data Science, FREng & IEEE Fellow (HKUST & University of Helsinki)
Ubiquitous ComputingMobile ComputingAugmented RealityData Science
CR

Chu-Ren Huang

Chair Professor, The Hong Kong Polytechnic University
computational linguisticscorpus linguisticsChinese linguisticslexical semantics