Score
Designs and implements evaluation methods, annotation guidelines, and analysis procedures that measure and interpret speech- and voice-based system performance while accounting for dialectal and culturally specific linguistic and paralinguistic variation. Builds dialect-aware test sets, scoring metrics, and interpretive frameworks to detect misrecognition, bias, or culturally significant vocal cues and to support culturally informed interpretation of model outputs.
Current AI systems face ambiguous evaluation in cross-cultural contexts due to the absence of a unified definition of cultural competence, often leading to misleading deployment decisions. This work addresses this gap by drawing on intercultural communication theory to propose, for the first time, a systematic three-tier framework—comprising cultural awareness, cultural sensitivity, and cultural competence—that integrates insights from intercultural communication studies with AI evaluation paradigms. Through theoretical modeling and conceptual analysis, the study formulates an operationalizable assessment framework for cultural competence in AI systems. This construct significantly enhances the validity, interpretability, and deployment safety of AI evaluations in multicultural environments.
Existing LLM cultural adaptation benchmarks lack ecological validity, failing to reflect authentic cross-cultural interaction scenarios. Method: We propose CulturaBench—the first evaluation framework grounded in sociocultural theory—centered on the dynamic evolution of linguistic style across situational, relational, and cultural contexts. It innovatively introduces three core dimensions for cross-cultural NLP assessment: conversational framing, stylistic sensitivity, and subjective correctness. A diverse, multicultural annotator cohort co-constructed a high-ecological-validity benchmark dataset, featuring multi-dimensional human annotations and contextualized dialogue test cases. Contribution/Results: Empirical evaluation reveals systematic deficiencies in mainstream LLMs regarding dynamic stylistic adaptation and implicit cultural norm comprehension. CulturaBench establishes a theory-driven, reproducible evaluation paradigm and benchmark resource for culturally aware dialogue systems.
This study systematically evaluates racial bias in four major commercial automatic speech recognition (ASR) systems. Using the Pacific Northwest English Corpus—featuring speakers from African American, White, Chicano, and Yakama communities—the authors quantify cross-ethnic transcription disparities via a novel phoneme error rate (PER) metric integrated with sociophonetic annotations. Results reveal significantly higher PER for African American speakers across all systems; critically, all models substantially underrepresent dialectal phenomena such as low-vowel mergers, confirming that inadequate acoustic modeling of sociophonetic variation constitutes the primary source of bias. The study introduces an analytical framework linking PER to fine-grained sociophonetic features, identifying vowel quality variation as a key determinant of performance disparity. These findings underscore the necessity of incorporating dialectal diversity into ASR training and evaluation to advance fairness and robustness in speech technology.
Existing cultural competence evaluations predominantly rely on decontextualized correctness judgments, failing to capture the depth of understanding and reasoning that large language models (LLMs) exhibit in authentic, multicultural contexts. To address this, we propose a “thick culture” evaluation paradigm—a context-aware, scenario-based framework for cultural assessment that emphasizes situationally grounded response generation. We introduce four fine-grained, low-variance metrics—coverage, specificity, semantic depth, and coherence—to quantify cultural reasoning rigorously. Integrating contextualized benchmarks with multidimensional automated evaluation, our experiments reveal that conventional “thin” evaluation significantly overestimates model capabilities and yields high result variance. In contrast, our framework stably discriminates subtle differences in cultural understanding across state-of-the-art models, delivering more reliable, interpretable, and actionable evaluation signals.
This work addresses the limitations of traditional sociodemographic prompting (SDP) in evaluating cultural alignment of large language models, as SDP is susceptible to confounding factors such as prompt sensitivity, decoding parameters, and task complexity, making it difficult to disentangle model bias from flaws in task design. To overcome these issues, the authors propose Inverse Sociodemographic Prompting (ISDP), which reframes cultural alignment assessment from a generation task into a discrimination task by prompting models to identify user groups based on real or simulated user behaviors. Experiments on the Goodreads-CSI dataset with models including Aya-23, Gemma-2, GPT-4o, and LLaMA-3.1 show that while models generally perform better on real user behaviors, their individual-level discriminative performance converges, revealing a significant bottleneck in current large language models’ capacity for deep, personalized cultural understanding.
This work proposes a multilingual framework for automatic intelligibility assessment of dysarthric speech, addressing the limitations of existing approaches that are often confined to a single language and struggle to model language-specific factors. The framework integrates universal phoneme recognition with language-specific phonemic mappings derived from contrastive phonological features, and incorporates sequence alignment to generate multidimensional intelligibility metrics. Notably, it introduces Phoneme Coverage (PhonCov)—a novel, alignment-free metric—that, together with Phone Error Rate (PER) and Phone Feature Error Rate (PFER), forms a comprehensive evaluation suite. Experiments on English, Spanish, Italian, and Tamil demonstrate that the framework effectively captures clinically relevant patterns of intelligibility degradation, with individual metrics benefiting differentially from phonemic mapping, alignment, or their combination.
This study addresses the pervasive lack of cultural competence in contemporary multilingual NLP models, which often fail to accurately interpret expressions deeply rooted in specific cultural contexts despite broad language coverage. Synthesizing insights from over 50 studies published between 2020 and 2026, the work advocates a paradigm shift from isolated language processing toward modeling the “communicative ecology,” integrating institutional norms, cultural scripts, and community practices as essential contextual dimensions. Through culturally aware evaluation benchmarks (e.g., Global-MMLU, CulturalBench), multimodal grounding of local knowledge, community-coconstructed datasets, and cultural alignment techniques, the research demonstrates that insufficient training data coverage is not the sole bottleneck—language choice, tokenization strategies, and translation benchmark design are equally critical. The paper calls for layered cultural evaluation frameworks and participatory alignment approaches to advance fair, inclusive, and culturally grounded NLP systems.
This study addresses the lack of objective and interpretable vocal biomarkers in mental health assessment by proposing a transparent, clinically interpretable speech analysis framework. It systematically integrates multidimensional perceptual features—including prosody, voice quality, semantic coherence, syntactic structure, and sarcasm—by jointly leveraging acoustic and linguistic information. Using an XGBoost model enhanced with SHAP and LIME for interpretability, the framework identifies key features such as jitter, shimmer, lexical-syntactic patterns, and affective intonation from real-world clinical data and multiple benchmark datasets. Experimental results demonstrate robust associations between vocal irregularities and symptom severity in depression, anxiety, and ADHD. Ablation studies further confirm the most discriminative feature subsets, offering reliable and explainable vocal biomarkers to support clinical evaluation.
This study addresses a critical gap in the evaluation of automatic speech recognition (ASR) systems by moving beyond conventional error-rate metrics to examine the emotional burden and adaptive costs users incur when ASR fails. Through qualitative research involving field interviews and open-ended narrative analysis across four U.S. dialect communities, the work integrates sociolinguistic and human-computer interaction theories to systematically investigate how ASR failures affect users’ emotions, behaviors, and self-perception. Findings reveal widespread experiences of frustration and self-doubt, alongside active linguistic accommodation strategies, underscoring significant deficiencies in cultural alignment and affective fairness. This research pioneers an expanded fairness framework for ASR that incorporates emotional impact, cultural inclusivity, and subjective user experience, thereby challenging the dominant paradigm that prioritizes accuracy alone.
This study addresses the underexplored issue of fairness in phoneme-based automatic speech recognition (ASR) systems across demographic dimensions such as race, age, gender, and accent. It presents the first systematic evaluation of group-level biases in two open-source IPA transcription models—WhisperIPA and ZIPA—using multilingual, demographically annotated corpora. To account for linguistically acceptable phonemic variation, the authors introduce a novel metric, Soft PER (Phoneme Error Rate), which relaxes strict phoneme matching. Experimental results demonstrate that even when accommodating such permissible variation, both models exhibit significant performance disparities across languages, genders, accents, ethnicities, and age groups. These findings underscore persistent fairness challenges in current IPA-based ASR systems and highlight the need for more equitable model development and evaluation practices.
This study addresses the lack of comprehensive evaluation frameworks for cross-lingual text-to-speech (TTS) systems in low-resource languages that jointly assess perceptual quality, speaker similarity, and acoustic fidelity. The authors propose a reproducible multi-metric benchmark integrating MUSHRA/ABX subjective listening tests, Resemblyzer-based speaker similarity scores, and objective measures such as Mel-cepstral distortion (MCD) and F0 RMSE. For the first time, four state-of-the-art TTS systems are rigorously evaluated across four distinct speech domains—formal, conversational, literary, and emotional—in a low-resource setting. Results reveal that emotional speech synthesis poses the greatest challenge (average MCD: 12.03 dB), while conversational speech achieves the highest acoustic fidelity, with significant performance variations observed across systems and domains. The complete evaluation toolkit and dataset are publicly released to advance standardized assessment in low-resource TTS research.