Score
Designs, builds, or evaluates systems that determine the language (and script or closely related dialect) of a given text or speech segment, producing language labels and confidence scores. Work includes handling monolingual and code-switched inputs, choosing features or models for short or noisy inputs, and diagnosing error modes.
Large language and speech models suffer from poor generalization to minority languages, dialects, and sociolinguistic variants due to skewed training data—exacerbating the digital divide and inequitable technology access. This project pioneers an integrated framework for linguistic justice, systematically bridging computational linguistics, sociolinguistics, and AI ethics. Methodologically, it combines corpus analysis, bias quantification, low-resource modeling, and interdisciplinary qualitative research. Key contributions include: (1) a reproducible, multi-dimensional evaluation metric suite for linguistic inclusivity; (2) a consensus-based white paper and 12 actionable, industry-deployable recommendations; and (3) an open-source dialect adaptation toolkit to advance standardized technical support for linguistic diversity. Collectively, these outcomes address the structural misalignment between model capabilities and real-world linguistic heterogeneity, offering both methodological rigor and practical pathways toward equitable AI.
This study addresses the limited effectiveness of automated code review in industrial C# projects. We propose a monolingual (C#-specific) supervised fine-tuning approach to systematically enhance model performance across three core tasks: code change quality assessment, review comment generation, and code improvement suggestion. Leveraging a curated C#-dedicated dataset, we fine-tune CodeReviewer, CodeLlama-7B, and DeepSeek-R1-Distill, validating results on both enterprise codebases and public benchmarks. Our key contribution is the empirical revelation that alignment between programming language semantics and natural language representations critically governs model capability—highlighting the synergistic importance of linguistic consistency and task-specific adaptation. Experiments demonstrate that monolingual fine-tuning significantly improves output accuracy and relevance, enabling near-parity with or partial superiority over static analysis tools on routine tasks. However, performance gaps persist against human reviewers in semantically complex, context-intensive scenarios requiring deep program understanding.
This study addresses the lack of empirical research on code-switching mechanisms (Spanish/English) in human–machine bilingual dialogue. We developed a task-oriented chatbot and conducted controlled map-task experiments to systematically evaluate distinct code-switching strategies—namely, rule-driven, grammatically compliant, and predictable versus random or ungrammatical switching. Through human-participant interaction studies, we provide the first empirical evidence that grammatical well-formedness and pattern predictability of code-switching significantly improve task completion efficiency and user experience; users consistently prefer structured, linguistically constrained language mixing. Our findings establish critical empirical foundations for designing multilingual AI dialogue systems, underscoring that controllability and linguistic plausibility of code-switching are essential for effective human–AI collaboration.
This study empirically evaluates large language models (LLMs) against industry-standard technical hiring assessments for algorithm and software engineering roles. Method: We administered realistic, industrial-grade programming, system design, and reasoning questions—commonly used by leading technology firms—to state-of-the-art LLMs (e.g., GPT-4, Claude 3, Gemini) and conducted multi-stage comparative analysis against official corporate reference solutions, assessing correctness, completeness, engineering soundness, and consistency. Contribution/Results: Our analysis reveals systematic structural gaps between LLM outputs and industrial expectations: no tested model met enterprise hiring thresholds. Critical deficiencies were observed in boundary-case handling, explicit modeling of resource constraints (e.g., time/space complexity, scalability), and maintainability-aware design. These findings challenge the prevailing assumption that LLMs can directly substitute for entry-level engineers. Moreover, this work introduces the first benchmark framework specifically tailored to industrial recruitment scenarios, providing empirically grounded insights for AI capability evaluation in real-world engineering hiring.
This study investigates whether professional translators without specialized training can reliably distinguish AI-generated Italian short stories from human-authored ones. In an offline experiment, 69 translators evaluated three anonymized texts—two produced by ChatGPT-4o and one written by a human—assessing their origin and providing justifications. As the first empirical examination of AI-text detection among a real-world cohort of professional translators, the research integrates quantitative scoring with qualitative analysis to uncover the dual influence of analytical reasoning and subjective preferences on identification accuracy. Findings reveal that only 16.2% of participants performed significantly above chance level. Effective diagnostic cues included low burstiness, narrative inconsistencies, and traces of English-language transfer, whereas high grammatical accuracy and emotional tone frequently led to misclassification.
This work addresses the frequent failure of natural language processing (NLP) projects in clinical settings, which often stems from a lack of systematic engineering practices and an overemphasis on algorithms at the expense of development rigor. To bridge this gap, the paper proposes a structured methodology grounded in the Systems Development Life Cycle (SDLC) framework to guide the end-to-end construction of NLP systems for extracting clinical information from electronic health records. By integrating SDLC principles throughout the NLP development pipeline, the approach counteracts the algorithm-centric bias prevalent in conventional tutorials and establishes a reproducible, generalizable development paradigm. This systematic integration enhances both the success rate and reliability of clinical text information extraction initiatives, offering a robust foundation for real-world deployment.
This work addresses the inconsistent detection and mitigation of toxic content by multilingual large language models across diverse linguistic and cultural contexts. It presents the first systematic synthesis of research on multilingual toxicity handling, proposing a comprehensive framework that encompasses threat modeling, task formulation, detection strategies—such as cross-lingual encoders, translation pipelines, and representation probing—and mitigation approaches, including data filtering, alignment tuning, decoding controls, and multilingual safeguards. The study identifies core challenges such as uneven language coverage and culturally contingent definitions of harm, while highlighting critical issues like fragmented evaluation protocols and the unintended suppression of legitimate expression. By elucidating these dimensions, the paper establishes a theoretical foundation and practical roadmap for achieving cross-lingual safety alignment in multilingual language models.
This study investigates the impact of multilingual prompts on code generation quality and adherence to programming conventions in large language models, revealing underlying linguistic biases. The authors construct the first high-quality multilingual programming benchmark encompassing Chinese, English, Hindi, Spanish, and Italian, featuring expert human-translated prompts and a multidimensional evaluation framework that includes unit tests, code metrics, static analysis, and lexical features. Using this benchmark, they systematically assess the performance of GPT-4o mini, DeepSeek, and Claude on Python and Java tasks. Their findings indicate that the effect of prompt language on code quality is contingent upon both the target programming language and the model architecture; notably, English prompts do not consistently yield superior results. Moreover, generated code frequently exhibits code-switching, with comments and string literals often mixing the prompt language and English.
This study addresses the limited support of current large code models and their evaluation methodologies for non-English natural language elements, such as code comments. The authors systematically evaluate five prominent models—including CodeGemma and CodeLlama—on comment generation across Dutch, English, Greek, Polish, and Chinese. They introduce the first multilingual code comment dataset comprising 12,500 human-annotated samples and propose a fine-grained error taxonomy encompassing 26 error categories. Their findings reveal a substantial degradation in comment quality for non-English languages, with linguistic errors increasing by up to 15.1×. Moreover, existing automatic evaluation methods, including neural metrics and LLM-as-a-judge approaches, prove unreliable in detecting linguistic and semantic inaccuracies, underscoring the irreplaceable role of human judgment in evaluating multilingual code generation.