Score
Designs and executes procedures to verify that translated or otherwise language-adapted textual content preserves the original meaning, conceptual equivalence, and cultural appropriateness across language versions. Builds and runs linguistic QA checks — such as proofreading, consistency and terminology checks, reconciliation of translation discrepancies, and final language-version verification — and documents defects and resolutions.
To address frequent grammatical and spelling errors in automatic sports news generation and the high cost of manual proofreading, this paper proposes a lightweight, plug-and-drop large language model (LLM)-based error correction module. For the first time, it integrates correction as a question-answering (QA) component into NLG pipelines, enabling real-time, multilingual grammar and spelling correction without fine-tuning the primary generative model. The approach balances accuracy and deployment efficiency while ensuring minimal system intrusion. We design a unified multilingual evaluation framework and validate performance on English, Spanish, and German sports news drafts. Experiments show that corrected outputs meet practical acceptability standards: critical error correction rate reaches 92.3%, and average manual proofreading time decreases by 67%. This work establishes a low-intrusion, highly compatible post-editing paradigm for NLG systems.
This study addresses the loss of student responses in multilingual assessment systems due to translation failures, which compromises reliability auditing. To mitigate this issue, the authors propose bypassing translation altogether by directly employing native-language multilingual sentence embeddings instead of translating responses into English. The approach is evaluated on 11 PIRLS constructed-response items using three state-of-the-art multilingual embedding models. Results demonstrate, for the first time, that native-language embeddings can reproduce translation-based reliability estimates without relying on machine translation, successfully recover responses previously excluded due to translation errors, and maintain comparable reliability metrics with no statistically significant degradation. This method enhances both the completeness and robustness of multilingual educational assessments.
Non-standard linguistic phenomena in user-generated content (UGC)—including misspellings, slang, character repetition, and emojis—undermine the consistency and fairness of translation quality evaluation, hindering equitable assessment of models and metrics. Method: We introduce the first fine-grained taxonomy encompassing 12 types of non-standard phenomena and 5 translation operations; propose a novel “guideline-aware controllable evaluation” paradigm that explicitly aligns translation objectives with human-authored evaluation guidelines; and conduct guideline analysis, LLM prompt sensitivity experiments, qualitative modeling, and cross-dataset validation. Contribution/Results: We demonstrate that LLM-based scores are highly sensitive to explicit instruction—when prompts conform to dataset-specific guidelines, BLEU and COMET scores improve by up to 12.3%. This work establishes a consensus evaluation framework anchored in translation guidelines, advancing standardized, interpretable, and guideline-grounded assessment for UGC translation.
This study addresses the uncontrolled quality of quality engineering (QE) artifacts—such as requirements specifications, test cases, and Behavior-Driven Development (BDD) scenarios—automatically generated by large language models (LLMs). We propose an iterative optimization framework integrating forward generation, backward generation, and rubric-guided scoring to enhance artifact quality along four dimensions: clarity, completeness, consistency, and testability. Our approach enables automated, quantitative, and reproducible quality assessment and improvement. Evaluated across 12 real-world projects, the method significantly improves output stability: it preserves high quality under high-quality inputs and substantially outperforms baselines under low-quality inputs. The core contribution is the first integration of backward generation with structured rubric-based guidance, establishing a closed-loop, artifact-centric quality enhancement paradigm for QE.
This work addresses inconsistencies arising from structural mismatches between natural language and formal languages during requirements formalization. It proposes a “consistency through formalization” principle, mandating strict logical alignment among natural language, the structured language FRETish, and the formal temporal logic MTL. Guided by this principle, the authors refine the FRETish-to-MTL translation pipeline in NASA’s FRET tool. Their approach uniquely integrates cross-layer consistency constraints into a collaborative framework combining large language models and formal verification tools. This integration not only uncovers and corrects multiple inconsistencies in the original translation but also demonstrates superior correctness and reliability, as substantiated by formal equivalence proofs and empirical evaluation.
This work addresses the limitations of traditional translation validation methods, which treat individual translations in isolation and struggle to assess trustworthiness within heterogeneous, multilingual, multi-path translation graphs. The authors propose a fidelity-graded translation calculus that models translations as composable, verifiable graph structures, dynamically establishing trust levels through inline validation and multi-path consistency checks, thereby enabling end-to-end self-certifying answers. Built upon a Lean 4 mechanized core, the hurdy-gurdy platform integrates 13 languages and 13 translator pairs, leveraging LLM agents to generate code and cross-validate semantic correctness. Experimental results demonstrate that this architecture substantially enhances scalability and trustworthiness—measured by joint coverage, branch consistency, certified unreachability, and escape rate—while decoupling trustworthiness from authorship.
Current automatic Machine Translation Quality Estimation (QE) systems lack reliability in real-world scenarios because their segment-level evaluations neglect critical dimensions such as discourse coherence, stylistic consistency, and rhetorical adequacy. Through theoretical analysis and empirical investigation, this study systematically uncovers structural limitations in QE—particularly concerning generalization capacity, data bias, overfitting, annotation noise, and the modeling of linguistic complexity—and, for the first time, identifies an inherent bottleneck rooted in the cognitive foundations of translation itself. These findings challenge the prevailing assumption that QE performance can be sufficiently improved merely by scaling up data or model capacity. The paper cautions against deploying existing QE systems as the sole basis for decision routing or bypassing human review in production environments and advocates redirecting future research toward automating human evaluation grounded in the Multidimensional Quality Metrics (MQM) framework.
This study investigates the impact of prompt design—integrating principles from fusion translation theory and varying prompt languages—on the quality of Spanish-to-Chinese news translation by large language models. Using GPT-5.2, the authors evaluate performance across 48 experimental conditions (comprising four prompt types, three prompt languages, and four editorial texts) through automatic metrics (BLEU and BERTScore-F1) and multidimensional human assessment via MQM. For the first time, translation theory is explicitly incorporated into prompt engineering, yielding significant improvements in expert human ratings, particularly in reducing “awkward style” errors. Results indicate that the BRIEF prompt achieves the highest MQM score (8.66 versus 7.84 for BASE), while the choice of prompt language exerts negligible influence, underscoring the critical role of theory-driven prompting in enhancing stylistic fluency in machine translation.
本文提出一种两阶段MQM引导的自动后编辑框架,通过诊断和修复提高特定领域机器翻译质量,优于单阶段方法。
This work addresses the challenge that existing proofreading tools struggle to simultaneously ensure factual accuracy, domain-specific technical correctness, and linguistic quality in educational textbooks. To this end, we propose AI Textbook Auditor—the first modular multi-agent system designed for comprehensive textbook auditing. The system operates dual parallel pipelines: one for fact and technical verification, and another for native PDF-based grammatical analysis, while a referee agent filters false positives to produce structured review reports. Integrating domain-customized prompt engineering, vision-aware PDF parsing (via PyMuPDF), web-augmented fact-checking, and rule-based false-positive suppression, the framework supports cross-disciplinary error categorization and native PDF processing. Evaluated on Romanian high school textbooks in computer science and history/social sciences, it identified 56 and 72 issues respectively, achieving an expert-validated precision of 62.5% and significantly enhancing manual review efficiency.