linguistic validation

Designs and executes procedures to verify that translated or otherwise language-adapted textual content preserves the original meaning, conceptual equivalence, and cultural appropriateness across language versions. Builds and runs linguistic QA checks — such as proofreading, consistency and terminology checks, reconciliation of translation discrepancies, and final language-version verification — and documents defects and resolutions.

linguisticvalidation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.03
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Integration of LLM Quality Assurance into an NLG System

Jan 27, 2025
CC
Ching-Yi Chen
🏛️ AX Semantics

To address frequent grammatical and spelling errors in automatic sports news generation and the high cost of manual proofreading, this paper proposes a lightweight, plug-and-drop large language model (LLM)-based error correction module. For the first time, it integrates correction as a question-answering (QA) component into NLG pipelines, enabling real-time, multilingual grammar and spelling correction without fine-tuning the primary generative model. The approach balances accuracy and deployment efficiency while ensuring minimal system intrusion. We design a unified multilingual evaluation framework and validate performance on English, Spanish, and German sports news drafts. Experiments show that corrected outputs meet practical acceptability standards: critical error correction rate reaches 92.3%, and average manual proofreading time decreases by 67%. This work establishes a low-intrusion, highly compatible post-editing paradigm for NLG systems.

Automated Text GenerationGrammar CorrectionSports News

This study addresses the loss of student responses in multilingual assessment systems due to translation failures, which compromises reliability auditing. To mitigate this issue, the authors propose bypassing translation altogether by directly employing native-language multilingual sentence embeddings instead of translating responses into English. The approach is evaluated on 11 PIRLS constructed-response items using three state-of-the-art multilingual embedding models. Results demonstrate, for the first time, that native-language embeddings can reproduce translation-based reliability estimates without relying on machine translation, successfully recover responses previously excluded due to translation errors, and maintain comparable reliability metrics with no statistically significant degradation. This method enhances both the completeness and robustness of multilingual educational assessments.

constructed-response itemsLinguistic-Integrated Reliability Auditingmultilingual sentence embeddings

Non-standard linguistic phenomena in user-generated content (UGC)—including misspellings, slang, character repetition, and emojis—undermine the consistency and fairness of translation quality evaluation, hindering equitable assessment of models and metrics. Method: We introduce the first fine-grained taxonomy encompassing 12 types of non-standard phenomena and 5 translation operations; propose a novel “guideline-aware controllable evaluation” paradigm that explicitly aligns translation objectives with human-authored evaluation guidelines; and conduct guideline analysis, LLM prompt sensitivity experiments, qualitative modeling, and cross-dataset validation. Contribution/Results: We demonstrate that LLM-based scores are highly sensitive to explicit instruction—when prompts conform to dataset-specific guidelines, BLEU and COMET scores improve by up to 12.3%. This work establishes a consensus evaluation framework anchored in translation guidelines, advancing standardized, interpretable, and guideline-grounded assessment for UGC translation.

Creating controllable evaluation frameworks aligned with translation guidelinesDeveloping guidelines for handling slang, errors, and emojis in translationEvaluating translation quality for non-standard user-generated content

Technique to Baseline QE Artefact Generation Aligned to Quality Metrics

Nov 18, 2025
EF
Eitan Farchi
🏛️ IBM Research | IBM Consulting

This study addresses the uncontrolled quality of quality engineering (QE) artifacts—such as requirements specifications, test cases, and Behavior-Driven Development (BDD) scenarios—automatically generated by large language models (LLMs). We propose an iterative optimization framework integrating forward generation, backward generation, and rubric-guided scoring to enhance artifact quality along four dimensions: clarity, completeness, consistency, and testability. Our approach enables automated, quantitative, and reproducible quality assessment and improvement. Evaluated across 12 real-world projects, the method significantly improves output stability: it preserves high quality under high-quality inputs and substantially outperforms baselines under low-quality inputs. The core contribution is the first integration of backward generation with structured rubric-based guidance, establishing a closed-loop, artifact-centric quality enhancement paradigm for QE.

Ensuring generated requirements and test cases meet quality metricsEstablishing baselines for automated QE artefact quality evaluationValidating LLM outputs through reverse generation and iterative refinement

This work addresses inconsistencies arising from structural mismatches between natural language and formal languages during requirements formalization. It proposes a “consistency through formalization” principle, mandating strict logical alignment among natural language, the structured language FRETish, and the formal temporal logic MTL. Guided by this principle, the authors refine the FRETish-to-MTL translation pipeline in NASA’s FRET tool. Their approach uniquely integrates cross-layer consistency constraints into a collaborative framework combining large language models and formal verification tools. This integration not only uncovers and corrects multiple inconsistencies in the original translation but also demonstrates superior correctness and reliability, as substantiated by formal equivalence proofs and empirical evaluation.

coherencyformal methodsformalisation

Latest Papers

What's happening recently
View more

This work addresses the limitations of traditional translation validation methods, which treat individual translations in isolation and struggle to assess trustworthiness within heterogeneous, multilingual, multi-path translation graphs. The authors propose a fidelity-graded translation calculus that models translations as composable, verifiable graph structures, dynamically establishing trust levels through inline validation and multi-path consistency checks, thereby enabling end-to-end self-certifying answers. Built upon a Lean 4 mechanized core, the hurdy-gurdy platform integrates 13 languages and 13 translator pairs, leveraging LLM agents to generate code and cross-validate semantic correctness. Experimental results demonstrate that this architecture substantially enhances scalability and trustworthiness—measured by joint coverage, branch consistency, certified unreachability, and escape rate—while decoupling trustworthiness from authorship.

compositional verificationfidelity-graded translationsmulti-language translation graph

Current automatic Machine Translation Quality Estimation (QE) systems lack reliability in real-world scenarios because their segment-level evaluations neglect critical dimensions such as discourse coherence, stylistic consistency, and rhetorical adequacy. Through theoretical analysis and empirical investigation, this study systematically uncovers structural limitations in QE—particularly concerning generalization capacity, data bias, overfitting, annotation noise, and the modeling of linguistic complexity—and, for the first time, identifies an inherent bottleneck rooted in the cognitive foundations of translation itself. These findings challenge the prevailing assumption that QE performance can be sufficiently improved merely by scaling up data or model capacity. The paper cautions against deploying existing QE systems as the sole basis for decision routing or bypassing human review in production environments and advocates redirecting future research toward automating human evaluation grounded in the Multidimensional Quality Metrics (MQM) framework.

Human EvaluationMachine TranslationNatural Language Processing

This study investigates the impact of prompt design—integrating principles from fusion translation theory and varying prompt languages—on the quality of Spanish-to-Chinese news translation by large language models. Using GPT-5.2, the authors evaluate performance across 48 experimental conditions (comprising four prompt types, three prompt languages, and four editorial texts) through automatic metrics (BLEU and BERTScore-F1) and multidimensional human assessment via MQM. For the first time, translation theory is explicitly incorporated into prompt engineering, yielding significant improvements in expert human ratings, particularly in reducing “awkward style” errors. Results indicate that the BRIEF prompt achieves the highest MQM score (8.66 versus 7.84 for BASE), while the choice of prompt language exerts negligible influence, underscoring the critical role of theory-driven prompting in enhancing stylistic fluency in machine translation.

journalistic translationlarge language modelsprompt language

This work addresses the challenge that existing proofreading tools struggle to simultaneously ensure factual accuracy, domain-specific technical correctness, and linguistic quality in educational textbooks. To this end, we propose AI Textbook Auditor—the first modular multi-agent system designed for comprehensive textbook auditing. The system operates dual parallel pipelines: one for fact and technical verification, and another for native PDF-based grammatical analysis, while a referee agent filters false positives to produce structured review reports. Integrating domain-customized prompt engineering, vision-aware PDF parsing (via PyMuPDF), web-augmented fact-checking, and rule-based false-positive suppression, the framework supports cross-disciplinary error categorization and native PDF processing. Evaluated on Romanian high school textbooks in computer science and history/social sciences, it identified 56 and 72 issues respectively, achieving an expert-validated precision of 62.5% and significantly enhancing manual review efficiency.

educational materialsfactual accuracylinguistic quality

Hot Scholars

NH

Nizar Habash

Professor of Computer Science, New York University Abu Dhabi
Natural Language ProcessingComputational LinguisticsArtificial Intelligence
AO

Alice Oh

KAIST Computer Science
machine learningNLPcomputational social science
AF

Alham Fikri Aji

MBZUAI, Monash Indonesia
MultilingualityLow-resource NLPLanguage ModelingMachine Translation
SK

Sylvain Kahane

University Paris Nanterre, Modyco & CNRS / Institut Universitaire de France
syntaxdependency grammartreebankquantitative typology