Score
Designs, builds, or analyzes software components and pipelines that correctly process, transform, and validate text encoded in Unicode — including encoding/decoding (UTF-8/16/32), handling code points, code units and surrogate pairs, grapheme-cluster segmentation and combining marks, bidirectional text, normalization (NFC/NFD/NFKC/NFKD), locale-aware case mapping and collation, and sanitization for storage, search, and display.
Byte-Pair Encoding (BPE) tokenizers suffer from two critical limitations in multilingual settings: (i) encoding penalties for non-Latin scripts due to UTF-8 byte fragmentation, and (ii) reduced robustness stemming from heuristic regular-expression-based preprocessing. Method: We propose SCRIPT, a structured pre-tokenization framework grounded in Unicode Script and General Category properties. SCRIPT replaces byte-level BPE with script-boundary-aware, rule-based pre-segmentation and enforces constrained BPE merges that preserve character integrity and eliminate cross-script encoding bias. Contribution/Results: Empirical evaluation shows that SCRIPT-BPE achieves token compression rates comparable to standard BPE while completely eliminating encoding penalties for non-Latin languages. Moreover, it significantly improves tokenization robustness—especially under noisy or malformed input—and enhances cross-lingual fairness by ensuring consistent, script-aware segmentation across diverse writing systems.
Byte-level subword tokenizers may generate invalid UTF-8 byte sequences, compromising the validity of large language model outputs, system stability, and security. Method: We formally model tokenization as a monoid operation—its first such theoretical treatment—and rigorously prove that if the vocabulary contains invalid UTF-8 substrings, any decoding (especially incremental decoding) can produce invalid byte sequences, and inconsistency between incremental and full-sequence decoding is inevitable. Through structural analysis of UTF-8 encoding, tokenizer behavior simulation, and empirical evaluation across mainstream models (e.g., Llama, Qwen) and inference engines (e.g., vLLM, Ollama), we identify this flaw in multiple production systems. Contribution/Results: We propose two mitigation strategies—vocabulary sanitization and runtime UTF-8 validation—achieving 100% detection of invalid sequences without performance degradation. Our open-source verification tool enables systematic auditing and hardening of tokenizer deployments.
Large language models (LLMs) exhibit pervasive output formatting bias in code translation tasks—generated outputs frequently contain extraneous natural-language explanations or formatting delimiters, causing standard evaluation metrics (e.g., computation accuracy, CA) to systematically underestimate true performance. Method: We systematically evaluate 11 instruction-tuned LLMs across five programming languages and find that 26.4%–73.7% of translations require post-hoc processing to extract clean code. To address this, we propose a robust code extraction method integrating regex-based parsing with prompt engineering. Contribution/Results: Our approach achieves a 92.73% average Code Extraction Success Rate (CSR) on a multilingual alignment benchmark, substantially improving evaluation fidelity. This work is the first to quantify the impact of formatting bias and establishes a new, generalizable, and robust code extraction paradigm—providing a reproducible, standardized evaluation benchmark for LLM-based code translation.
The underlying mechanisms of code generation errors in large language models (LLMs) remain poorly understood. Method: Leveraging the HumanEval benchmark, this work systematically analyzes errors produced by six state-of-the-art LLMs and introduces, for the first time, a multidimensional, fine-grained error taxonomy integrating both semantic and syntactic dimensions. Using open coding and thematic analysis—augmented by statistical testing and qualitative root-cause attribution—the study identifies over ten recurrent error patterns, including logical flaws, boundary condition failures, and API misuse. Contribution/Results: The analysis reveals that LLM errors exhibit nontriviality, cross-line dependencies, and dispersed distribution—uncovering latent, deep-seated errors even in high-pass-rate tasks. It further demonstrates a nonlinear positive correlation between error frequency and task complexity. This taxonomy provides an interpretable, extensible theoretical foundation and empirical grounding for error localization, diagnosis, and repair in LLM-generated code.
Traditional Byte-Pair Encoding (BPE) tokenization introduces token redundancy in low-resource languages, degrading the performance of small-scale models. Method: This paper proposes a BPE configuration method integrating hyperparameter optimization and compressed sensing. It systematically searches key BPE hyperparameters—including vocabulary size and merge iterations—and jointly evaluates configurations using intrinsic metrics (e.g., token count) and extrinsic task performance (generation and classification). Contribution/Results: The study provides the first empirical evidence that BPE configuration significantly impacts multilingual modeling for low-resource languages. Experiments across diverse languages and model scales show that optimal configurations reduce token counts by 12.7% on average and improve downstream task accuracy by 1.8–3.4 percentage points for small models. These gains substantially enhance modeling efficiency and generalization capability in low-resource settings.
This work addresses the challenge that byte-level language models often generate invalid UTF-8 sequences when producing rare or previously unseen characters, thereby compromising the reliability of multilingual text generation. The authors train a 355-million-parameter, byte-level language model on 80 billion tokens of multilingual data and introduce a perplexity-independent evaluation protocol to assess structural validity of UTF-8 sequences. Their findings reveal that achieving convergence in UTF-8 validity requires approximately twice as much training data as needed for perplexity stabilization—42 billion versus 21 billion tokens. Notably, in context-free generation, rare characters exhibit higher structural validity than common ones, challenging conventional assumptions in representation learning. These results demonstrate that reliable UTF-8 sequence generation constitutes a distinct modeling capability beyond mere perplexity minimization.
This work addresses the limitations of existing Unicode code point–based text evaluation methods, which often fail to accurately measure character-level errors in complex writing systems where a single grapheme frequently comprises multiple code points. To overcome this, the authors introduce grapheme-kit, an open-source Python library that, for the first time, extends widely used NLP evaluation metrics—such as edit distance and similarity—to the grapheme cluster level. By adhering to Unicode standards for grapheme cluster identification, composition, and decomposition, the proposed approach significantly improves evaluation accuracy for tasks like OCR on scripts with complex orthographies, including Tamil and Sinhala. This advancement provides a precise, grapheme-aware toolkit for text processing in low-resource languages.
This study addresses the excessive token overhead and constrained context windows caused by UTF-8 encoding for non-English scripts in multilingual large language models. To this end, we propose a universal byte-level encoding scheme featuring a novel dual-alphabet dynamic routing mechanism that hybridizes UTF-8 and UTF-16 encodings to optimize multi-byte character processing. Crucially, this method modifies only the underlying byte representations without altering BPE merge rules, thereby achieving lossless and efficient cross-script encoding. Experimental results demonstrate that our approach substantially reduces token counts for high-premium scripts and enhances effective context utilization. Furthermore, it accelerates inference while preserving model quality, effectively mitigating cross-lingual performance disparities.
This study addresses the limited support of current large code models and their evaluation methodologies for non-English natural language elements, such as code comments. The authors systematically evaluate five prominent models—including CodeGemma and CodeLlama—on comment generation across Dutch, English, Greek, Polish, and Chinese. They introduce the first multilingual code comment dataset comprising 12,500 human-annotated samples and propose a fine-grained error taxonomy encompassing 26 error categories. Their findings reveal a substantial degradation in comment quality for non-English languages, with linguistic errors increasing by up to 15.1×. Moreover, existing automatic evaluation methods, including neural metrics and LLM-as-a-judge approaches, prove unreliable in detecting linguistic and semantic inaccuracies, underscoring the irreplaceable role of human judgment in evaluating multilingual code generation.
This study addresses the cross-platform display inconsistencies and retrieval failures in Yoruba caused by the absence of precomposed characters in Unicode, highlighting the structural limitations of existing NFC normalization strategies. By combining empirical analysis with multi-platform text processing techniques, this work systematically documents rendering failures in digital environments. Moving beyond conventional approaches that rely solely on combining character sequences, it proposes submitting a formal request to the Unicode Consortium for the direct encoding of essential characters. The primary contribution lies in identifying the structural root cause of this technical barrier and providing an actionable, fundamental encoding remediation to advance the digital representation of African languages.