unicode text processing

Designs, builds, or analyzes software components and pipelines that correctly process, transform, and validate text encoded in Unicode — including encoding/decoding (UTF-8/16/32), handling code points, code units and surrogate pairs, grapheme-cluster segmentation and combining marks, bidirectional text, normalization (NFC/NFD/NFKC/NFKD), locale-aware case mapping and collation, and sanitization for storage, search, and display.

unicodetextprocessing

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.24
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Byte-Pair Encoding (BPE) tokenizers suffer from two critical limitations in multilingual settings: (i) encoding penalties for non-Latin scripts due to UTF-8 byte fragmentation, and (ii) reduced robustness stemming from heuristic regular-expression-based preprocessing. Method: We propose SCRIPT, a structured pre-tokenization framework grounded in Unicode Script and General Category properties. SCRIPT replaces byte-level BPE with script-boundary-aware, rule-based pre-segmentation and enforces constrained BPE merges that preserve character integrity and eliminate cross-script encoding bias. Contribution/Results: Empirical evaluation shows that SCRIPT-BPE achieves token compression rates comparable to standard BPE while completely eliminating encoding penalties for non-Latin languages. Moreover, it significantly improves tokenization robustness—especially under noisy or malformed input—and enhances cross-lingual fairness by ensuring consistent, script-aware segmentation across diverse writing systems.

BPE tokenizers struggle with multilingual script handlingPretokenization fragility due to complex regular expressionsUTF-8 byte conversion issues in token creation

UTF-8 Plumbing: Byte-level Tokenizers Unavoidably Enable LLMs to Generate Ill-formed UTF-8

Nov 05, 2025
PF
Preston Firestone
🏛️ University of Illinois Urbana-Champaign

Byte-level subword tokenizers may generate invalid UTF-8 byte sequences, compromising the validity of large language model outputs, system stability, and security. Method: We formally model tokenization as a monoid operation—its first such theoretical treatment—and rigorously prove that if the vocabulary contains invalid UTF-8 substrings, any decoding (especially incremental decoding) can produce invalid byte sequences, and inconsistency between incremental and full-sequence decoding is inevitable. Through structural analysis of UTF-8 encoding, tokenizer behavior simulation, and empirical evaluation across mainstream models (e.g., Llama, Qwen) and inference engines (e.g., vLLM, Ollama), we identify this flaw in multiple production systems. Contribution/Results: We propose two mitigation strategies—vocabulary sanitization and runtime UTF-8 validation—achieving 100% detection of invalid sequences without performance degradation. Our open-source verification tool enables systematic auditing and hardening of tokenizer deployments.

Byte-level tokenizers enable LLMs to generate invalid UTF-8 sequencesIll-formed UTF-8 output breaks applications expecting valid text encodingIncremental token decoding produces different results than full-sequence conversion

Large language models (LLMs) exhibit pervasive output formatting bias in code translation tasks—generated outputs frequently contain extraneous natural-language explanations or formatting delimiters, causing standard evaluation metrics (e.g., computation accuracy, CA) to systematically underestimate true performance. Method: We systematically evaluate 11 instruction-tuned LLMs across five programming languages and find that 26.4%–73.7% of translations require post-hoc processing to extract clean code. To address this, we propose a robust code extraction method integrating regex-based parsing with prompt engineering. Contribution/Results: Our approach achieves a 92.73% average Code Extraction Success Rate (CSR) on a multilingual alignment benchmark, substantially improving evaluation fidelity. This work is the first to quantify the impact of formatting bias and establishes a new, generalizable, and robust code extraction paradigm—providing a reproducible, standardized evaluation benchmark for LLM-based code translation.

Evaluating LLM code translation suffers from output format biasesNon-code elements in outputs interfere with performance assessment metricsProposing methods to extract source code for reliable model evaluation

Towards Understanding the Characteristics of Code Generation Errors Made by Large Language Models

Jun 13, 2024
ZW
Zhijie Wang
🏛️ University of Alberta | University of Illinois, Urbana-Champaign | The University of Tokyo | Purdue University

The underlying mechanisms of code generation errors in large language models (LLMs) remain poorly understood. Method: Leveraging the HumanEval benchmark, this work systematically analyzes errors produced by six state-of-the-art LLMs and introduces, for the first time, a multidimensional, fine-grained error taxonomy integrating both semantic and syntactic dimensions. Using open coding and thematic analysis—augmented by statistical testing and qualitative root-cause attribution—the study identifies over ten recurrent error patterns, including logical flaws, boundary condition failures, and API misuse. Contribution/Results: The analysis reveals that LLM errors exhibit nontriviality, cross-line dependencies, and dispersed distribution—uncovering latent, deep-seated errors even in high-pass-rate tasks. It further demonstrates a nonlinear positive correlation between error frequency and task complexity. This taxonomy provides an interpretable, extensible theoretical foundation and empirical grounding for error localization, diagnosis, and repair in LLM-generated code.

Analyze code generation errors by LLMsClassify semantic and syntactic error characteristicsExplore error correlation with task complexity

Traditional Byte-Pair Encoding (BPE) tokenization introduces token redundancy in low-resource languages, degrading the performance of small-scale models. Method: This paper proposes a BPE configuration method integrating hyperparameter optimization and compressed sensing. It systematically searches key BPE hyperparameters—including vocabulary size and merge iterations—and jointly evaluates configurations using intrinsic metrics (e.g., token count) and extrinsic task performance (generation and classification). Contribution/Results: The study provides the first empirical evidence that BPE configuration significantly impacts multilingual modeling for low-resource languages. Experiments across diverse languages and model scales show that optimal configurations reduce token counts by 12.7% on average and improve downstream task accuracy by 1.8–3.4 percentage points for small models. These gains substantially enhance modeling efficiency and generalization capability in low-resource settings.

Compression-optimized tokenization benefits low-resource languagesImproved performance in multilingual NLP tasksOptimal BPE configuration reduces token count

Latest Papers

What's happening recently
View more

This work addresses the challenge that byte-level language models often generate invalid UTF-8 sequences when producing rare or previously unseen characters, thereby compromising the reliability of multilingual text generation. The authors train a 355-million-parameter, byte-level language model on 80 billion tokens of multilingual data and introduce a perplexity-independent evaluation protocol to assess structural validity of UTF-8 sequences. Their findings reveal that achieving convergence in UTF-8 validity requires approximately twice as much training data as needed for perplexity stabilization—42 billion versus 21 billion tokens. Notably, in context-free generation, rare characters exhibit higher structural validity than common ones, challenging conventional assumptions in representation learning. These results demonstrate that reliable UTF-8 sequence generation constitutes a distinct modeling capability beyond mere perplexity minimization.

byte-level tokenizationlanguage modelsrare characters

This work addresses the limitations of existing Unicode code point–based text evaluation methods, which often fail to accurately measure character-level errors in complex writing systems where a single grapheme frequently comprises multiple code points. To overcome this, the authors introduce grapheme-kit, an open-source Python library that, for the first time, extends widely used NLP evaluation metrics—such as edit distance and similarity—to the grapheme cluster level. By adhering to Unicode standards for grapheme cluster identification, composition, and decomposition, the proposed approach significantly improves evaluation accuracy for tasks like OCR on scripts with complex orthographies, including Tamil and Sinhala. This advancement provides a precise, grapheme-aware toolkit for text processing in low-resource languages.

complex scriptsgraphemelexical metrics

This study addresses the excessive token overhead and constrained context windows caused by UTF-8 encoding for non-English scripts in multilingual large language models. To this end, we propose a universal byte-level encoding scheme featuring a novel dual-alphabet dynamic routing mechanism that hybridizes UTF-8 and UTF-16 encodings to optimize multi-byte character processing. Crucially, this method modifies only the underlying byte representations without altering BPE merge rules, thereby achieving lossless and efficient cross-script encoding. Experimental results demonstrate that our approach substantially reduces token counts for high-premium scripts and enhances effective context utilization. Furthermore, it accelerates inference while preserving model quality, effectively mitigating cross-lingual performance disparities.

BBPEbyte-level encodingcross-script token disparity

This study addresses the limited support of current large code models and their evaluation methodologies for non-English natural language elements, such as code comments. The authors systematically evaluate five prominent models—including CodeGemma and CodeLlama—on comment generation across Dutch, English, Greek, Polish, and Chinese. They introduce the first multilingual code comment dataset comprising 12,500 human-annotated samples and propose a fine-grained error taxonomy encompassing 26 error categories. Their findings reveal a substantial degradation in comment quality for non-English languages, with linguistic errors increasing by up to 15.1×. Moreover, existing automatic evaluation methods, including neural metrics and LLM-as-a-judge approaches, prove unreliable in detecting linguistic and semantic inaccuracies, underscoring the irreplaceable role of human judgment in evaluating multilingual code generation.

code comment evaluationLLM evaluationmultilingual code generation

This study addresses the cross-platform display inconsistencies and retrieval failures in Yoruba caused by the absence of precomposed characters in Unicode, highlighting the structural limitations of existing NFC normalization strategies. By combining empirical analysis with multi-platform text processing techniques, this work systematically documents rendering failures in digital environments. Moving beyond conventional approaches that rely solely on combining character sequences, it proposes submitting a formal request to the Unicode Consortium for the direct encoding of essential characters. The primary contribution lies in identifying the structural root cause of this technical barrier and providing an actionable, fundamental encoding remediation to advance the digital representation of African languages.

diacriticsnormalizationprecomposed characters

Hot Scholars

XN

Xuefei Ning

Tsinghua University
Other Interesting Toys for MeReasoningLearningEfficient Deep Learning
NA

Nouar AlDahoul

PHD, AI Research Scientist in New York University, Abu Dhabi, UAE
Social Science-Large Language Models-Machine learning-Computer Vision-Internet of Things
KW

Kasun Wickramasinghe

University of Moratuwa
Natural Language ProcessingArtificial IntelligenceMachine LearningLow Resource Languages
SR

Surangika Ranathunga

Senior Lecturer, School of Mathematical and Computational Sciences, Massey University, New Zealand
Natural Language ProcessingMachine LearningLarge Language Models