normalize data representations

Designs, builds, or evaluates preprocessing and normalization components that transform numeric and textual representations into canonical, constrained, or comparable forms. This includes converting model scores to probability distributions (softmax), scaling numerical features and weights, and canonicalizing strings (Unicode, orthography, casing, punctuation and other textual variants) to enforce sum/partition constraints and reduce lexical mismatch.

normalizedatarepresentations

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.39
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Language Models over Canonical Byte-Pair Encodings

Jun 09, 2025
TV
Tim Vieira
🏛️ ETH Zürich | Mila | McGill University

Existing BPE-based language models assign non-zero probabilities to numerous “non-canonical” tokens—i.e., token sequences that are decodable into valid strings but cannot be generated by a deterministic tokenizer—causing probability leakage and modeling inefficiency. This work is the first to systematically formulate and implement **token-level canonicity constraints** for language models, distinguishing between *conditional canonicity* (inference-time reweighting) and *constructive canonicity* (parameter-level enforcement), thereby eliminating exponential-scale invalid probability mass at its source. Leveraging canonicity-aware model design, test-time reweighting, and rigorous likelihood evaluation, we demonstrate consistent and significant improvements in held-out log-likelihood across multiple architectures and corpora. Our results validate that canonicity constraints are both theoretically well-founded and practically effective, yielding trainable performance gains without architectural overhaul.

Language models assign probability to invalid token encodingsNoncanonical token strings waste probability massProposing methods to enforce canonical token outputs

This work addresses the vulnerability of pretrained vision models to affine transformations—such as rotation and scaling—that preserve semantic class but often lead to misclassification. The authors propose a test-time normalization method that maps inputs into a canonical form aligned with the training distribution, without modifying or retraining the classifier. Their key insight is framing this robustness challenge within an out-of-distribution (OOD) detection framework: OOD scores guide the search for beneficial affine transformations, and a gating mechanism applies these transformations only when necessary. Through systematic evaluation of over 20 OOD scoring functions and nine search strategies, they find that distance-based scores combined with random search followed by local optimization yield the best performance. The approach consistently enhances robustness across diverse benchmarks—including handwritten characters, sketches, natural images, and 3D point clouds—while preserving accuracy on in-distribution data.

affine transformationsout-of-distribution detectionrobustness

ReForm: Reflective Autoformalization with Prospective Bounded Sequence Optimization

Oct 28, 2025
GC
Guoxin Chen
🏛️ Renmin University of China | Tongyi Lab | Alibaba Group

Automated formalization suffers from semantic distortion: large language models often generate syntactically correct but semantically inaccurate formal statements, lacking human experts’ reflective reasoning and iterative refinement capabilities. To address this, we propose ReForm, a reflective automated formalization framework featuring a novel generation–evaluation–self-correction loop integrated with prospective bounded sequence optimization (PBSO), sequence-level reinforcement learning reward modeling, and a semantic consistency assessment mechanism. To support training and evaluation, we introduce ConsistencyCheck—the first human-annotated benchmark explicitly designed to measure semantic fidelity in formalization. Experiments show that ReForm achieves an average 17.2-percentage-point improvement over the strongest baselines across four mainstream benchmarks. Moreover, ConsistencyCheck reveals that 38.5% of expert-provided formalizations contain semantic errors, underscoring the task’s inherent difficulty and validating ReForm’s effectiveness and conceptual novelty.

Addressing LLMs' lack of self-reflection during autoformalizationDeveloping iterative refinement for machine-verifiable formal statementsImproving semantic fidelity in natural-to-formal mathematics translation

Broken Tokens? Your Language Model can Secretly Handle Non-Canonical Tokenizations

Jun 23, 2025
BS
Brian Siyuan Zheng
🏛️ University of Washington | Stanford University

This work investigates the robustness of language models to out-of-distribution, nonstandard tokenization schemes—such as character-level splitting, random segmentation, and right-aligned numeric grouping—not encountered during training. Leveraging 20 diverse benchmark tasks, we systematically evaluate both base and instruction-finetuned (IF) models on semantic understanding and generation fluency. Results demonstrate that IF models exhibit strong robustness: retaining 93.4% of original performance under random tokenization and 90.8% under character-level tokenization. Crucially, carefully designed nonstandard tokenization yields measurable task gains—up to +14% in string manipulation and code comprehension, and +33% in large-number arithmetic accuracy. We further establish, for the first time, that instruction finetuning is the primary source of this robustness. Moreover, we show that inference-time tokenization interventions serve as a lightweight, training-free mechanism for performance enhancement.

Assess LM robustness to unseen non-canonical tokenizationsExplore performance gains from alternative tokenization schemesInvestigate source of robustness in instruction-tuned models

This work addresses the inefficiency of calibration data selection in post-training compression of large language models by proposing ZipCal, a model-agnostic data filtering method that leverages the Zipfian power-law distribution without relying on model-specific signals. By analyzing word frequency distributions, ZipCal constructs calibration sets with high lexical diversity at linear computational complexity. Experimental results demonstrate that ZipCal significantly outperforms random sampling across multiple pruning and quantization benchmarks, achieving performance comparable to state-of-the-art perplexity-based methods while reducing computational overhead by an average of approximately 240×.

calibration datadata curationmodel compression

Latest Papers

What's happening recently
View more

This study investigates whether latent representations from heterogeneous text embedding models can be transferred via simple transformations to enable direct AI-to-AI communication without decoding into human-readable text. For the first time, we systematically evaluate the effectiveness and limitations of linear mappings as lightweight translators across nine diverse models varying in architecture, pooling strategy, and training objective, using real-world textual data. Through comprehensive metrics—including Centered Kernel Alignment (CKA) similarity, downstream task transfer performance, fidelity, and retrieval accuracy—we find that simple transformations succeed only between partially compatible model pairs and largely fail otherwise. These results indicate that semantic transfer across heterogeneous embedding spaces cannot be universally achieved through alignment alone, as compatibility is jointly constrained by architectural design, training objectives, pooling mechanisms, and data distribution.

heterogeneous embeddingslatent universalitymodel compatibility

This study addresses the lack of multidimensional, systematic tools for evaluating text simplification by large language models that simultaneously meet research and educational requirements. To bridge this gap, we propose an interactive human-in-the-loop web application enabling parallel simplification and real-time comparative analysis across diverse prompt–model (P×M) configurations, tailored to any target CEFR proficiency level. The core innovation lies in a visualization mechanism that integrates a hierarchical semantic alignment engine with a linear bias heuristic (λ), substantially reducing cognitive load during manual evaluation and facilitating reproducible, structured annotations. The system combines LLM APIs, semantic alignment algorithms, and a responsive front-end framework. Both source code and a live demonstration platform are publicly released, and the tool is readily applicable to downstream NLP dataset construction.

human-in-the-loopLarge Language Modelsprompt evaluation

Existing approaches to automatic formalization often overlook the hierarchical logical structure inherent in mathematical statements. This work proposes the DSR framework, which achieves modular formalization by decomposing statements, constructing operator trees, and iteratively refining and repairing subtrees. It introduces, for the first time, the topological structure of operator trees to guide error localization and correction, and presents PRIME, a high-quality benchmark of formalized theorems. By integrating neural-symbolic systems, large language models, and formal verification, DSR significantly outperforms current methods under the same computational budget, establishing a new state-of-the-art in automatic formalization.

autoformalizationformal languagehierarchical logic

This work addresses the representational bias and geometric distortion commonly induced by heterogeneous tasks during pre-finetuning of domain-adaptive text embeddings, which often degrade downstream performance. To mitigate this issue, the authors propose REZE, a novel framework that formalizes representation shift control as a core principle of pre-finetuning. REZE decomposes feature-space relationships between anchor-positive pairs to identify task-variant directions and applies adaptive soft shrinkage to constrain undesirable shifts. Notably, this approach incurs no inference overhead while effectively preserving task-invariant semantic structures. Extensive experiments across multiple embedding backbones and domain benchmarks demonstrate that REZE consistently outperforms standard pre-finetuning and post-hoc regularization methods. Embedding space analysis further reveals that the induced representation shifts align closely with—and remain stable relative to—the original data manifold.

domain adaptationpre-finetuningrepresentation shift

Hot Scholars

AH

Ali Hamdi

Computer Science, MSA University
Computer VisionDeep LearningText Mining
MS

Maosong Sun

Professor of Computer Science and Technology, Tsinghua University
Natural Language ProcessingArtificial IntelligenceSocial Computing
GS

Grigori Sidorov

Professor of Computational Linguistics, Instituto Politécnico Nacional (IPN), Mexico
Computational LinguisticsNatural Language ProcessingArtificial IntelligenceMachine Learning
SR

Surangika Ranathunga

Senior Lecturer, School of Mathematical and Computational Sciences, Massey University, New Zealand
Natural Language ProcessingMachine LearningLarge Language Models
TC

Ting Cai

University of Wisconsin-Madison
Machine LearningData Science