Score
Designs, builds, or evaluates preprocessing and normalization components that transform numeric and textual representations into canonical, constrained, or comparable forms. This includes converting model scores to probability distributions (softmax), scaling numerical features and weights, and canonicalizing strings (Unicode, orthography, casing, punctuation and other textual variants) to enforce sum/partition constraints and reduce lexical mismatch.
Existing BPE-based language models assign non-zero probabilities to numerous “non-canonical” tokens—i.e., token sequences that are decodable into valid strings but cannot be generated by a deterministic tokenizer—causing probability leakage and modeling inefficiency. This work is the first to systematically formulate and implement **token-level canonicity constraints** for language models, distinguishing between *conditional canonicity* (inference-time reweighting) and *constructive canonicity* (parameter-level enforcement), thereby eliminating exponential-scale invalid probability mass at its source. Leveraging canonicity-aware model design, test-time reweighting, and rigorous likelihood evaluation, we demonstrate consistent and significant improvements in held-out log-likelihood across multiple architectures and corpora. Our results validate that canonicity constraints are both theoretically well-founded and practically effective, yielding trainable performance gains without architectural overhaul.
This work addresses the vulnerability of pretrained vision models to affine transformations—such as rotation and scaling—that preserve semantic class but often lead to misclassification. The authors propose a test-time normalization method that maps inputs into a canonical form aligned with the training distribution, without modifying or retraining the classifier. Their key insight is framing this robustness challenge within an out-of-distribution (OOD) detection framework: OOD scores guide the search for beneficial affine transformations, and a gating mechanism applies these transformations only when necessary. Through systematic evaluation of over 20 OOD scoring functions and nine search strategies, they find that distance-based scores combined with random search followed by local optimization yield the best performance. The approach consistently enhances robustness across diverse benchmarks—including handwritten characters, sketches, natural images, and 3D point clouds—while preserving accuracy on in-distribution data.
Automated formalization suffers from semantic distortion: large language models often generate syntactically correct but semantically inaccurate formal statements, lacking human experts’ reflective reasoning and iterative refinement capabilities. To address this, we propose ReForm, a reflective automated formalization framework featuring a novel generation–evaluation–self-correction loop integrated with prospective bounded sequence optimization (PBSO), sequence-level reinforcement learning reward modeling, and a semantic consistency assessment mechanism. To support training and evaluation, we introduce ConsistencyCheck—the first human-annotated benchmark explicitly designed to measure semantic fidelity in formalization. Experiments show that ReForm achieves an average 17.2-percentage-point improvement over the strongest baselines across four mainstream benchmarks. Moreover, ConsistencyCheck reveals that 38.5% of expert-provided formalizations contain semantic errors, underscoring the task’s inherent difficulty and validating ReForm’s effectiveness and conceptual novelty.
This work investigates the robustness of language models to out-of-distribution, nonstandard tokenization schemes—such as character-level splitting, random segmentation, and right-aligned numeric grouping—not encountered during training. Leveraging 20 diverse benchmark tasks, we systematically evaluate both base and instruction-finetuned (IF) models on semantic understanding and generation fluency. Results demonstrate that IF models exhibit strong robustness: retaining 93.4% of original performance under random tokenization and 90.8% under character-level tokenization. Crucially, carefully designed nonstandard tokenization yields measurable task gains—up to +14% in string manipulation and code comprehension, and +33% in large-number arithmetic accuracy. We further establish, for the first time, that instruction finetuning is the primary source of this robustness. Moreover, we show that inference-time tokenization interventions serve as a lightweight, training-free mechanism for performance enhancement.
This work addresses the inefficiency of calibration data selection in post-training compression of large language models by proposing ZipCal, a model-agnostic data filtering method that leverages the Zipfian power-law distribution without relying on model-specific signals. By analyzing word frequency distributions, ZipCal constructs calibration sets with high lexical diversity at linear computational complexity. Experimental results demonstrate that ZipCal significantly outperforms random sampling across multiple pruning and quantization benchmarks, achieving performance comparable to state-of-the-art perplexity-based methods while reducing computational overhead by an average of approximately 240×.
This study investigates whether latent representations from heterogeneous text embedding models can be transferred via simple transformations to enable direct AI-to-AI communication without decoding into human-readable text. For the first time, we systematically evaluate the effectiveness and limitations of linear mappings as lightweight translators across nine diverse models varying in architecture, pooling strategy, and training objective, using real-world textual data. Through comprehensive metrics—including Centered Kernel Alignment (CKA) similarity, downstream task transfer performance, fidelity, and retrieval accuracy—we find that simple transformations succeed only between partially compatible model pairs and largely fail otherwise. These results indicate that semantic transfer across heterogeneous embedding spaces cannot be universally achieved through alignment alone, as compatibility is jointly constrained by architectural design, training objectives, pooling mechanisms, and data distribution.
This study addresses the lack of multidimensional, systematic tools for evaluating text simplification by large language models that simultaneously meet research and educational requirements. To bridge this gap, we propose an interactive human-in-the-loop web application enabling parallel simplification and real-time comparative analysis across diverse prompt–model (P×M) configurations, tailored to any target CEFR proficiency level. The core innovation lies in a visualization mechanism that integrates a hierarchical semantic alignment engine with a linear bias heuristic (λ), substantially reducing cognitive load during manual evaluation and facilitating reproducible, structured annotations. The system combines LLM APIs, semantic alignment algorithms, and a responsive front-end framework. Both source code and a live demonstration platform are publicly released, and the tool is readily applicable to downstream NLP dataset construction.
Existing approaches to automatic formalization often overlook the hierarchical logical structure inherent in mathematical statements. This work proposes the DSR framework, which achieves modular formalization by decomposing statements, constructing operator trees, and iteratively refining and repairing subtrees. It introduces, for the first time, the topological structure of operator trees to guide error localization and correction, and presents PRIME, a high-quality benchmark of formalized theorems. By integrating neural-symbolic systems, large language models, and formal verification, DSR significantly outperforms current methods under the same computational budget, establishing a new state-of-the-art in automatic formalization.
为了解决OCR文本处理中信息丢失和重复工作的问题,提出了一种名为Enriched Text的方法,通过保留元数据并进行多语言标注来优化大规模文本处理。
This work addresses the representational bias and geometric distortion commonly induced by heterogeneous tasks during pre-finetuning of domain-adaptive text embeddings, which often degrade downstream performance. To mitigate this issue, the authors propose REZE, a novel framework that formalizes representation shift control as a core principle of pre-finetuning. REZE decomposes feature-space relationships between anchor-positive pairs to identify task-variant directions and applies adaptive soft shrinkage to constrain undesirable shifts. Notably, this approach incurs no inference overhead while effectively preserving task-invariant semantic structures. Extensive experiments across multiple embedding backbones and domain benchmarks demonstrate that REZE consistently outperforms standard pre-finetuning and post-hoc regularization methods. Embedding space analysis further reveals that the induced representation shifts align closely with—and remain stable relative to—the original data manifold.