Score
Designs and implements analyses, pipelines, and visualizations that compute and compare frequency counts and distributions of contiguous token sequences (n-grams) across corpora or system outputs; builds statistical tests, metrics, and detectors to find rare or anomalous n-gram patterns, measure repetition and diversity, and diagnose n-gram–based failure modes.
In large language model (LLM) pretraining, the relationship between training data and model behavior remains difficult to analyze efficiently; existing data debugging workflows are fragmented, highly coupled, and lack interactive support. Method: We propose DataLens—a lightweight, plugin-based framework for structured, interactive data understanding and editing—enabling search, sampling, editing, and import/export of pretraining data without modifying training code. Its modular backend supports mainstream Megatron-style frameworks (e.g., GPT-NeoX, Megatron-LM, NeMo) and unifies heterogeneous data processing pipelines. Contribution/Results: Open-sourced with comprehensive documentation, tutorials, and demonstration videos, DataLens significantly improves data accessibility, interpretability, and pretraining development efficiency. It lowers the barrier to adopting high-quality data tooling in LLM research and engineering, enabling rapid, iterative data-centric experimentation.
Existing text similarity tools—including large language models—struggle to distinguish superficial lexical overlap from genuine semantic similarity among underlying entities. To address this, we propose a non-parametric similarity analysis framework based on weighted n-grams. Our method incorporates a language-frequency penalty to correct statistical biases in English corpora, ensuring similarity scores reflect semantically related entities rather than surface-level word repetition. All computational steps are fully traceable and interpretable, and results support visualization (e.g., word clouds) for empirical validation. Extensive experiments across diverse domains—including biographies, scientific literature, and historical texts—demonstrate that the framework consistently identifies deep, cross-document entity-level semantic similarity. Results are deterministic and fully reproducible. An open-source implementation is publicly available.
This work investigates the *n*-gram novelty of language model (LM) generations—i.e., the proportion of *n*-grams in generated text absent from the training corpus. To quantify this, we introduce *n*-novelty, a formal metric, and develop Rusty-DAWG: the first Rust implementation of a Directed Acyclic Word Graph (DAWG) enabling O(1) average-case *n*-gram lookup for arbitrary *n*. Using Pythia models and systematic *n*-gram frequency analysis, we empirically demonstrate two key findings: (1) for *n* > 4, LM outputs exhibit *lower* *n*-gram novelty than human-written text, contradicting intuition; and (2) training *n*-gram frequency correlates strongly and negatively with model completion loss. These results challenge assumptions about LM creativity and memorization. We publicly release Rusty-DAWG to support reproducible, scalable analysis of pretraining data provenance and LM memory behavior, providing both a novel methodological tool and empirical grounding for future research on LM generalization and memorization.
A lack of standardized, reproducible methods for quantifying textual diversity in large language models (LLMs) hinders rigorous evaluation of generation quality and cross-model or cross-corpus comparisons. Method: We propose the first systematic framework for text diversity evaluation, empirically validating convergent validity of diversity metrics and identifying a minimal, complete metric set—comprising compression ratio (zlib/lz4), long n-gram self-repetition rate, Self-BLEU, and BERTScore—that exhibits low inter-metric correlation and complementary multidimensional coverage. Contribution/Results: We release *diversity*, an open-source Python library enabling efficient computation and interactive visualization. Empirical analysis demonstrates that lightweight compression-based metrics robustly substitute for computationally expensive n-gram homogeneity scores. The framework substantially enhances interpretability, comparability, and practical utility of diversity assessment in LLM research.
Traditional Byte-Pair Encoding (BPE) tokenization introduces token redundancy in low-resource languages, degrading the performance of small-scale models. Method: This paper proposes a BPE configuration method integrating hyperparameter optimization and compressed sensing. It systematically searches key BPE hyperparameters—including vocabulary size and merge iterations—and jointly evaluates configurations using intrinsic metrics (e.g., token count) and extrinsic task performance (generation and classification). Contribution/Results: The study provides the first empirical evidence that BPE configuration significantly impacts multilingual modeling for low-resource languages. Experiments across diverse languages and model scales show that optimal configurations reduce token counts by 12.7% on average and improve downstream task accuracy by 1.8–3.4 percentage points for small models. These gains substantially enhance modeling efficiency and generalization capability in low-resource settings.
Intermittent failures in continuous integration (CI) pipelines are notoriously difficult to diagnose, leading to wasted resources and reduced development efficiency. This work proposes FlaXifyer, a few-shot learning approach that integrates the interpretable AI technique LogSift to fine-tune pretrained language models on pipeline logs using only 12 labeled examples per failure class. The method simultaneously predicts failure categories and pinpoints critical log entries indicative of root causes. Evaluated on 2,458 real-world CI failures, FlaXifyer achieves a Macro F1 score of 84.3% and a Top-2 accuracy of 92.0%, reducing the required log inspection effort by 74.4%. Furthermore, it successfully identifies the underlying fault in 87% of cases, demonstrating its effectiveness in accelerating failure diagnosis with minimal labeled data.
Static analysis of C/C++ programs faces significant challenges due to pointer aliasing, multi-level indirection, function pointers, and type ambiguity induced by `typedef`. To address these, this paper proposes an end-to-end, compiler-agnostic, interprocedural, type-aware static analysis system. Methodologically, it leverages Clang LibTooling to construct a unified multi-view intermediate representation—integrating AST, CFG, and DFG—and introduces custom CFG/DFG construction algorithms alongside an alias-aware type inference module, enabling the first interprocedural, type-sensitive data-flow modeling. Contributions include: (1) statement-level control-flow and type-aware data-flow graph generation for uncompiled code; (2) empirical validation on real-world open-source projects demonstrating high-fidelity modeling of complex semantic dependencies; and (3) provision of interpretable, structurally enriched input representations for downstream software engineering tasks such as vulnerability detection and code completion.
This study addresses the challenges posed by AI-generated code in academic integrity, hiring assessments, and software security by proposing a multi-perspective invariant representation learning framework for robust code provenance detection across multiple programming languages and generative models. The approach integrates structural prefixes, lexical normalization, symmetric KL divergence consistency loss, token dropout, and hybrid content augmentation, along with a class-weighting strategy to mitigate performance degradation caused by extreme class imbalance in multi-class settings. Fine-tuned from UniXcoder-base, the model achieves a macro F1 score of 0.845 on binary classification tasks and significantly improves the multi-class macro F1 from 0.086 to 0.345—a relative gain of 301%—demonstrating its effectiveness and strong generalization capability.
This study addresses the prevalent yet underexplored issue of refactorable repetitive step subsequences in Behavior-Driven Development (BDD) tests, for which no automated identification and classification methods previously existed. We propose the first end-to-end framework that leverages Sentence-BERT, UMAP, and HDBSCAN to perform semantic clustering on Gherkin corpora, thereby uncovering recurring fragments. These fragments are then annotated manually to train an XGBoost classifier that ranks their refactoring potential and assigns them to one of three established refactoring patterns. Applying our approach across 339 repositories, we identify 692,020 repetitive patterns and release the first large-scale annotated dataset, along with a complete toolchain and evaluation benchmark. Experimental results demonstrate that our XGBoost classifier achieves an F1 score of 0.891, significantly outperforming rule-based baselines and LLM-based judges, with 75% of test scenarios containing high-potential refactoring candidates.