analyze n-gram frequency

Designs and implements analyses, pipelines, and visualizations that compute and compare frequency counts and distributions of contiguous token sequences (n-grams) across corpora or system outputs; builds statistical tests, metrics, and detectors to find rare or anomalous n-gram patterns, measure repetition and diversity, and diagnose n-gram–based failure modes.

analyzen-gramfrequency

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.11
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

TokenSmith: Streamlining Data Editing, Search, and Inspection for Large-Scale Language Model Training and Interpretability

Jul 25, 2025
MA
Mohammad Aflah Khan
🏛️ Max Planck Institute for Software Systems | University of Southern California

In large language model (LLM) pretraining, the relationship between training data and model behavior remains difficult to analyze efficiently; existing data debugging workflows are fragmented, highly coupled, and lack interactive support. Method: We propose DataLens—a lightweight, plugin-based framework for structured, interactive data understanding and editing—enabling search, sampling, editing, and import/export of pretraining data without modifying training code. Its modular backend supports mainstream Megatron-style frameworks (e.g., GPT-NeoX, Megatron-LM, NeMo) and unifies heterogeneous data processing pipelines. Contribution/Results: Open-sourced with comprehensive documentation, tutorials, and demonstration videos, DataLens significantly improves data accessibility, interpretability, and pretraining development efficiency. It lowers the barrier to adopting high-quality data tooling in LLM research and engineering, enabling rapid, iterative data-centric experimentation.

Enhancing dataset inspection and search in pretraining workflowsProviding accessible tools for model behavior and data relationship understandingSimplifying data editing and analysis for large-scale language model training

Existing text similarity tools—including large language models—struggle to distinguish superficial lexical overlap from genuine semantic similarity among underlying entities. To address this, we propose a non-parametric similarity analysis framework based on weighted n-grams. Our method incorporates a language-frequency penalty to correct statistical biases in English corpora, ensuring similarity scores reflect semantically related entities rather than surface-level word repetition. All computational steps are fully traceable and interpretable, and results support visualization (e.g., word clouds) for empirical validation. Extensive experiments across diverse domains—including biographies, scientific literature, and historical texts—demonstrate that the framework consistently identifies deep, cross-document entity-level semantic similarity. Results are deterministic and fully reproducible. An open-source implementation is publicly available.

Automatically compare text documents for meaningful insightsDevelop n-gram framework to uncover subject-level document similaritiesProvide explainable textual similarities beyond surface-level comparisons

Evaluating n-Gram Novelty of Language Models Using Rusty-DAWG

Jun 18, 2024
WM
William Merrill
🏛️ New York University | Allen Institute for AI | University of Washington

This work investigates the *n*-gram novelty of language model (LM) generations—i.e., the proportion of *n*-grams in generated text absent from the training corpus. To quantify this, we introduce *n*-novelty, a formal metric, and develop Rusty-DAWG: the first Rust implementation of a Directed Acyclic Word Graph (DAWG) enabling O(1) average-case *n*-gram lookup for arbitrary *n*. Using Pythia models and systematic *n*-gram frequency analysis, we empirically demonstrate two key findings: (1) for *n* > 4, LM outputs exhibit *lower* *n*-gram novelty than human-written text, contradicting intuition; and (2) training *n*-gram frequency correlates strongly and negatively with model completion loss. These results challenge assumptions about LM creativity and memorization. We publicly release Rusty-DAWG to support reproducible, scalable analysis of pretraining data provenance and LM memory behavior, providing both a novel methodological tool and empirical grounding for future research on LM generalization and memorization.

Analyzing factors affecting text novelty in language modelsDeveloping Rusty-DAWG tool for efficient n-gram searchEvaluating novelty of LM-generated texts versus training data

Standardizing the Measurement of Text Diversity: A Tool and a Comparative Analysis of Scores

Mar 01, 2024
CS
Chantal Shaib
🏛️ Northeastern University | Adobe

A lack of standardized, reproducible methods for quantifying textual diversity in large language models (LLMs) hinders rigorous evaluation of generation quality and cross-model or cross-corpus comparisons. Method: We propose the first systematic framework for text diversity evaluation, empirically validating convergent validity of diversity metrics and identifying a minimal, complete metric set—comprising compression ratio (zlib/lz4), long n-gram self-repetition rate, Self-BLEU, and BERTScore—that exhibits low inter-metric correlation and complementary multidimensional coverage. Contribution/Results: We release *diversity*, an open-source Python library enabling efficient computation and interactive visualization. Empirical analysis demonstrates that lightweight compression-based metrics robustly substitute for computationally expensive n-gram homogeneity scores. The framework substantially enhances interpretability, comparability, and practical utility of diversity assessment in LLM research.

Evaluating convergent validity of existing diversity scores is neededIdentifying repetitive structures in large text corpora is challengingStandardizing text diversity measurement lacks a universal method

Traditional Byte-Pair Encoding (BPE) tokenization introduces token redundancy in low-resource languages, degrading the performance of small-scale models. Method: This paper proposes a BPE configuration method integrating hyperparameter optimization and compressed sensing. It systematically searches key BPE hyperparameters—including vocabulary size and merge iterations—and jointly evaluates configurations using intrinsic metrics (e.g., token count) and extrinsic task performance (generation and classification). Contribution/Results: The study provides the first empirical evidence that BPE configuration significantly impacts multilingual modeling for low-resource languages. Experiments across diverse languages and model scales show that optimal configurations reduce token counts by 12.7% on average and improve downstream task accuracy by 1.8–3.4 percentage points for small models. These gains substantially enhance modeling efficiency and generalization capability in low-resource settings.

Compression-optimized tokenization benefits low-resource languagesImproved performance in multilingual NLP tasksOptimal BPE configuration reduces token count

Latest Papers

What's happening recently
View more

Intermittent failures in continuous integration (CI) pipelines are notoriously difficult to diagnose, leading to wasted resources and reduced development efficiency. This work proposes FlaXifyer, a few-shot learning approach that integrates the interpretable AI technique LogSift to fine-tune pretrained language models on pipeline logs using only 12 labeled examples per failure class. The method simultaneously predicts failure categories and pinpoints critical log entries indicative of root causes. Evaluated on 2,458 real-world CI failures, FlaXifyer achieves a Macro F1 score of 84.3% and a Top-2 accuracy of 92.0%, reducing the required log inspection effort by 74.4%. Furthermore, it successfully identifies the underlying fault in 87% of cases, demonstrating its effectiveness in accelerating failure diagnosis with minimal labeled data.

automated triagecontinuous integrationfailure diagnosis

ATLAS: Automated Tree-based Language Analysis System for C and C++ source programs

Dec 13, 2025
JM
Jaid Monwar Chowdhury
🏛️ Bangladesh University of Engineering and Technology

Static analysis of C/C++ programs faces significant challenges due to pointer aliasing, multi-level indirection, function pointers, and type ambiguity induced by `typedef`. To address these, this paper proposes an end-to-end, compiler-agnostic, interprocedural, type-aware static analysis system. Methodologically, it leverages Clang LibTooling to construct a unified multi-view intermediate representation—integrating AST, CFG, and DFG—and introduces custom CFG/DFG construction algorithms alongside an alias-aware type inference module, enabling the first interprocedural, type-sensitive data-flow modeling. Contributions include: (1) statement-level control-flow and type-aware data-flow graph generation for uncompiled code; (2) empirical validation on real-world open-source projects demonstrating high-fidelity modeling of complex semantic dependencies; and (3) provision of interpretable, structurally enriched input representations for downstream software engineering tasks such as vulnerability detection and code completion.

Generates control and data flow graphs for C/C++ programsHandles compilable and non-compilable multi-file C/C++ projectsProduces unified multi-view code representations for program analysis

This study addresses the challenges posed by AI-generated code in academic integrity, hiring assessments, and software security by proposing a multi-perspective invariant representation learning framework for robust code provenance detection across multiple programming languages and generative models. The approach integrates structural prefixes, lexical normalization, symmetric KL divergence consistency loss, token dropout, and hybrid content augmentation, along with a class-weighting strategy to mitigate performance degradation caused by extreme class imbalance in multi-class settings. Fine-tuned from UniXcoder-base, the model achieves a macro F1 score of 0.845 on binary classification tasks and significantly improves the multi-class macro F1 from 0.086 to 0.345—a relative gain of 301%—demonstrating its effectiveness and strong generalization capability.

AI-generated codeclass imbalancecode attribution

This study addresses the prevalent yet underexplored issue of refactorable repetitive step subsequences in Behavior-Driven Development (BDD) tests, for which no automated identification and classification methods previously existed. We propose the first end-to-end framework that leverages Sentence-BERT, UMAP, and HDBSCAN to perform semantic clustering on Gherkin corpora, thereby uncovering recurring fragments. These fragments are then annotated manually to train an XGBoost classifier that ranks their refactoring potential and assigns them to one of three established refactoring patterns. Applying our approach across 339 repositories, we identify 692,020 repetitive patterns and release the first large-scale annotated dataset, along with a complete toolchain and evaluation benchmark. Experimental results demonstrate that our XGBoost classifier achieves an F1 score of 0.891, significantly outperforming rule-based baselines and LLM-based judges, with 75% of test scenarios containing high-potential refactoring candidates.

Behaviour-Driven Developmentextraction-worthinessrefactoring pattern selection

Hot Scholars

AA

Antonis Antoniades

Ph.D. Student, University of California Santa Barbara
Machine LearningArtificial IntelligenceNeuroscience
XW

Xiting Wang

Associate Professor, Renmin University of China
Explainable AIAI AlignmentVisual AnalyticsTrustworthy AI
QY

Qingqing Ye

Assistant Professor, The Hong Kong Polytechnic University
data privacy and securityadversarial machine learning
SV

Svitlana Vakulenko

Vienna University of Economics and Business
RAGConversational SearchInformation Seeking
JY

Jinyoung Yeo

Yonsei University
Natural Language ProceesingLarge Language ModelsAI Agents