n-gram and pattern mining

Designs and implements methods and tools to extract, score, and analyze contiguous token sequences (n-grams) and broader lexical or structural patterns from text or token streams, and to mine frequent, discriminative, or sequential patterns and association rules. Builds lexicon-based feature sets, pattern dictionaries, and rule-based detectors or feature engineering pipelines that integrate mined n-grams and lexicons for downstream analysis or modeling.

n-gramandpatternmining

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.14
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Text2Struct: A Machine Learning Pipeline for Mining Structured Data from Text

Dec 18, 2022
CZ
Chaochao Zhou
🏛️ Northwestern University Feinberg School of Medicine | University of Illinois Urbana-Champaign

The lack of standardized annotation schemas and general-purpose extraction methods for unstructured medical text hinders structured data mining. Method: This paper proposes an open-domain, end-to-end framework for numerical triplet (value–measure–unit) extraction. It introduces the first template-free, general numerical association annotation scheme and jointly models medical text preprocessing, sequence labeling, and relation extraction to simultaneously identify and align numerical values, clinical measures, and associated units. Results: Evaluated on a thrombectomy literature dataset, the framework achieves a Dice coefficient of 0.82; random sampling validation confirms high accuracy in value–entity relation matching. This work establishes a transferable methodology and practical paradigm for structuring template-free medical texts.

Extracting structured data from unstructured textsLack of annotation scheme and training datasetMining metrics and units associated with numerals

Mining Frequent Structures in Conceptual Models

Jun 11, 2024
MF
Mattia Fumagalli
🏛️ Free University of Bozen-Bolzano | University of Twente | University of Catania

The absence of automated, systematic approaches for modeling pattern recognition in conceptual modeling hinders improvements in knowledge representation and modeling quality. Method: This paper introduces frequent subgraph mining—systematically applied to conceptual modeling for the first time—and proposes a cross-language (OntoUML/ArchiMate), multi-criteria structural pattern discovery framework. It integrates a gSpan variant with graph editing, graph isomorphism testing, and pattern abstraction techniques to build an extensible, exploratory analysis tool. Contribution/Results: Evaluated on two authoritative datasets, the method successfully identifies highly reusable structural patterns. It demonstrates effectiveness in assessing modeling practices, supporting language evolution, and optimizing model quality—thereby filling a critical research gap in automated pattern mining for conceptual modeling.

Knowledge RepresentationMature MethodStructural Patterns

Generating Inputs for Grammar Mining using Dynamic Symbolic Execution

Aug 05, 2025
AP
Andreas Pointner
🏛️ University of Applied Sciences Upper Austria | Johannes Kepler University

In grammar reverse-engineering of legacy parsers, insufficient input samples often lead to incomplete grammar coverage—particularly missing edge cases or deprecated features. Method: This paper proposes an automated input generation approach based on dynamic symbolic execution (DSE), the first to apply DSE to grammar mining. We design a three-stage decoupled input generation framework and an iterative expansion strategy to effectively mitigate DSE’s inherent limitations in handling structured inputs. Crucially, our method requires no prior input samples and systematically triggers deep parser behaviors. Results: Evaluated on 11 real-world benchmarks, our generated grammars achieve precision and recall comparable to state-of-the-art methods, while significantly improving detection of subtle semantic features and historical edge-case usages.

Generating diverse inputs for grammar mining automaticallyImproving grammar coverage by capturing edge casesOvercoming limitations of Dynamic Symbolic Execution for parsers

Building Entity Association Mining Framework for Knowledge Discovery

Jun 02, 2025
AR
Anshika Rawal
🏛️ Fidelity Investments

This paper addresses three key challenges in entity relation mining from unstructured text: (1) difficulty in extracting meaningful entity associations, (2) weak support for dynamic knowledge graph construction, and (3) poor cross-domain adaptability. To this end, we propose the first domain-agnostic, modular, and plug-and-play end-to-end framework for entity relation mining. Methodologically, it integrates multi-source, configurable entity extraction—including DBpedia Spotlight, spaCy NER, and dictionary- or rule-based phrase extraction—with co-occurrence graph modeling, frequency-based statistics, and a novel quantitative relation scoring mechanism for ranked association inference. The framework further supports document filtering, dynamic knowledge graph augmentation, and rapid adaptation to diverse business scenarios. Empirical evaluation on two financial tasks—brand-product discovery and supplier risk monitoring—demonstrates its effectiveness, significantly reducing development redundancy while improving prototyping efficiency and system reusability.

Developing a reusable framework for text mining applicationsExtracting patterns from unstructured text for business decisionsMining entity associations to enrich knowledge graphs

Existing graph mining methods primarily focus on topological subgraph discovery and lack a unified mechanism for jointly modeling syntactic and semantic aspects of association rules over attributed graphs. This paper proposes the MINE GRAPH RULE operator, the first to enable integrated syntactic–semantic expression of graph association rules in an attributed graph database—implemented as an extension to Neo4j. Syntactically, conditions are specified via Cypher-like queries; semantically, rule quality is evaluated using support and confidence metrics. The operator tightly couples graph structure with attribute semantics, leverages Neo4j’s native query optimization, and incorporates relational association rule pruning strategies to ensure efficiency and portability. Experiments demonstrate strong scalability across multidimensional parameters. An open-source plugin implementing the operator significantly enhances both the expressiveness and practical utility of graph association rule mining.

Defining MINE GRAPH RULE operator for property graph miningExpressing syntax and semantics for graph association rulesImplementing and evaluating operator on Neo4j with real-world data

Latest Papers

What's happening recently
View more

This study addresses the high technical barrier in graph rule mining, which conventionally demands specialized expertise in graph theory and query languages. To overcome this limitation, we propose a zero-code interactive framework leveraging large language models (LLMs). By employing prompt engineering to bridge user intent with property graph mining workflows, this work introduces a novel mechanism enabling LLMs to directly generate MINE GRAPH RULE queries, thereby facilitating the automated generation, optimization, and interpretation of complex relational rules. This research effectively translates expert-level graph mining capabilities into natural language interactions, substantially lowering the technical threshold for practitioners. Consequently, the proposed framework significantly enhances both the efficiency and accuracy of complex graph association rule mining for non-expert users, democratizing access to advanced graph analytics.

Graph Rule MiningLarge Language ModelsProperty Graphs

Intermediate N-Gramming: Deterministic and Fast N-Grams For Large N and Large Datasets

Nov 18, 2025
RR
Ryan R. Curtin
🏛️ Booz Allen Hamilton | University of Maryland, Baltimore County | CrowdStrike | Laboratory for Physical Sciences

Extracting high-frequency n-grams—especially for large n—from massive datasets poses significant challenges in terms of accuracy, efficiency, and determinism. Method: This paper proposes Intergrams, a hardware-aware multi-pass algorithm that exploits the power-law distribution of n-gram frequencies. It generates candidate n-grams from frequent (n−1)-grams, applies frequency-based pruning, and incorporates low-level optimizations to progressively shrink the search space across iterative passes. Theoretical analysis guides algorithm design to ensure exactness and strong scalability. Results: On real-world large-scale datasets, Intergrams achieves 10.3×–33× speedup over the state-of-the-art method. It is the first deterministic approach to break the performance bottleneck for extracting high-frequency n-grams with large n, while guaranteeing correctness and scalability.

Deterministic fast n-gram extraction with hardware-optimized multi-pass algorithmEfficiently computing top-k frequent n-grams for large n and datasetsOvercoming exponential growth of n-gram features in computational processing

This study addresses the widespread assumption in large language models that tokens serve as stable units of measurement, despite significant variation in token length across different tokenizers and text domains—a discrepancy that introduces bias in model evaluation and billing. For the first time, this work systematically quantifies the sequence compression behavior of mainstream tokenizers under diverse text distributions through large-scale empirical analysis, revealing the high variability of token lengths. By challenging the common simplifying assumption that token length is approximately constant, the research elucidates the limitations of treating tokens as a universal metric. These findings provide a more accurate theoretical foundation for assessing model performance, estimating computational resources, and designing fairer usage-based billing mechanisms.

large language modelstext compressiontoken count

This study addresses the fundamental question of whether linguistic resources—such as dictionaries and grammars—should be constructed through meticulous manual curation or via scalable, automated methods. By systematically comparing representative resources like WordNet, FrameNet, and TAG, and integrating insights from both theoretical linguistics and natural language processing practice, the paper evaluates the trade-offs between these approaches in terms of semantic richness, construction efficiency, and downstream application performance. The findings indicate that manually crafted resources offer fine-grained semantic detail but incur high development costs, whereas automatically generated ones provide strong scalability at the expense of informational depth. A hybrid strategy combining both paradigms emerges as the most viable path forward. This work thus proposes a new paradigm for linguistic resource development that balances quality and efficiency, fostering synergistic advancement between linguistic theory and language technology.

grammarshandcraftingindustrialization

Hot Scholars

LU

Lyle Ungar

University of Pennsylvania
machine learningcomputational linguisticscomputational social science
SC

Sharath Chandra Guntuku

University of Pennsylvania
Digital HealthComputational PsychologySocial ListeningApplied Machine Learning
CD

Christine de Kock

NLP researcher, Melbourne University
natural language processingmachine learning
AS

Abeed Sarker

Emory University School of Medicine
Natural Language ProcessingBiomedical InformaticsHealth Data ScienceApplied Machine Learning