Score
Designs and implements methods and tools to extract, score, and analyze contiguous token sequences (n-grams) and broader lexical or structural patterns from text or token streams, and to mine frequent, discriminative, or sequential patterns and association rules. Builds lexicon-based feature sets, pattern dictionaries, and rule-based detectors or feature engineering pipelines that integrate mined n-grams and lexicons for downstream analysis or modeling.
The lack of standardized annotation schemas and general-purpose extraction methods for unstructured medical text hinders structured data mining. Method: This paper proposes an open-domain, end-to-end framework for numerical triplet (value–measure–unit) extraction. It introduces the first template-free, general numerical association annotation scheme and jointly models medical text preprocessing, sequence labeling, and relation extraction to simultaneously identify and align numerical values, clinical measures, and associated units. Results: Evaluated on a thrombectomy literature dataset, the framework achieves a Dice coefficient of 0.82; random sampling validation confirms high accuracy in value–entity relation matching. This work establishes a transferable methodology and practical paradigm for structuring template-free medical texts.
The absence of automated, systematic approaches for modeling pattern recognition in conceptual modeling hinders improvements in knowledge representation and modeling quality. Method: This paper introduces frequent subgraph mining—systematically applied to conceptual modeling for the first time—and proposes a cross-language (OntoUML/ArchiMate), multi-criteria structural pattern discovery framework. It integrates a gSpan variant with graph editing, graph isomorphism testing, and pattern abstraction techniques to build an extensible, exploratory analysis tool. Contribution/Results: Evaluated on two authoritative datasets, the method successfully identifies highly reusable structural patterns. It demonstrates effectiveness in assessing modeling practices, supporting language evolution, and optimizing model quality—thereby filling a critical research gap in automated pattern mining for conceptual modeling.
In grammar reverse-engineering of legacy parsers, insufficient input samples often lead to incomplete grammar coverage—particularly missing edge cases or deprecated features. Method: This paper proposes an automated input generation approach based on dynamic symbolic execution (DSE), the first to apply DSE to grammar mining. We design a three-stage decoupled input generation framework and an iterative expansion strategy to effectively mitigate DSE’s inherent limitations in handling structured inputs. Crucially, our method requires no prior input samples and systematically triggers deep parser behaviors. Results: Evaluated on 11 real-world benchmarks, our generated grammars achieve precision and recall comparable to state-of-the-art methods, while significantly improving detection of subtle semantic features and historical edge-case usages.
This paper addresses three key challenges in entity relation mining from unstructured text: (1) difficulty in extracting meaningful entity associations, (2) weak support for dynamic knowledge graph construction, and (3) poor cross-domain adaptability. To this end, we propose the first domain-agnostic, modular, and plug-and-play end-to-end framework for entity relation mining. Methodologically, it integrates multi-source, configurable entity extraction—including DBpedia Spotlight, spaCy NER, and dictionary- or rule-based phrase extraction—with co-occurrence graph modeling, frequency-based statistics, and a novel quantitative relation scoring mechanism for ranked association inference. The framework further supports document filtering, dynamic knowledge graph augmentation, and rapid adaptation to diverse business scenarios. Empirical evaluation on two financial tasks—brand-product discovery and supplier risk monitoring—demonstrates its effectiveness, significantly reducing development redundancy while improving prototyping efficiency and system reusability.
Existing graph mining methods primarily focus on topological subgraph discovery and lack a unified mechanism for jointly modeling syntactic and semantic aspects of association rules over attributed graphs. This paper proposes the MINE GRAPH RULE operator, the first to enable integrated syntactic–semantic expression of graph association rules in an attributed graph database—implemented as an extension to Neo4j. Syntactically, conditions are specified via Cypher-like queries; semantically, rule quality is evaluated using support and confidence metrics. The operator tightly couples graph structure with attribute semantics, leverages Neo4j’s native query optimization, and incorporates relational association rule pruning strategies to ensure efficiency and portability. Experiments demonstrate strong scalability across multidimensional parameters. An open-source plugin implementing the operator significantly enhances both the expressiveness and practical utility of graph association rule mining.
This study addresses the high technical barrier in graph rule mining, which conventionally demands specialized expertise in graph theory and query languages. To overcome this limitation, we propose a zero-code interactive framework leveraging large language models (LLMs). By employing prompt engineering to bridge user intent with property graph mining workflows, this work introduces a novel mechanism enabling LLMs to directly generate MINE GRAPH RULE queries, thereby facilitating the automated generation, optimization, and interpretation of complex relational rules. This research effectively translates expert-level graph mining capabilities into natural language interactions, substantially lowering the technical threshold for practitioners. Consequently, the proposed framework significantly enhances both the efficiency and accuracy of complex graph association rule mining for non-expert users, democratizing access to advanced graph analytics.
Extracting high-frequency n-grams—especially for large n—from massive datasets poses significant challenges in terms of accuracy, efficiency, and determinism. Method: This paper proposes Intergrams, a hardware-aware multi-pass algorithm that exploits the power-law distribution of n-gram frequencies. It generates candidate n-grams from frequent (n−1)-grams, applies frequency-based pruning, and incorporates low-level optimizations to progressively shrink the search space across iterative passes. Theoretical analysis guides algorithm design to ensure exactness and strong scalability. Results: On real-world large-scale datasets, Intergrams achieves 10.3×–33× speedup over the state-of-the-art method. It is the first deterministic approach to break the performance bottleneck for extracting high-frequency n-grams with large n, while guaranteeing correctness and scalability.
研究通过J-Miner从微调后的语言模型分类器中提取决策知识,并将其编码为可执行形式,以实现高精度的任务决策再现和知识转移。
This study addresses the widespread assumption in large language models that tokens serve as stable units of measurement, despite significant variation in token length across different tokenizers and text domains—a discrepancy that introduces bias in model evaluation and billing. For the first time, this work systematically quantifies the sequence compression behavior of mainstream tokenizers under diverse text distributions through large-scale empirical analysis, revealing the high variability of token lengths. By challenging the common simplifying assumption that token length is approximately constant, the research elucidates the limitations of treating tokens as a universal metric. These findings provide a more accurate theoretical foundation for assessing model performance, estimating computational resources, and designing fairer usage-based billing mechanisms.
This study addresses the fundamental question of whether linguistic resources—such as dictionaries and grammars—should be constructed through meticulous manual curation or via scalable, automated methods. By systematically comparing representative resources like WordNet, FrameNet, and TAG, and integrating insights from both theoretical linguistics and natural language processing practice, the paper evaluates the trade-offs between these approaches in terms of semantic richness, construction efficiency, and downstream application performance. The findings indicate that manually crafted resources offer fine-grained semantic detail but incur high development costs, whereas automatically generated ones provide strong scalability at the expense of informational depth. A hybrid strategy combining both paradigms emerges as the most viable path forward. This work thus proposes a new paradigm for linguistic resource development that balances quality and efficiency, fostering synergistic advancement between linguistic theory and language technology.