Score
Designs and implements methods to extract and analyze lexical and n-gram features from natural language text—e.g., token n-grams, word counts, frequency and co-occurrence statistics—and to produce representations that incorporate syntactic tags and semantic role annotations. These features and analyses are used to build or evaluate models for tasks such as text classification, information extraction, and other NLP applications.
Existing text similarity tools—including large language models—struggle to distinguish superficial lexical overlap from genuine semantic similarity among underlying entities. To address this, we propose a non-parametric similarity analysis framework based on weighted n-grams. Our method incorporates a language-frequency penalty to correct statistical biases in English corpora, ensuring similarity scores reflect semantically related entities rather than surface-level word repetition. All computational steps are fully traceable and interpretable, and results support visualization (e.g., word clouds) for empirical validation. Extensive experiments across diverse domains—including biographies, scientific literature, and historical texts—demonstrate that the framework consistently identifies deep, cross-document entity-level semantic similarity. Results are deterministic and fully reproducible. An open-source implementation is publicly available.
This work addresses few-shot named entity recognition (NER) by challenging the end-to-end joint modeling paradigm. Drawing on generative grammar theory, we propose a “extract-then-classify” decoupled framework: entity extraction is treated as a syntactic task—requiring no semantic information—while classification is delegated to pre-trained language models (PLMs) or large language models (LLMs) as a semantic task. Empirical analysis reveals that rare words—particularly proper nouns—serve as critical syntactic cues; high-precision extraction is achieved using only shallow syntactic features (e.g., POS tags, dependency relations, and n-grams), with word embeddings or contextualized semantic representations yielding no performance gain. On benchmarks including CoNLL-2003, our extraction module achieves state-of-the-art F1 scores; ablation studies confirm that incorporating semantic features does not improve extraction accuracy. To our knowledge, this is the first study grounding syntactic–semantic separation in formal linguistics, providing both theoretical justification and empirical validation for decoupled modeling, while elucidating the root cause of failures in multi-task joint parsing.
Biological sequence analysis across multi-omics—genomics, transcriptomics, and proteomics—faces fundamental challenges in modeling DNA, RNA, and protein sequences with appropriate granularity, evolutionary awareness, and functional interpretability. Method: This study systematically investigates the adaptation mechanisms and application boundaries of NLP techniques to biological sequences. We comprehensively map the architectural evolution from word2vec to Transformer and Hyena-based models, propose a cross-scale tokenization strategy, and introduce a task-driven evaluation framework. Contribution/Results: Through empirical validation on structural prediction, functional annotation, and gene expression modeling, we quantitatively benchmark model performance across sequence modeling fidelity, evolutionary signal capture, and functional generalization. Our analysis identifies precise performance ceilings and domain-specific applicability for each architecture. The work establishes a methodology for AI-native biological sequence modeling, enabling a paradigm shift toward precision biology grounded in foundational language modeling principles.
This work addresses the need for lightweight syntactic structure analysis by proposing a part-of-speech (POS)-free method to quantify sentence structural balance. Methodologically, it innovatively replaces conventional POS tags with character-level ASCII encodings to represent syntactic constituents, integrating lexical category alignment and PCA-based dimensionality reduction to construct a low-resource, interpretable structural balance metric. Experiments across 11 corpora demonstrate that Grok-generated text exhibits near-normal syntactic distribution (confirmed via Shapiro–Wilk and Anderson–Darling tests), indicating high structural balance. Among the remaining ten corpora, four also pass these normality tests. These results validate the method’s effectiveness and generalizability for preliminary text quality assessment and stylistic analysis, particularly in resource-constrained settings. The approach advances syntactic evaluation by eliminating reliance on external POS taggers while preserving interpretability and computational efficiency.
Constructing structured knowledge from multilingual (English/French/Spanish) social media texts in tourism faces challenges of linguistic diversity and prohibitively high annotation costs. Method: This paper proposes a multitask NLP framework tailored for few-shot learning, introducing the first fine-grained, multilingual, domain-specific dataset for tourism. It systematically evaluates and integrates few-shot learning, pattern-based prompting, and parameter-efficient fine-tuning, jointly leveraging sequence labeling and semantic resource alignment. Contribution/Results: The framework achieves performance on par with fully supervised baselines using only 15 samples for sentiment analysis, 160 for location recognition, and 200 for topic concept extraction (315 classes), effectively overcoming low-resource bottlenecks. It establishes a scalable, low-dependency automation paradigm for cross-lingual tourism sentiment analysis and knowledge graph construction.
This study investigates whether recent large language models (LLMs), while improving instruction alignment, concurrently sacrifice linguistic diversity. To this end, we introduce a novel evaluation paradigm that integrates ecological and information-theoretic diversity metrics within the formal framework of Head-Driven Phrase Structure Grammar (HPSG). We systematically compare syntactic structures and lexical type distributions between two generations of LLMs and human-authored English news texts. Our analysis reveals that newer, alignment-optimized models exhibit significantly reduced syntactic and lexical diversity compared to both their predecessors and contemporaneous human writing, which remains stable over time. These findings not only demonstrate a previously underappreciated side effect of alignment training—namely, linguistic simplification—but also establish a new methodological approach for assessing the expressive capacity of large language models.
This study addresses the bottleneck in acquiring lexical knowledge from Arabic–English dictionaries by proposing an integrated approach that combines n-gram modeling, keyword-in-context (KWIC) analysis, and rule-based information extraction. For the first time, this method systematically extracts morphological, syntactic, and semantic knowledge automatically from the Al-Mawrid bilingual dictionary. Leveraging punctuation patterns and heuristic strategies, the approach effectively identifies synonym sets, hyponymy-hypernymy relations, and domain-specific labels. Experimental results demonstrate high precision across all extraction tasks, with particularly strong recall for synonyms, thereby confirming that the Al-Mawrid dictionary encodes a rich repository of structured linguistic knowledge amenable to automated harvesting.
This study addresses the scarcity of multilingual corpora supporting emerging concepts—such as “non-technological innovation”—in the humanities and social sciences (HSS). To tackle this, we propose a hybrid multilingual corpus construction methodology that integrates corporate websites and annual reports, combining automatic language identification, domain-adapted content filtering, relevant paragraph extraction, expert-lexicon-driven contextual block identification, thematic annotation, and enriched structured metadata. Our key contribution is the first systematic construction of a high-quality, multilingual corpus specifically designed for HSS emerging concepts, accompanied by a parallel English supervised dataset with fine-grained thematic labels. The resulting resource enables cross-lingual lexical variation analysis, training of multilingual NLP models, and empirical social science research. It is both reusable and extensible, effectively bridging a critical gap between domain-specific knowledge modeling and computational linguistics applications.
This study addresses the systematic stylistic and semantic deviations of large language model (LLM)-generated text from human writing, particularly its lack of literary expressiveness. For the first time, it systematically identifies consistent n-gram distribution patterns in LLM outputs and demonstrates, through combined statistical analysis and qualitative comparison, that this stylistic impoverishment is closely linked to constrained semantic expression. Challenging the long-standing assumption that style and semantics are separable, the work elucidates their intrinsic coupling mechanism, offering a novel perspective for understanding the “non-stylistic” nature of LLM-generated text.
This study systematically investigates the intrinsic semantic and syntactic properties of mainstream word embedding methods—such as Word2Vec and GloVe—and their performance disparities across diverse natural language processing tasks. By establishing a unified evaluation framework that integrates publicly available pretrained embeddings with standard benchmark datasets, the work conducts empirical comparisons on canonical tasks including semantic similarity and analogical reasoning. The findings delineate the performance boundaries and optimal application scenarios for each embedding model, offering practitioners reliable guidance for model selection in real-world settings. Furthermore, the analysis deepens the understanding of the inherent limitations of static word representations, highlighting critical constraints in capturing contextual and compositional linguistic phenomena.