Score
Designs and constructs composite numeric indices that quantify stylistic properties of text by selecting, normalizing, weighting, and aggregating multiple stylistic features (lexical, syntactic, and higher-level signals) into an operational metric. Builds and validates variants tied to LLM-associated style signals, detects temporal shifts in the index, and analyzes relationships between the style index and related complexity or other quantitative metrics.
This study addresses key limitations in AI-generated text linguistics research—namely, insufficient cross-lingual and cross-model coverage, and a lack of systematic investigation into prompt sensitivity. We systematically review corpus- and computational linguistics literature (2020–2024) to develop a multidimensional classification framework encompassing lexical, syntactic, semantic, stylistic, and diversity dimensions. Through mixed-methods corpus analysis, we identify consistent linguistic patterns in AI outputs: heightened formality and impersonality, elevated noun and article usage, scarcity of adjectives and adverbs, low lexical diversity, and high repetition. Our contribution is threefold: (1) the first integrated empirical synthesis across languages (including Chinese, German, Spanish), models (beyond GPT-series), and genres; (2) explicit demonstration of how prompt engineering critically shapes output characteristics; and (3) provision of a theoretically grounded, methodologically robust foundation for multilingual AI-text detection, evaluation, and controllable generation.
This survey systematically reviews authorship analysis (AA) research from 2015 to 2024, focusing on the core tasks of author attribution and author verification. To address these, we synthesize methodological advances across feature engineering (e.g., n-grams, stylometric features), classical machine learning (e.g., SVM, Random Forest), deep learning (e.g., RNNs, Transformers), and large language models (via fine-tuning and prompt engineering), organizing insights into a four-dimensional framework: *method–feature–dataset–challenge*. Our contribution is threefold: first, we provide the first comprehensive taxonomy of multi-paradigm AA approaches, clarifying their evolutionary trajectories and applicability boundaries; second, we explicitly identify critical research gaps—including low-resource language processing, multilingual adaptability, cross-domain generalization, and detection of AI-generated text; third, we offer a principled theoretical foundation and actionable guidelines for developing robust, multilingual, and interpretable AA systems.
This work investigates how large language models (LLMs) distinguish author style from genre style, and the underlying representational mechanisms. We employ a multi-faceted analytical framework—including syntactic masking, attention head attribution, neuron activation tracking, controlled ablation, and style-classification fine-tuning—to systematically compare recognition pathways for these two stylistic dimensions. Results demonstrate that LLMs exhibit strong discriminative capability for both authorship and genre, yet rely on fundamentally distinct strategies: author style is primarily captured through local syntactic patterns and context-sensitive lexical choices, whereas genre style depends more heavily on global structural cues. Notably, pronoun usage and word order emerge as highly discriminative features across hierarchical model layers in both tasks. The study further reveals the critical role of fine-grained linguistic features—such as syntactic fine-tuning—in stylistic representation, providing interpretable evidence for how LLMs model stylistic variation.
This study systematically investigates stylistic differences between human- and large language model (LLM)-generated texts across genres, models, and decoding strategies to inform responsible LLM deployment. Leveraging Biber’s multidimensional framework of register variation, the authors conduct a large-scale comparative analysis of texts produced by 11 LLMs across eight genres and four decoding strategies. The findings reveal that model type and genre exert substantially stronger influences on textual style than prompting or decoding choices. Notably, chat-oriented models exhibit pronounced clustering in stylistic space, and key linguistic features of LLM-generated text demonstrate robustness across generation conditions, with genre effects consistently outweighing those of text origin.
This study addresses the systematic stylistic and semantic deviations of large language model (LLM)-generated text from human writing, particularly its lack of literary expressiveness. For the first time, it systematically identifies consistent n-gram distribution patterns in LLM outputs and demonstrates, through combined statistical analysis and qualitative comparison, that this stylistic impoverishment is closely linked to constrained semantic expression. Challenging the long-standing assumption that style and semantics are separable, the work elucidates their intrinsic coupling mechanism, offering a novel perspective for understanding the “non-stylistic” nature of LLM-generated text.
Existing content preservation evaluation methods for text style/attribute transfer—relying on lexical or semantic similarity metrics or current LLM-based evaluators—fail to model stylistic conditionality, resulting in low correlation with human judgments. Method: The authors propose, for the first time, that content preservation assessment must be *conditioned on style transfer*, and introduce a zero-shot evaluation method based on next-token conditional likelihood. They further construct a human-annotated benchmark specifically designed for meta-evaluation alignment to systematically validate the necessity of conditional modeling across multiple style transfer tasks. Contribution/Results: Experiments demonstrate that the proposed method significantly outperforms baseline approaches, achieving an average 23% improvement in correlation with human judgments. This work establishes conditional modeling as essential for accurate, human-aligned content preservation evaluation in style transfer.
A lack of standardized, reproducible methods for quantifying textual diversity in large language models (LLMs) hinders rigorous evaluation of generation quality and cross-model or cross-corpus comparisons. Method: We propose the first systematic framework for text diversity evaluation, empirically validating convergent validity of diversity metrics and identifying a minimal, complete metric set—comprising compression ratio (zlib/lz4), long n-gram self-repetition rate, Self-BLEU, and BERTScore—that exhibits low inter-metric correlation and complementary multidimensional coverage. Contribution/Results: We release *diversity*, an open-source Python library enabling efficient computation and interactive visualization. Empirical analysis demonstrates that lightweight compression-based metrics robustly substitute for computationally expensive n-gram homogeneity scores. The framework substantially enhances interpretability, comparability, and practical utility of diversity assessment in LLM research.
This study addresses the underexplored capacity of large language models to comprehend and generate aesthetic stylistic expressions in cross-cultural contexts, particularly their ability to employ culturally resonant linguistic strategies. Focusing on cross-cultural stylistic variations in film, television, and advertising texts from Hong Kong and Mainland China, this work introduces C4STYLI, a high-quality bilingual benchmark dataset, and proposes a dual evaluation framework that integrates both style recognition and generation. Through structural ablation studies and logistic regression probing analyses, the research reveals that models predominantly rely on surface-level linguistic features and lack deep structural understanding—especially in recognizing Hong Kong–specific stylistic conventions. A significant discrepancy is observed between model performance in recognition versus generation tasks, and both diverge markedly from human judgments, highlighting critical limitations in current models’ ability to capture cross-cultural stylistic nuances.
This study addresses the longstanding reliance on subjective judgment in narrative quality assessment by introducing a computational framework grounded in 33 quantifiable linguistic features spanning lexical, syntactic, and semantic dimensions. For the first time, this work systematically applies multidimensional quantitative stylometric indicators to the automatic evaluation of narrative quality. Leveraging natural language processing, clustering analysis, and similarity matrix construction, the proposed model achieves near-perfect discrimination between texts authored by professional editors and self-published writers. Furthermore, it significantly outperforms existing evaluation metrics on a manually annotated dataset, thereby overcoming the limitations inherent in traditional story-level assessment approaches.
This work addresses the limitations of current large language model–based style descriptions, which are often plagued by hallucination, bias, and a lack of interpretability and practical utility. To overcome these issues, the authors propose style-eliciting prompts as an interpretable interface for representing textual style. They construct a dataset comprising 1,010 fine-grained style attributes and train a decoder to reconstruct these prompts from implicit style representations. This approach enables, for the first time, structured and controllable interpretation of text style, supporting tasks such as style recovery, imitation, and guidance. The method significantly outperforms strong baselines across three key evaluations—style prompt recovery, style-controlled text generation, and alignment with human-perceived style—demonstrating enhanced accuracy in style description and improved controllability in generation.
Current copyright detection technologies primarily identify verbatim copying and fail to address “substantial similarity” as protected under EU copyright law—such as stylistic elements and narrative structures—creating a compliance gap. This work proposes PSALM, a novel framework that operationalizes the EU’s legal standard of substantial similarity into a computable, multidimensional assessment system. Leveraging an LLM-as-a-judge architecture, PSALM systematically evaluates infringement risk across ten dimensions, including computational overlap, stylistic resemblance, content alignment, and statutory exceptions. Experiments reveal that fine-tuned models often exhibit systematic stylistic appropriation beyond literal copying; while negative preference optimization reduces overall similarity, residual stylistic traces remain detectable, exposing limitations in current unlearning approaches.