Linguistic Characteristics of AI-Generated Text: A Survey

📅 2025-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses key limitations in AI-generated text linguistics research—namely, insufficient cross-lingual and cross-model coverage, and a lack of systematic investigation into prompt sensitivity. We systematically review corpus- and computational linguistics literature (2020–2024) to develop a multidimensional classification framework encompassing lexical, syntactic, semantic, stylistic, and diversity dimensions. Through mixed-methods corpus analysis, we identify consistent linguistic patterns in AI outputs: heightened formality and impersonality, elevated noun and article usage, scarcity of adjectives and adverbs, low lexical diversity, and high repetition. Our contribution is threefold: (1) the first integrated empirical synthesis across languages (including Chinese, German, Spanish), models (beyond GPT-series), and genres; (2) explicit demonstration of how prompt engineering critically shapes output characteristics; and (3) provision of a theoretically grounded, methodologically robust foundation for multilingual AI-text detection, evaluation, and controllable generation.

Technology Category

Natural Language Processing: Sentiment Analysis, Stylistic Analysis, and Argument MiningMachine Learning: Large Multimodal Models (LMMs)Humans and AI: AI for Accessibility

Application Category

Search and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchSocial Networks and Social Media: Generative AI / large language models and their impact on social systemsWeb Mining and Content Analysis: Mining multimedia, multimodal, multilingual, cross-lingual Web data
📝 Abstract
Large language models (LLMs) are solidifying their position in the modern world as effective tools for the automatic generation of text. Their use is quickly becoming commonplace in fields such as education, healthcare, and scientific research. There is a growing need to study the linguistic features present in AI-generated text, as the increasing presence of such texts has profound implications in various disciplines such as corpus linguistics, computational linguistics, and natural language processing. Many observations have already been made, however a broader synthesis of the findings made so far is required to provide a better understanding of the topic. The present survey paper aims to provide such a synthesis of extant research. We categorize the existing works along several dimensions, including the levels of linguistic description, the models included, the genres analyzed, the languages analyzed, and the approach to prompting. Additionally, the same scheme is used to present the findings made so far and expose the current trends followed by researchers. Among the most-often reported findings is the observation that AI-generated text is more likely to contain a more formal and impersonal style, signaled by the increased presence of nouns, determiners, and adpositions and the lower reliance on adjectives and adverbs. AI-generated text is also more likely to feature a lower lexical diversity, a smaller vocabulary size, and repetitive text. Current research, however, remains heavily concentrated on English data and mostly on text generated by the GPT model family, highlighting the need for broader cross-linguistic and cross-model investigation. In most cases authors also fail to address the issue of prompt sensitivity, leaving much room for future studies that employ multiple prompt wordings in the text generation phase.
Problem

Research questions and friction points this paper is trying to address.

Analyzing linguistic features of AI-generated text across disciplines
Synthesizing existing research on AI text characteristics systematically
Identifying gaps in cross-linguistic and cross-model AI text analysis
Innovation

Methods, ideas, or system contributions that make the work stand out.

Survey categorizes research by linguistic levels and models
Analyzes AI text features across genres and languages
Identifies need for cross-linguistic and prompt variation studies
🔎 Similar Papers
2024-06-21Journal of Artificial Intelligence ResearchCitations: 6
💼 Related Jobs
No related jobs found.
L
Luka Terčon
Faculty of Arts, University of Ljubljana; Faculty of Computer and Information Science, University of Ljubljana
K
Kaja Dobrovoljc
Faculty of Arts, University of Ljubljana; Jožef Stefan Institute