Do LLMs write like humans? Variation in grammatical and rhetorical styles

📅 2024-10-21
🏛️ Proceedings of the National Academy of Sciences of the United States of America
📈 Citations: 15
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether instruction-tuned large language models (LLMs) adhere to human genre conventions in written style. Methodologically, it employs a linguistics-informed quantitative stylistic analysis—integrating corpus-based statistics with controlled comparative experiments—to systematically characterize LLM output against human-authored texts. Results reveal a previously undocumented, intrinsic “noun-dense” bias in LLMs: their generated text consistently deviates from human norms across grammatical complexity, informational density, and contextual appropriateness—even under informal prompting—and fails to replicate human rhetorical pacing and lexical selection patterns. This finding introduces an interpretable, measurable linguistic dimension for LLM text evaluation, moving beyond opaque, black-box assessment paradigms. It provides both theoretical grounding and empirical evidence for advancing controllable stylistic generation and human–AI collaborative writing systems.

Technology Category

Natural Language Processing: Sentiment Analysis, Stylistic Analysis, and Argument MiningMachine Learning: Large Multimodal Models (LMMs)Humans and AI: Learning Human Values and Preferences

Application Category

User Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systemsSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactions
📝 Abstract
Significance As large language models (LLMs) have grown in power and become more widely available, research has focused on their ability to complete various tasks and the biases they exhibit when doing so. In this study, we instead examine their writing style in detail. We show that instruction-tuned models, which are trained to answer questions and solve problems, have a distinct noun-heavy, informationally dense writing style, even when prompted to match the style of informal speech and writing. These findings suggest that instruction-tuned models generate text that does not align with genre conventions familiar to human audiences, and demonstrate the value of linguistic variables in evaluating the output of LLMs.
Problem

Research questions and friction points this paper is trying to address.

Analyzing rhetorical style differences between LLMs and humans
Identifying systematic grammatical variations in AI-generated text
Detecting LLM output through advanced linguistic feature analysis
Innovation

Methods, ideas, or system contributions that make the work stand out.

Analyzed rhetorical styles using Biber's features
Constructed parallel corpora from human and LLM texts
Identified systematic stylistic differences across model variants
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Carnegie Mellon University | New Jersey Institute of Technology | Heinz College of Information Systems and Public Policy
Alex Reinhart
Alex Reinhart
Carnegie Mellon University
Statistics & Data Science
D
David West Brown
Department of English, Carnegie Mellon University
B
Ben Markey
Department of English, Carnegie Mellon University
M
Michael Laudenbach
Department of Humanities & Social Sciences, New Jersey Institute of Technology
K
Kachatad Pantusen
Department of Statistics & Data Science, Carnegie Mellon University
K
Kachatad Pantusen
Heinz College of Information Systems and Public Policy, Carnegie Mellon University
R
Ronald Yurko
Department of Statistics & Data Science, Carnegie Mellon University
G
Gordon Weinberg
Department of Statistics & Data Science, Carnegie Mellon University