🤖 AI Summary
This study investigates whether instruction-tuned large language models (LLMs) adhere to human genre conventions in written style. Methodologically, it employs a linguistics-informed quantitative stylistic analysis—integrating corpus-based statistics with controlled comparative experiments—to systematically characterize LLM output against human-authored texts. Results reveal a previously undocumented, intrinsic “noun-dense” bias in LLMs: their generated text consistently deviates from human norms across grammatical complexity, informational density, and contextual appropriateness—even under informal prompting—and fails to replicate human rhetorical pacing and lexical selection patterns. This finding introduces an interpretable, measurable linguistic dimension for LLM text evaluation, moving beyond opaque, black-box assessment paradigms. It provides both theoretical grounding and empirical evidence for advancing controllable stylistic generation and human–AI collaborative writing systems.
📝 Abstract
Significance As large language models (LLMs) have grown in power and become more widely available, research has focused on their ability to complete various tasks and the biases they exhibit when doing so. In this study, we instead examine their writing style in detail. We show that instruction-tuned models, which are trained to answer questions and solve problems, have a distinct noun-heavy, informationally dense writing style, even when prompted to match the style of informal speech and writing. These findings suggest that instruction-tuned models generate text that does not align with genre conventions familiar to human audiences, and demonstrate the value of linguistic variables in evaluating the output of LLMs.