π€ AI Summary
This study addresses the systematic stylistic and semantic deviations of large language model (LLM)-generated text from human writing, particularly its lack of literary expressiveness. For the first time, it systematically identifies consistent n-gram distribution patterns in LLM outputs and demonstrates, through combined statistical analysis and qualitative comparison, that this stylistic impoverishment is closely linked to constrained semantic expression. Challenging the long-standing assumption that style and semantics are separable, the work elucidates their intrinsic coupling mechanism, offering a novel perspective for understanding the βnon-stylisticβ nature of LLM-generated text.
π Abstract
Prior work on LLM-generated text has demonstrated quantitative and qualitative departures from text produced by humans. LLM-generated texts differ from human writing in style, resulting in a characteristic textual "feel," while the semantic range of LLMs is much restricted compared to that of humans. In this contribution, I note simple but consistent patterns in the statistical distribution of n-grams within LLM-generated text. Via qualitative analysis of these n-grams, I reveal deficiencies in LLM style. Because higher-order n-grams correlate to semantic content, I conclude that questions of style and semantics are not cleanly separable.