🤖 AI Summary
This study investigates whether professional translators without specialized training can reliably distinguish AI-generated Italian short stories from human-authored ones. In an offline experiment, 69 translators evaluated three anonymized texts—two produced by ChatGPT-4o and one written by a human—assessing their origin and providing justifications. As the first empirical examination of AI-text detection among a real-world cohort of professional translators, the research integrates quantitative scoring with qualitative analysis to uncover the dual influence of analytical reasoning and subjective preferences on identification accuracy. Findings reveal that only 16.2% of participants performed significantly above chance level. Effective diagnostic cues included low burstiness, narrative inconsistencies, and traces of English-language transfer, whereas high grammatical accuracy and emotional tone frequently led to misclassification.
📝 Abstract
This study investigates whether professional translators can reliably identify short stories generated in Italian by artificial intelligence (AI) without prior specialized training. Sixty-nine translators took part in an in-person experiment, where they assessed three anonymized short stories - two written by ChatGPT-4o and one by a human author. For each story, participants rated the likelihood of AI authorship and provided justifications for their choices. While average results were inconclusive, a statistically significant subset (16.2%) successfully distinguished the synthetic texts from the human text, suggesting that their judgements were informed by analytical skill rather than chance. However, a nearly equal number misclassified the texts in the opposite direction, often relying on subjective impressions rather than objective markers, possibly reflecting a reader preference for AI-generated texts. Low burstiness and narrative contradiction emerged as the most reliable indicators of synthetic authorship, with unexpected calques, semantic loans and syntactic transfer from English also reported. In contrast, features such as grammatical accuracy and emotional tone frequently led to misclassification. These findings raise questions about the role and scope of synthetic-text editing in professional contexts.