Who Wrote It Is Not Enough: Detecting Who Contributed the Insight

πŸ“… 2026-10-05
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limitation of existing research that predominantly focuses on text authorship detection while neglecting the attribution of intellectual contributions underlying insights in scientific discourse. To this end, it proposes a novel task termed β€œinsight provenance tracing,” which aims to precisely identify whether an insight originates from humans, large language models (LLMs), or hybrid sources. Methodologically, the authors construct the InsightProv dataset by leveraging GPT-4o for simulated data generation and sentence-level annotation, and introduce a two-stage adversarial learning framework designed to effectively mitigate model reliance on superficial shortcut signals. The findings reveal fundamental disparities between AI and human agents in knowledge introduction and failure modes, demonstrating that the origin of ideas is more persistently identifiable than surface phrasing. Ultimately, this work establishes a new paradigm for evaluating intellectual contributions in academia.
πŸ“ Abstract
As LLMs increasingly assist scientific writing and peer review, detecting who wrote the text is no longer sufficient: we need to determine who contributed the underlying insight. We introduce Insight Provenance, the task of identifying whether a review insight originates from a human, an LLM, or their hybrid contribution. We construct InsightProv-v0 from 4,057 scientific papers and 12,660 human reviews, simulating different levels of LLM involvement with GPT-4o, Gemini, and DeepSeek and annotating provenance at the sentence level. We show that strong performance on raw data can be misleading, as models exploit linguistic and textual-authorship shortcuts that degrade substantially under progressively debiased evaluation. We therefore propose a two-stage adversarial framework that suppresses shortcut signals while preserving provenance-relevant information. Beyond detection, extensive analyses reveal what makes intellectual authorship identifiable: paper grounding and neighboring review context provide complementary provenance signals, while human, hybrid, and AI insights systematically differ in their information sources and failure modes. Most strikingly, AI insights predominantly remain close to generic or paper-provided information, whereas human insights more often introduce external knowledge and independent judgment. These findings suggest that while wording can be rewritten by an LLM, the provenance of an idea leaves a deeper and more persistent signal.
Problem

Research questions and friction points this paper is trying to address.

Insight Provenance
Large Language Models
Scientific Peer Review
Intellectual Authorship
Contribution Detection
Innovation

Methods, ideas, or system contributions that make the work stand out.

Insight Provenance
Adversarial Framework
Shortcut Suppression
LLM Detection
Scientific Peer Review
πŸ”Ž Similar Papers
No similar papers found.