LA-CPD: Local-Evidence-Aware Change-Point Detection for Human-LLM Authorship Segmentation

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of accurately localizing authorship boundaries in human-AI co-authored documents, where noise complicates attribution. We propose a precise provenance tracing method based on change-point detection and dynamic programming. This approach converts sentence-level detection scores into coherent paragraph partitions by integrating length-weighted residuals with a window-based double-mean comparison to suppress spurious boundaries caused by content fluctuations. Furthermore, the Akaike Information Criterion (AIC) and a local evidence-aware algorithm are introduced to optimize cut-point selection. Experimental results demonstrate that sentence-level accuracy improves from 0.747 to 0.796, significantly enhancing boundary localization precision and the identification of spans generated by large language models.
📝 Abstract
As LLM-generated text becomes increasingly human-like, accurately localizing LLM-authored spans in human-LLM co-authored documents is important for attribution and accountability in cases involving copyright infringement, fraud, and other harmful uses of AI-generated content. Sentence-level detectors provide local authorship evidence, but content variation can cause score fluctuations even among sentences from the same source, creating spurious boundaries. Recovering a coherent document partition therefore remains challenging when both the number and locations of authorship transitions are unknown. We propose Local-Evidence-Aware Change-Point Detection (LA-CPD), a structured method that transforms noisy sentence-level score sequences into coherent authorship segments. Given scores from a frozen local detector, LA-CPD combines a length-weighted within-segment residual with a windowed two-mean contrast to capture segment consistency and sustained changes around candidate cut points. Dynamic programming optimizes cut locations for each candidate count, while an AIC-style criterion selects the final partition, yielding sentence labels, authorship boundaries, and maximal LLM-authored spans. On a held-out human-LLM co-authored test set, LA-CPD outperforms WCP+AIC, increasing sentence-level accuracy from 0.747 to 0.796 while improving boundary localization and LLM-span delineation.
Problem

Research questions and friction points this paper is trying to address.

Authorship Segmentation
Change-Point Detection
Human-LLM Co-authorship
LLM-generated Text
Attribution
Innovation

Methods, ideas, or system contributions that make the work stand out.

Change-Point Detection
Authorship Segmentation
Dynamic Programming
Local-Evidence-Aware
Human-LLM Co-authorship
💼 Related Jobs
No related jobs found.
Q
Qing Yang
Guilin University of Electronic Technology, China
Z
Zhenyu Mao
Guilin University of Electronic Technology, China
Zixiang Luo
Zixiang Luo
The Hong Kong University of Science and Technology
Computational Neuroscience
Z
Zezheng Wu
Guilin University of Electronic Technology, China
X
Xinghe Cheng
Jinan University, China
Qinggang Zhang
Qinggang Zhang
The Hong Kong Polytechnic University
Knowledge GraphsLarge Language ModelsRetrieval-Augmented GenerationText-to-SQL
J
Jingwei Zhang
Guilin University of Electronic Technology, China
J
Jiapu Wang
Nanjing University of Science and Technology, China