C-HAT-Bench: Benchmarking Chinese AI-Text Detection Beyond Fully Generated Text

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the overestimated reliability of existing detectors in Chinese human-machine collaborative text scenarios and the lack of cross-domain, multi-generation benchmarks. We construct the first Chinese benchmark for detecting such texts, linking 5,000 source documents to over 240,000 variants. Leveraging six generative models, we transcend binary detection limitations by introducing prefix continuation and three collaboration modes. Twenty-one detectors are systematically evaluated across four protocols: zero-shot detection, supervised detection, boundary localization, and cross-condition generalization. Results demonstrate that average AUROC declines by 12.0% under collaboration modes, with a maximum drop of 44.4%, alongside significantly asymmetric performance transferability. This work fills a critical gap in cross-condition generalization evaluation for Chinese machine-generated text detection.
📝 Abstract
Large Language Models (LLMs) increasingly participate in writing by modifying or extending human drafts, causing machine involvement to vary in both form and extent. Yet most Machine-Generated Text (MGT) detectors are evaluated only on fully human-written versus fully AI-generated text. Because human--AI collaboration can weaken or redistribute cues associated with machine generation, strong performance under this binary setting may overstate detector reliability. This mismatch remains underexplored in Chinese: detection cues are shaped by tokenization and language-specific text distributions, yet controlled resources spanning production settings, domains, and generators remain limited. To fill this gap, we present a Chinese Human-AI Collaborative Text Detection Benchmark (C-HAT-Bench), a unified benchmark that links $5,000$ human-written source texts from five domains to more than $240,000$ variants produced using six generative models under Prefix-Conditioned Continuation as a reference setting and three collaborative production modes. We evaluate $21$ detectors through four protocols spanning zero-shot and pretrained supervised document-level detection, boundary localization, and cross-condition generalization. Relative to Prefix-Conditioned Continuation, mean AUROC across document-level detectors is $12.0\%$ lower on the collaborative production modes, with the largest detector-specific relative decrease reaching $44.4\%$. Transfer across collaborative production modes is also asymmetric, indicating that performance in a given production setting is not a reliable predictor of performance in other production settings.
Problem

Research questions and friction points this paper is trying to address.

Machine-Generated Text Detection
Human-AI Collaboration
Chinese Benchmark
Large Language Models
Text Attribution
Innovation

Methods, ideas, or system contributions that make the work stand out.

Human-AI Collaborative Text Detection
Chinese Benchmark
Machine-Generated Text
Cross-condition Generalization
Boundary Localization
🔎 Similar Papers
2024-06-21Journal of Artificial Intelligence ResearchCitations: 6
💼 Related Jobs
No related jobs found.
Q
Qing Yang
Guilin University of Electronic Technology, China
Zixiang Luo
Zixiang Luo
The Hong Kong University of Science and Technology
Computational Neuroscience
Z
Zhenyu Mao
Guilin University of Electronic Technology, China
Z
Zezheng Wu
Guilin University of Electronic Technology, China
X
Xinghe Cheng
Jinan University, China
H
Haibo Chen
Nanjing University of Science and Technology, China
Qinggang Zhang
Qinggang Zhang
The Hong Kong Polytechnic University
Knowledge GraphsLarge Language ModelsRetrieval-Augmented GenerationText-to-SQL
J
Jiapu Wang
Nanjing University of Science and Technology, China
J
Jingwei Zhang
Guilin University of Electronic Technology, China