🤖 AI Summary
This study addresses the high sensitivity of existing AI-generated text detectors to decoding distributions, noting that current evasion methods predominantly rely on training or external feedback signals. To overcome this limitation, this work proposes EASE, a pioneering training-free and detector-agnostic framework for entropy-adaptive distribution shaping. By leveraging only the predictive entropy of the source large language model as an internal signal, the method dynamically modulates logit perturbation intensity and sampling temperature, thereby achieving stealthy generation without requiring any external feedback. Extensive evaluations demonstrate that EASE significantly reduces detection rates across multiple models and diverse detectors while preserving text generation quality, all with negligible inference overhead.
📝 Abstract
AI-generated text (AIGT) detection can be sensitive to the decoding choices of the source large language model (LLM). We observe that perturbing next-token logits or adjusting sampling temperature can reduce detection performance, providing a clear signal of detector vulnerability to decoding-time distribution changes. Building on this observation, we propose EASE (Entropy-Adaptive Distribution Shaping for Evasion), a training-free and detector-agnostic framework for evading AIGT detectors. EASE computes predictive entropy directly from the source LLM's next-token distribution and uses it to adapt both logit perturbation and sampling temperature, without detector feedback or model fine-tuning. Experiments across three source LLMs and multiple detectors demonstrate consistent reductions in detection performance, with negligible degradation in text quality and negligible inference overhead.