Fighting Fire with Fire: On the Feasibility of Protecting Exercises Against AI Cheating

📅 2026-08-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the growing concern of students exploiting generative AI to complete assignments, which obscures their true competencies and undermines independent thinking. To counter this, the authors propose a novel anti-cheating method that embeds subtle adversarial perturbations into images of multimodal multiple-choice questions. These perturbations are designed to systematically induce mainstream multimodal large language models—such as ChatGPT, Claude, and Gemini—into producing predetermined incorrect answers under black-box conditions. The resulting response patterns serve as statistically detectable “answer fingerprints.” This work pioneers the application of adversarial machine learning to academic integrity verification, leveraging surrogate models to optimize perturbations and employing hypothesis testing to reliably identify rote copying behavior, thereby establishing a scalable technical safeguard for scholarly honesty in the age of AI.
📝 Abstract
The widespread adoption of generative AI enables students to outsource cognitive effort to increasingly capable assistants, creating an illusion of competence while undermining the independent reasoning that education aims to cultivate. We investigate whether adversarial machine learning can be repurposed to protect educational exercises against such corrosive reliance. Our approach uses multimodal multiple-choice questions whose visual components can be protected with subtle visual perturbations that steer AI solvers toward designated incorrect answers. These responses form a statistical fingerprint: students who blindly copy a solver reproduce the induced answer pattern more frequently than genuine students. We study the feasibility of this paradigm under realistic black-box assistant assumptions using three of the most common state-of-the-art multimodal language models: Anthropic's Claude, Google's Gemini, and OpenAI's ChatGPT. By using accessible surrogate models, we optimize adversarial perturbations that induce consistent response patterns. Those patterns enable principled detection through statistical hypothesis testing. These findings establish both the promise and the limitations of fighting machine-assisted reasoning with the vulnerabilities of the machines themselves.
Problem

Research questions and friction points this paper is trying to address.

AI cheating
educational integrity
generative AI
student assessment
adversarial protection
Innovation

Methods, ideas, or system contributions that make the work stand out.

adversarial machine learning
AI cheating detection
multimodal educational assessment
visual perturbations
statistical fingerprinting