AI Watchdog: Agent Interfaces for Detecting and Defending Against Manipulative Dark Patterns in AI Conversations

📅 2026-08-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究通过AI Watchdog界面监测对话中五种操纵性暗模式,并以即时警告方式有效减少用户对含有暗模式的AI建议的遵从度。
📝 Abstract
Conversational AI increasingly shapes consequential decisions, yet users have limited support for recognizing and resisting manipulation. We present AI Watchdog, a browser-based agent interface that monitors live conversations, detects five dark-pattern categories, including sycophancy, brand bias, anthropomorphization, sneaking, and harmful generation, and alerts users when they occur. Its open-weight turn-level classifier supports independent deployment and a path toward local inference, preserving user privacy while remaining separate from the conversational AI. We evaluated AI Watchdog in a preregistered, five-condition between-subjects experiment (N = 150) comparing a no-intervention control with four configurations varying nudge timing (prebunking vs. just-in-time) and engagement mode (without vs. with cognitive forcing). Results show that participants rarely flagged manipulative turns across all conditions, and post-task awareness did not differ significantly across groups. However, just-in-time warnings without cognitive forcing were the only intervention to significantly reduce compliance with AI-steered recommendations containing dark patterns, lowering compliance from 71.7% to 53.7%, an 18 percentage-point reduction. Exploratory analyses further showed that lower misinformation susceptibility was associated with greater flagging but not lower compliance, while higher AI trust was associated with greater compliance and lower reported awareness. Together, these findings suggest that explicit recognition of conversational dark patterns and behavioral resistance to AI steering may be distinct outcomes, motivating further investigation of timely, low-friction defensive interfaces.
Problem

Research questions and friction points this paper is trying to address.

manipulative dark patterns
conversational AI
user support
Innovation

Methods, ideas, or system contributions that make the work stand out.

AI Watchdog
dark patterns
turn-level classifier
just-in-time warnings
local inference
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
R
Rachel Poonsiriwong
MIT Media Lab, Cambridge, MA, USA
C
Chayapatr Archiwaranguprok
MIT Media Lab, Cambridge, MA, USA
C
Constanze Albrecht
MIT Media Lab, Cambridge, MA, USA
M
Monchai Lertsutthiwong
KASIKORN Labs, Nonthaburi, Thailand
Pattie Maes
Pattie Maes
Professor of Media Arts and Sciences, MIT
human computer interactionartificial intelligencedigital health
Pat Pataranutaporn
Pat Pataranutaporn
PhD Researcher at MIT Media Lab, Massachusetts Institute of Technology
Human-AI InteractionCyborg PsychologyHuman FlourishingBio-Digital Interfaces