Safety Nudges: User-Facing Interventions for Real-Time AI Risk Awareness

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究通过开发Safety Nudges工具,实时标记聊天机器人对话中的风险行为,以提高用户对AI潜在危害的意识,尽管这未直接改变用户行为。
📝 Abstract
Conversational AI systems can pose safety risks to their users such as hallucination, sycophancy, overconfidence, and anthropomorphism, but these risks are difficult for users to detect during everyday use. We introduce Safety Nudges, a browser-based tool that provides lightweight, in situ flags when concerning behavior is detected in chatbot conversations. We evaluated Safety Nudges in a two-week field study with 45 frequent chatbot users, collecting interaction logs, surveys, and feedback on individual nudges. Participants found the tool useful, clear, and minimally disruptive, with nearly all users reporting an increased awareness of potential AI harms, though we found that this improved awareness alone did not necessarily lead to discernible behavioral changes. Our results suggest that user facing safety nudges can complement model-level safeguards by helping people critically evaluate AI responses in context, while highlighting the importance of relevance, calibration, and user control in nudge design for conversational AI safety.
Problem

Research questions and friction points this paper is trying to address.

Safety Risks
Conversational AI
User Awareness
Innovation

Methods, ideas, or system contributions that make the work stand out.

Safety Nudges
real-time AI risk awareness
in situ flags
chatbot conversations
user control
🔎 Similar Papers
No similar papers found.