🤖 AI Summary
This paper identifies critical, underexamined sociotechnical challenges in AI red teaming: its implicit value assumptions, opaque labor organization, and substantial psychological burden on red team members. Employing qualitative social science methods—including critical comparative analysis with content moderation practices and the Value Sensitive Design framework—the study systematically reframes red teaming as a sociotechnical practice requiring interdisciplinary scrutiny, rather than a purely technical security assessment. Its core contributions are threefold: (1) it moves beyond technocentric paradigms to expose structural risks embedded in current red teaming practices; (2) it advances a human-centered approach to AI safety governance; and (3) it calls for substantive collaboration between computer science and the social sciences to preempt the ethical harms and labor exploitation already evident in content moderation—thereby supporting the development of sustainable, accountable AI evaluation systems.
📝 Abstract
As generative AI technologies find more and more real-world applications, the importance of testing their performance and safety seems paramount."Red-teaming"has quickly become the primary approach to test AI models--prioritized by AI companies, and enshrined in AI policy and regulation. Members of red teams act as adversaries, probing AI systems to test their safety mechanisms and uncover vulnerabilities. Yet we know far too little about this work or its implications. This essay calls for collaboration between computer scientists and social scientists to study the sociotechnical systems surrounding AI technologies, including the work of red-teaming, to avoid repeating the mistakes of the recent past. We highlight the importance of understanding the values and assumptions behind red-teaming, the labor arrangements involved, and the psychological impacts on red-teamers, drawing insights from the lessons learned around the work of content moderation.