The Effect of Perceived Race and Gender on Police Language Use: Experimental Evidence from VR Simulations

📅 2026-08-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether police officers exhibit systematically greater verbal disrespect toward virtual characters perceived as Black males in virtual reality (VR) simulations. By treating character race and gender as experimental manipulations, the work introduces a novel causal inference framework that integrates large language model (LLM)-derived textual features into a mixed-effects model combined with inverse probability of treatment weighting (IPTW) to estimate average treatment effects (ATE) from multilevel textual data. Results reveal that officers direct significantly more disrespectful language toward Black male avatars—particularly when portrayed as suspects—thereby increasing the risk of conversational breakdown. The study also demonstrates the validity and innovation of the proposed LLM-augmented causal inference approach for quantifying linguistic bias in interactive contexts.
📝 Abstract
Against the backdrop of violence in police interactions with the U.S. public, we explore how deferentially police officers speak to virtual characters depicted as Black adult males in vir- tual reality (VR) simulations. We evaluate the effect of seeing and communicating with these characters through a causal in- ference lens, where the assignment of the Black man character to a police officer and simulation is the treatment variable. Our (marginal) average treatment effect AT E measures the social impact of the character on the deference of officer statements with each turn of the conversation. Soberingly, we find that most officers speak less deferentially to Black man characters, except for White, biracial, and multiracial female officers, es- pecially in settings where the VR character was known to be a suspect. Across a full conversation of a typical VR scene, these marginal AT Es can result in notable changes in def- erence of tone (two to several points difference on a scale of 0-10), above and beyond that due to the initial effect of per- ceiving a Black male character. Even more disconcerting is that this can contribute to conversation breakdowns that po- tentially result in violence or danger to both the public and the police. We also explored the capabilities of large language models (LLMs) for ATE estimation. From our methods com- parison analysis, including model validation against synthetic data, we provide unique scientific insights on LLM-assisted methodologies for ATE estimation. As such, for ATE esti- mation with multilevel data with text, we recommend mixed effects models with the inverse propensity treatment weighted (iptw) approach, which utilized an LLM for text feature cre- ation. While we also tested LLMs for finetuning prediction models ultimately for ATE estimation, we conclude they are an area for further development and refinement.
Problem

Research questions and friction points this paper is trying to address.

police language
perceived race
perceived gender
deference
conversation breakdown
Innovation

Methods, ideas, or system contributions that make the work stand out.

causal inference
large language models
virtual reality simulations
average treatment effect
text-based multilevel modeling
🔎 Similar Papers
No similar papers found.