๐ค AI Summary
This study investigates the relationship between linguistic features in written narratives by residents of low-security correctional facilities and their recidivism risk, while quantifying peer effects on language useโparticularly interactive and feedback-oriented expressions. Methodologically, it employs transformer-based large language models (LLMs) to generate text embeddings and integrates zero-shot classification to enhance predictive interpretability. A novel multivariate peer effect estimation framework is proposed, accommodating sparse social networks, latent variables, and multiple correlated outcomes, while addressing network endogeneity. Empirically, LLM-derived embeddings improve recidivism prediction accuracy by 30% over baseline models; statistically significant and robust peer effects are identified in linguistic interactions. This work represents the first integration of LLM-based textual representations with rigorous causal peer effect modeling, yielding new empirical evidence and methodological tools for judicial risk assessment and the study of social influence mechanisms within carceral environments.
๐ Abstract
We find AI embeddings obtained using a pre-trained transformer-based Large Language Model (LLM) of 80,000-120,000 written affirmations and correction exchanges among residents in low-security correctional facilities to be highly predictive of recidivism. The prediction accuracy is 30% higher with embedding vectors than with only pre-entry covariates. However, since the text embedding vectors are high-dimensional, we perform Zero-Shot classification of these texts to a low-dimensional vector of user-defined classes to aid interpretation while retaining the predictive power. To shed light on the social dynamics inside the correctional facilities, we estimate peer effects in these LLM-generated numerical representations of language with a multivariate peer effect model, adjusting for network endogeneity. We develop new methodology and theory for peer effect estimation that accommodate sparse networks, multivariate latent variables, and correlated multivariate outcomes. With these new methods, we find significant peer effects in language usage for interaction and feedback.