Recidivism and Peer Influence with LLM Text Embeddings in Low Security Correctional Facilities

๐Ÿ“… 2025-09-24
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study investigates the relationship between linguistic features in written narratives by residents of low-security correctional facilities and their recidivism risk, while quantifying peer effects on language useโ€”particularly interactive and feedback-oriented expressions. Methodologically, it employs transformer-based large language models (LLMs) to generate text embeddings and integrates zero-shot classification to enhance predictive interpretability. A novel multivariate peer effect estimation framework is proposed, accommodating sparse social networks, latent variables, and multiple correlated outcomes, while addressing network endogeneity. Empirically, LLM-derived embeddings improve recidivism prediction accuracy by 30% over baseline models; statistically significant and robust peer effects are identified in linguistic interactions. This work represents the first integration of LLM-based textual representations with rigorous causal peer effect modeling, yielding new empirical evidence and methodological tools for judicial risk assessment and the study of social influence mechanisms within carceral environments.

Technology Category

Natural Language Processing: (Large) Language ModelsMachine Learning: Large Multimodal Models (LMMs)Data Mining & Knowledge Management: Conversational Systems for Recommendation & Retrieval

Application Category

User Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSocial Networks and Social Media: Generative AI / large language models and their impact on social systemsGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
๐Ÿ“ Abstract
We find AI embeddings obtained using a pre-trained transformer-based Large Language Model (LLM) of 80,000-120,000 written affirmations and correction exchanges among residents in low-security correctional facilities to be highly predictive of recidivism. The prediction accuracy is 30% higher with embedding vectors than with only pre-entry covariates. However, since the text embedding vectors are high-dimensional, we perform Zero-Shot classification of these texts to a low-dimensional vector of user-defined classes to aid interpretation while retaining the predictive power. To shed light on the social dynamics inside the correctional facilities, we estimate peer effects in these LLM-generated numerical representations of language with a multivariate peer effect model, adjusting for network endogeneity. We develop new methodology and theory for peer effect estimation that accommodate sparse networks, multivariate latent variables, and correlated multivariate outcomes. With these new methods, we find significant peer effects in language usage for interaction and feedback.
Problem

Research questions and friction points this paper is trying to address.

Predicting recidivism using LLM text embeddings from correctional facility communications
Interpreting high-dimensional text data through zero-shot classification methods
Estimating peer effects in language usage within correctional facility social networks
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM embeddings predict recidivism from correctional texts
Zero-shot classification reduces dimensionality while preserving accuracy
Novel peer effect model analyzes language influence in networks
๐Ÿ’ผ Related Jobs
No related jobs found.
Shanjukta Nath
Shanjukta Nath
Assistant Professor at University of Georgia
Labor EconomicsDevelopment EconomicsEconometrics.
Jiwon Hong
Jiwon Hong
Senior Research Fellow, University of Auckland
Acute diseaselymphaticsextracellular vesiclecancermicrogravity
J
Jae Ho Chang
Department of Statistics, The Ohio State University
K
Keith Warren
College of Social Work, The Ohio State University
S
Subhadeep Paul
Department of Statistics, The Ohio State University