Measuring Human Value Expression in Social Media Texts: Calibrated LLM Annotation and Encoder Transfer

📅 2026-06-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of reliably and scalably measuring expressions of human basic values—grounded in Schwartz’s theory—in unstructured social media text. To this end, the authors propose a theory-constrained calibration mechanism for large language model (LLM) annotations, integrating multilingual prompt engineering, iterative error analysis, and expert-validated rules to generate soft labels that preserve semantic structure and uncertainty. These calibrated labels are then transferred to an encoder-based model to enable large-scale value prediction. The approach significantly reduces value misclassification rates, enhances agreement with expert annotations, and improves alignment with the theoretical value structure, thereby establishing a scalable framework for value detection that maintains strong fidelity to Schwartz’s theory.
📝 Abstract
Measuring subjective constructs in naturally occurring social media text requires annotation procedures that are theoretically grounded, empirically validated, and transferable to an encoder model for scalable prediction. Using non-English social media posts annotated according to Schwartz's theory of basic human values, we investigate how different LLMs, prompts, and instruction languages operationalize the expression of values in text. We argue that although texts may permit multiple plausible interpretations, theory-based value definitions can constrain interpretations and reduce spurious value attributions. Beyond precision, recall, and F1, we evaluate structural alignment between values, error structure, confidence-ambiguity relations, and annotation stability. We show that different LLMs produce different value interpretations. Iterative prompt calibration through error analysis reduces misattributions and improves alignment with expert annotations. We also derive targeted expert verification rules from recurrent error structures and use them during corpus annotation. Finally, we show that LLM annotations can be transferred to an encoder model through soft-label training, retaining theory-based value interpretations and information about uncertainty in value expression.
Problem

Research questions and friction points this paper is trying to address.

human values
social media text
value measurement
annotation
theory-based interpretation
Innovation

Methods, ideas, or system contributions that make the work stand out.

calibrated LLM annotation
value expression measurement
prompt calibration
soft-label transfer
theory-guided NLP
🔎 Similar Papers
No similar papers found.
M
Maria Milkova
Independent researcher, Lisbon, Portugal
M
Maksim Rudnev
University of Waterloo, ON, Canada