Detecting Spin in Clinical Trials with Large Language Models

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of automatically detecting outcome spin and outcome switching in clinical trials by proposing a large language model (LLM)-based automated detection framework. The proposed approach integrates prompt engineering, token probability classification, majority voting mechanisms, and semantic similarity computation to precisely identify outcome switching while generating interpretable textual explanations. Experimental results demonstrate that the framework achieves an accuracy of 0.90 and an F1 score of 0.78 on the test set, significantly outperforming baseline models. By innovatively combining the reasoning capabilities of LLMs with semantic features, this work enables high-precision spin detection alongside natural language explanation generation, providing an effective assistive tool for clinical review processes.
📝 Abstract
Spin in clinical trials includes reporting practices that distort the presentation of results. This is particularly critical in medicine, where spin is present in more than 50% of randomized controlled trials that fail to reach statistical significance. The comparison of primary and reported outcomes is crucial for detecting several types of spin, including outcome switching. We used 300 pairs of outcomes labeled with semantic similarity to develop a system for automatic detection of outcome switching. We evaluated baseline text similarity models and open-source LLMs using generated similarity scores and the Youden index to determine the classification threshold. The proposed approach involves prompt engineering, classification based on token probabilities, and majority voting for the final decision. The results on the test set of 2,496 examples with an F1 score of 0.78 and an accuracy of 0.90 outperform baseline text similarity models but trail behind fine-tuned versions of BERT. We used LLMs to generate natural language explanations for the classified instances and manually assessed their quality.
Problem

Research questions and friction points this paper is trying to address.

Clinical trials
Spin detection
Outcome switching
Large Language Models
Semantic similarity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Large Language Models
Outcome Switching
Prompt Engineering
Token Probabilities
Clinical Trials
🔎 Similar Papers
No similar papers found.