STRIVE: Probing Reasoning Limits in Graded Plausibility Generation and Evaluation

📅 2026-08-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high cost of manually constructing slot-controlled event sets for psycholinguistic research. The authors propose STRIVE, a framework that leverages large language models (e.g., GPT-5.1) to systematically generate controlled event sets by varying only a single slot within a fixed event structure, spanning levels of plausibility and difficulty. STRIVE introduces a novel global reasoning scratchpad mechanism and an evaluator-guided iterative refinement strategy, substantially improving generation quality and alignment with human judgments. Experimental results show that the合格 rate of generated event sets increases from 16.7% to 75.0%, with significantly enhanced consistency between the evaluator and human raters. However, accuracy remains at 57% under implausible–difficult conditions, indicating that boundary cases still require manual intervention.
📝 Abstract
Event knowledge concerns who does what to whom. Psycholinguists use event-plausibility judgments to examine how this knowledge supports human language processing. To isolate plausibility effects, these studies require controlled event sets in which one event slot varies across plausibility levels while all other event features remain fixed. Constructing such sets manually is labor-intensive. We therefore introduce STRIVE, an LLM-based framework for jointly generating and evaluating controlled event sets crossing plausibility class (plausible vs. implausible) with intended classification difficulty (easy vs. hard). Given a verb, STRIVE constructs a shared event frame, then produces one event per condition by varying one slot while holding all others fixed. In experiments with six models across 60 verbs, GPT-5.1 produced high-quality sets only 16.7% of the time using the baseline generation prompt. Adding a global reasoning scratchpad and evaluator-guided refinement raised this rate to 75.0%. Greater reasoning effort also improved evaluator--human agreement. Nevertheless, events near the plausibility boundary remain most difficult. They elicit the greatest human disagreement, and the best evaluator reaches only 57% accuracy on the implausible-hard condition, indicating a need for human input. Overall, STRIVE offers a scalable approach to reducing manual effort by automating initial event-set generation and evaluation for psycholinguistic studies.
Problem

Research questions and friction points this paper is trying to address.

event plausibility
controlled event sets
graded plausibility
psycholinguistics
plausibility judgment
Innovation

Methods, ideas, or system contributions that make the work stand out.

controlled event generation
plausibility evaluation
reasoning scratchpad
evaluator-guided refinement
psycholinguistic probing