🤖 AI Summary
This study investigates how textual style variations introduced by LLM-based paraphrasing affect the decision stability of multimodal claim verification. Leveraging eleven open-source vision-language models, we design two rewriting strategies—natural polishing and controlled single-point lexical injection—and systematically evaluate their effects on verification accuracy and probability distributions through controlled experiments and statistical significance analysis. Our findings reveal that claim verification is more robust than rating manipulation; while most models exhibit no significant accuracy degradation and grammar-correction rewrites exert minimal influence, ambiguity-oriented phrasing induces widespread and statistically significant probability shifts. This work provides empirical evidence for enhancing the textual robustness of multimodal verification systems.
📝 Abstract
LLMs are known to introduce stylistic changes into generated text, yet how these stylistic shifts affect model decisions on scientific tasks remains underexplored. In this paper, we focus on multimodal claim verification, where the goal is to determine whether a textual claim is grounded in a given piece of evidence. We apply two rewriting strategies: natural rewriting, which simulates how researchers routinely use LLMs to polish academic text, and controlled injection, which inserts a single LLM-associated word to isolate the effect of vocabulary choice. We evaluate 11 open-weight models spanning five VLM families and ranging from 2B to 38B parameters. We find that models are robust to these modifications: most show no significant drop in accuracy, and compared to prior work on review-score manipulation, verification appears far more stable. However, consistent probability shifts do occur. Hedging-oriented conditions produce significant shifts across nearly all models, while boosting conditions show a weaker effect and general polishing conditions (e.g., grammar correction, fluency improvement) have little effect.