Ai-Facilitated Analysis of Abstracts and Conclusions: Flagging Unsubstantiated Claims and Ambiguous Pronouns

📅 2025-06-16
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses two critical credibility risks in academic texts—unsubstantiated assertions and ambiguous pronominal references—by proposing the first dual-dimensional LLM analysis framework targeting information integrity and linguistic clarity. Methodologically, it introduces a hierarchical reasoning structured prompting scheme integrating high-level semantic parsing and coreference resolution mechanisms, rigorously evaluated across multiple rounds on Gemini Pro 2.5 and ChatGPT Plus o3. Its key contribution lies in empirically uncovering significant performance impacts of syntactic roles, task types, and context interactions—advancing fine-grained adaptation of LLMs for scholarly credibility assessment. Experimental results show 95% accuracy in noun-phrase assertion identification and perfect (100%) coreference resolution in abstracts; however, substantial inter-model disagreement emerges in adjective-modifier judgment (0% vs. 95%), highlighting inherent challenges in fine-grained linguistic analysis.

Technology Category

Natural Language Processing: Discourse, Pragmatics & Argument MiningMachine Learning: Large Multimodal Models (LMMs)Cognitive Modeling & Cognitive Systems: Conceptual Inference and Reasoning

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Web evaluation methodologies and metricsGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
We present and evaluate a suite of proof-of-concept (PoC), structured workflow prompts designed to elicit human-like hierarchical reasoning while guiding Large Language Models (LLMs) in high-level semantic and linguistic analysis of scholarly manuscripts. The prompts target two non-trivial analytical tasks: identifying unsubstantiated claims in summaries (informational integrity) and flagging ambiguous pronoun references (linguistic clarity). We conducted a systematic, multi-run evaluation on two frontier models (Gemini Pro 2.5 Pro and ChatGPT Plus o3) under varied context conditions. Our results for the informational integrity task reveal a significant divergence in model performance: while both models successfully identified an unsubstantiated head of a noun phrase (95% success), ChatGPT consistently failed (0% success) to identify an unsubstantiated adjectival modifier that Gemini correctly flagged (95% success), raising a question regarding potential influence of the target's syntactic role. For the linguistic analysis task, both models performed well (80-90% success) with full manuscript context. In a summary-only setting, however, ChatGPT achieved a perfect (100%) success rate, while Gemini's performance was substantially degraded. Our findings suggest that structured prompting is a viable methodology for complex textual analysis but show that prompt performance may be highly dependent on the interplay between the model, task type, and context, highlighting the need for rigorous, model-specific testing.
Problem

Research questions and friction points this paper is trying to address.

Identifying unsubstantiated claims in scholarly summaries
Flagging ambiguous pronoun references for clarity
Evaluating LLM performance on syntactic and contextual tasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Structured workflow prompts for hierarchical reasoning
Targeting unsubstantiated claims and ambiguous pronouns
Model-specific testing for prompt performance
🔎 Similar Papers
No similar papers found.