Score
Use close reading to produce detailed, evidence-based analyses of texts by annotating specific words, phrases, and contexts; comparing concepts and structures across passages; identifying and interpreting evaluative language and rhetorical moves; and synthesizing those observations into grounded interpretive claims about the text.
This work addresses the absence of standardized evaluation metrics for literary close reading—the ability to perform nuanced, interpretive textual analysis—in large language models (LLMs). We introduce KRISTEVA, the first benchmark explicitly designed for interpretive reasoning via close reading. Comprising 1,331 multiple-choice questions derived from authentic classroom materials, KRISTEVA features three progressively complex task categories: stylistic feature extraction, contextual knowledge retrieval, and style–context multi-hop reasoning. Crucially, it formalizes the humanities’ core competency of “close reading” as a quantifiable, task-based assessment. Empirical evaluation across 11 subtasks reveals that current state-of-the-art LLMs underperform human experts on 10 tasks (accuracy: 49.7%–69.7%), exposing fundamental limitations in literary interpretive reasoning. KRISTEVA thus fills a critical gap in evaluating deep textual understanding and literary interpretation—capabilities previously unaddressed by existing NLP benchmarks.
This study investigates the impact of AI assistance on human performance in close reading of poetry and associated reading enjoyment. Through a preregistered randomized controlled experiment (N = 400), it compares three conditions: no AI support, a single AI interpretation, and multiple AI interpretations. The research provides the first empirical evidence of a “less-is-more” effect in AI-assisted cultural interpretation: a single AI interpretation significantly enhances both close reading performance and reading enjoyment, whereas multiple interpretations, while further improving performance, diminish enjoyment due to increased overreliance on AI. These findings offer critical empirical insights into the boundaries and optimization pathways of human–AI collaboration in hermeneutic tasks within the humanities.
This work addresses a critical limitation in current reading augmentation systems, which predominantly prioritize information transmission while neglecting the reader’s active role in interpretation and meaning-making during scholarly reading—thereby constraining critical thinking and creative transformation. To counter this efficiency-driven paradigm, the project introduces the concept of “creative reading,” reconceptualizing reading as a generative process co-constituted by the reader’s evolving identity and pluralistic interpretations. Integrating literary narrative theory, humanistic HCI approaches, and design principles from creativity support tools, the study proposes an interaction mechanism centered on provocation. It articulates the first design space for reading augmentation explicitly oriented toward long-term cognitive development and the emergence of diverse meanings, offering both theoretical grounding and practical pathways to support readers’ transformative engagement with texts.
The exponential growth of scientific literature has led to information overload for researchers, while existing large language model (LLM)-generated summaries often suffer from verbosity and over-generalization, impeding deep comprehension of source texts. Method: We propose an AI reading assistant embedded with expert scholarly reading paradigms, whose primary output is a structured “literature map”—a navigable, hierarchical导读—not a reductive summary. Our approach employs a prompt-driven, domain-adapted system architecture that integrates discipline-specific reading strategies (e.g., problem–method–evidence–inference decomposition) to enable targeted information extraction and multi-level presentation. Contribution/Results: Empirical evaluation demonstrates that our system produces significantly more structured, actionable, and faithful literature maps than general-purpose LLMs. It enhances reading efficiency and fosters critical analytical capabilities without compromising textual fidelity, establishing a novel paradigm for scholarly reading support tools.
This paper addresses the insufficient modeling of textual implicit semantics by proposing a novel paradigm that explicitly transforms “subtext” into a verifiable set of propositions. Methodologically, it systematically leverages large language models to generate implicit inference propositions from text, which are then validated for plausibility by human annotators; the resulting proposition semantics are subsequently fused into the original text representations. Key contributions include: (1) introducing the first annotated framework for implicit propositions tailored to social science tasks; and (2) demonstrating that this explicit modeling significantly outperforms literal-only representations across three distinct tasks—argument similarity assessment, public opinion interpretation, and legislative behavior simulation—with average improvements of 12.7% in F1 or accuracy. Results substantiate the effectiveness and generalizability of structured implicit semantic modeling for enhancing human-like semantic understanding.
This work addresses the limitation of existing cultural benchmarks, which predominantly assess factual knowledge while neglecting deeper reasoning capabilities such as explaining, substantiating, and revising cultural references. Using literary interpretation as the evaluation context, this study proposes an evidence-centered benchmark that systematically examines models’ deep cultural understanding through the cross-contextual identification and reconstruction of cultural references. Methodologically, it integrates literary data analysis, contextual resources, and expert feedback mechanisms. The framework is validated using Danish literature case studies to evaluate models’ cultural robustness and interpretive depth while preserving legitimate scholarly disagreement. Ultimately, this research establishes an evaluation paradigm that transcends conventional metrics, advancing AI development toward systems capable of sophisticated cultural reasoning.
This study addresses the challenge of reliably capturing narrative themes and their dynamic evolution in small-scale poetic corpora, where traditional topic models often yield unstable results. To overcome this limitation, the authors propose a lightweight hybrid framework that integrates unsupervised Latent Dirichlet Allocation (LDA) with supervised sparse Partial Least Squares Discriminant Analysis (sPLS-DA). By incorporating multi-seed consensus strategies and narrative hub analysis—and deliberately filtering out prosodic and other surface-level linguistic features—the approach enables a computationally rigorous close reading of *Eugene Onegin*. The method substantially enhances topic stability and literary interpretability within limited corpora, successfully identifying five coherent themes that align meaningfully with the poem’s emotional trajectory and narrative arc. This work thus establishes a transparent, reproducible paradigm for computational analysis of densely layered literary texts.
研究通过半结构化访谈探索读者对AI生成脚注的需求和偏好,旨在解决静态脚注无法满足所有读者问题的情况。
This work addresses the challenge of extracting scientific hypotheses and their supporting statistical evidence from research papers, a task hindered by the documents’ length and dispersed information. The authors propose a two-stage retrieve-and-extract framework that incorporates a paper-structure-aware context selection strategy to effectively link key findings in abstracts with corresponding hypotheses and evidence in the main text. Through systematic evaluation of various configurations—including standard RAG, re-ranking, fine-tuned retrievers, and large language model extractors—and by decoupling retrieval from extraction performance using oracle passages, the study demonstrates that high-quality contextual passages substantially improve hypothesis extraction. However, extracting statistical evidence remains challenging, revealing limitations in current models when processing hybrid numerical-textual statements.
This work addresses the challenge researchers often face in balancing novelty with effective grounding in existing literature when developing new ideas, as well as the lack of tools that support dynamic interaction between emerging concepts and relevant scholarly works. The paper introduces a novel “literature-driven idea pivoting” mechanism—a closed-loop framework that integrates idea drafting, dynamic literature retrieval, semantic clustering, and generative critical feedback to enable co-evolution of research ideas and the literature space. The system performs context-aware analysis of partial idea content and provides real-time improvement suggestions based on clusters of relevant papers. Experimental results demonstrate that this approach significantly enhances the quality of user-generated ideas and strengthens researchers’ ability to comprehend and leverage the scholarly context effectively.