Score
Designs, implements, and evaluates methods for interpreting and representing the meaning and structure of content—such as text, images, audio, or video—by extracting entities, relations, topics, intent, sentiment, summaries, or semantic segments. Produces representations, annotations, or predictions that enable search, classification, summarization, question answering, or other downstream analyses.
Multimodal document structural understanding lacks formal theoretical foundations. Method: This paper proposes a novel categorical modeling paradigm grounded in category theory: documents are formalized as categories whose objects are document elements and morphisms are question-answer pairs; information orthogonal decomposition and composability constraints are introduced to enable mathematical representation and rate-distortion analysis of document content. Building upon this, we develop a measurable, enumerable information-theoretic evaluation framework—enabling unsupervised text exegesis expansion and consistency-aware abstractive summarization—and optimize large language models via RLVR (Reconstruction-Labeling-Verification-Refinement) self-supervision. Contribution/Results: This work is the first to systematically integrate category theory into document structure modeling; it establishes the principle of information orthogonality and unifies structural, semantic, and generative aspects under a quantitative framework. Empirical results demonstrate significant improvements in summary quality and generation consistency, enabling theory-driven, self-supervised enhancement of multimodal foundation models.
This work addresses the challenge of named entity extraction from news videos, where diverse on-screen text layouts and the infeasibility of manual annotation hinder reliable performance. To this end, the authors introduce the first balanced, graphics-focused annotated dataset for news videos and propose an interpretable, modular, deterministic multimodal pipeline. The framework synergistically combines rule-based methods with deep learning, featuring a high-precision graphic detector (mAP@0.5 of 95.8%) and a hallucination-free named entity recognition module. Operating under strict auditability and zero-hallucination constraints, the system achieves 79.9% precision and 74.4% recall. A user study further reveals that 59% of viewers struggle to identify on-screen names during fast-paced broadcasts, underscoring the practical relevance of the proposed approach.
This study investigates multimodal large language models’ (MLLMs) capacity to comprehend abstract psychological concepts—such as self-disclosure—in YouTube短视频 depicting depression. Method: Leveraging 725 annotated keyframes, we conduct a systematic human–AI semantic interpretation comparison, examining how conceptual operationalization granularity, semantic complexity, and video modality diversity affect human–AI semantic alignment. Using LLaVA-1.6-Mistral-7B, we integrate qualitative interpretability analysis with cross-modal concept alignment evaluation. Contribution/Results: We propose a tailored prompting strategy for abstract psychological concepts and a human-centered multimodal evaluation paradigm. Contrary to intuition, excessive operational granularity degrades alignment; we identify critical dimensions governing consistency. Our work delivers a reproducible methodological framework and practical guidelines for AI-driven video content analysis in computational social science.
This paper addresses the insufficient modeling of textual implicit semantics by proposing a novel paradigm that explicitly transforms “subtext” into a verifiable set of propositions. Methodologically, it systematically leverages large language models to generate implicit inference propositions from text, which are then validated for plausibility by human annotators; the resulting proposition semantics are subsequently fused into the original text representations. Key contributions include: (1) introducing the first annotated framework for implicit propositions tailored to social science tasks; and (2) demonstrating that this explicit modeling significantly outperforms literal-only representations across three distinct tasks—argument similarity assessment, public opinion interpretation, and legislative behavior simulation—with average improvements of 12.7% in F1 or accuracy. Results substantiate the effectiveness and generalizability of structured implicit semantic modeling for enhancing human-like semantic understanding.
Deep learning models achieve state-of-the-art performance in NLP and information retrieval, yet their opacity severely hinders trustworthy deployment. This paper presents the first systematic, cross-model (word embeddings, RNNs/LSTMs, Transformers, BERT) and cross-task (text classification, question answering, document ranking) survey of interpretability methods in NLP/IR. We propose a structured taxonomy covering major paradigms—including feature attribution (e.g., LIME, SHAP), attention analysis, surrogate modeling, saliency mapping, and counterfactual explanation. Our framework constitutes the most comprehensive synthesis of textual interpretability techniques to date. We rigorously identify critical limitations—particularly the lack of standardized evaluation protocols and insufficient task-specific adaptation—and highlight key research gaps. The work establishes both theoretical foundations and practical guidelines for developing interpretable, reliable NLP systems.
This work addresses the limitations of existing video analysis methods, which rely on coarse-grained AI summaries and struggle to capture the structural evolution and semantic relationships in long-form videos. To overcome this, the authors propose a multi-level large language model (LLM) framework that first performs global semantic modeling over full video transcripts, then conducts context-aware sentence-by-sentence parsing. Innovatively, the framework employs an LLM as a discriminator to cluster sentences based on semantic similarity, enabling fine-grained semantic segmentation. By integrating both global and local contextual information, the approach supports the generation of interpretable visualizations—including semantic structure graphs and relevance heatmaps—thereby significantly enhancing the depth and explainability of video understanding. This method is particularly well-suited for applications such as educational content analysis and lecture replay, where structured semantic insight is essential.
This study addresses the limited interactivity and domain adaptability of existing clustering methods for digital humanities scholars working with large-scale unstructured documents. To bridge this gap, the authors propose an analysis-perspective-driven interactive document clustering framework. This framework enables users to define initial semantic lenses through prompt rewriting and instruction embedding, and integrates interactive visualization, on-the-fly cluster adjustment, and online fine-tuning of embedding models into a closed-loop human-in-the-loop feedback process. The approach supports an interpretable, intervenable, and iterative clustering experience, empowering researchers to efficiently uncover latent semantic structures—such as thematic patterns or sentiment signals—and thereby generate high-quality structured data to support in-depth humanities inquiry.
This study addresses the challenges of losing diverse perspectives and lacking traceability in the analysis of large-scale heterogeneous textual corpora. To this end, it proposes a structured reading approach grounded in large language models, which defers irreversible information compression by sequentially performing insight extraction, semantic clustering, theme generation, and iterative omission detection. This pipeline explicitly preserves divergent viewpoints, thereby enhancing both coverage and auditability of the analytical process. Evaluated on a corpus of 152 industrial policy documents, the method successfully extracted over 17,500 structured insights and constructed a comprehensive thematic map. The implementation has been open-sourced as the first end-to-end framework supporting large-scale qualitative synthesis.
This study addresses the challenging task of automatically identifying and segmenting legal conditions (Tatbestand) from legal consequences (Rechtsfolge) in German statutory texts. To facilitate research on this structural parsing problem, the authors introduce ANNOTARES, the first fine-grained annotated dataset covering three major German legal codes, enabling cross-code generalization studies. The work systematically evaluates a range of approaches, including rule-based baselines, CRF, BiLSTM, BiLSTM-CRF, and Transformer architectures based on BERT and large language models. Experimental results demonstrate that BERT-based and large language models significantly outperform traditional methods in capturing the complex syntactic structures inherent in legal texts, thereby confirming the effectiveness of pretrained language models for extracting logical structures in legal documents.
This work addresses the challenge posed by the lack of structured semantic representations in legal case records, which hinders the performance of downstream legal AI tasks. To overcome this limitation, the authors propose LeDA, a web-based annotation platform that supports dynamic label creation without requiring a predefined ontology. LeDA enables annotators to iteratively discover and define legal concepts during the annotation process, while incorporating collaborative multi-user annotation and an arbitration mechanism to resolve disagreements. The system was successfully deployed on judgments from the Supreme Court of India, where three annotators constructed a “bag-of-concepts” semantic representation. This representation effectively facilitates precedent retrieval and judgment prediction, significantly enhancing the structured understanding and semantic processing of legal texts.