Score
Designs and implements analysis systems that measure and interpret semantic content of textual data at multiple granularities—documents, segments, topics, and words—using terminology analysis, topic/segment modeling, and multilevel textual representations. Builds flexible, generalizable measurement pipelines that integrate embeddings, LLM outputs, and segment- or topic-based analyses to produce semantic metrics, labels, or summaries for corpora.
Taxonomy construction for classifying unstructured text (e.g., personal goal statements) is time-intensive, prone to researcher bias, and suffers from poor reproducibility. Method: This paper proposes a human–AI collaborative, iterative text analysis paradigm that integrates top-down and bottom-up strategies to enable dynamic taxonomy generation, evaluation, refinement, and validation. Leveraging prompt engineering, it facilitates multi-turn collaboration between domain researchers and large language models (LLMs), with human feedback driving iterative taxonomy optimization. Intercoder reliability is quantified using Cohen’s κ within a structured coding framework. Results: Empirical evaluation in a life-domain dataset achieves κ > 0.85, significantly enhancing analytical efficiency, reliability, and reproducibility. This work pioneers deep integration of LLMs into the qualitative analysis closed loop, offering a novel methodology for low-bias, high-fidelity open-text classification.
Traditional unsupervised text analysis suffers from poor interpretability and weak semantic coherence in data-scarce domains. To address this, we propose Recursive Topic Partitioning (RTP), the first framework that deeply integrates problem-driven binary semantic tree construction with large language models (LLMs), explicitly encoding clustering logic and transforming analytical pathways into structured prompts—thereby enabling interpretable clustering and controllable generation in a closed loop. RTP synergizes recursive semantic segmentation, contextual reasoning, and prompt engineering to jointly support topic discovery and data synthesis. Experiments demonstrate that RTP-generated semantic trees significantly outperform keyword-based methods (e.g., BERTopic) in structural quality and coherence; it achieves substantial gains in few-shot classification tasks; and it enables precise, semantics-guided controllable text generation aligned with user-specified semantic features.
This study addresses the challenge of applying unstructured text data directly to psychometric analysis. We propose a novel paradigm wherein documents are treated as respondents and words as test items; context-aware word embeddings are extracted using encoder-only Transformer models to construct contextualized response data. Subsequently, multivariate factor analytic techniques—including exploratory factor analysis and bifactor modeling—are employed to uncover latent knowledge dimensions and structural patterns. Our key contribution lies in the first systematic integration of large language model–generated contextual embeddings into a psychometric framework, enabling naturalistic, interpretable measurement of semantic variation in text. Experiments on the Wiki STEM corpus successfully identify coherent, interpretable knowledge dimensions, demonstrating the method’s validity and generalizability for text analysis in education, psychology, and legal domains.
This paper addresses critical challenges in the deep integration of large language models (LLMs) with visual analytics—namely, ambiguous task boundaries, fragmented technical approaches, and absent ethical safeguards. Methodologically, it establishes the first LLM-empowered visual analytics technology taxonomy and SWOT framework, proposing a classification system spanning 12 core tasks. It empirically benchmarks leading platforms (e.g., LIDA, Chat2VIS) and multimodal models (e.g., ChartLlama, CharXIV), while innovatively unifying natural language understanding, text-to-chart generation, and human-AI collaborative interaction—augmented by dual-dimensional constraints: ethical principles and methodological rigor. The study identifies three fundamental bottlenecks: computational overhead, representational bias, and data privacy risks. Its contributions include a theoretically grounded, practice-oriented foundation for developing trustworthy, interpretable, and human-centered visualization AI systems.
With the exponential growth of scientific literature, automated extraction of key concepts remains challenging, particularly due to poor cross-disciplinary adaptability. Method: This paper proposes a lightweight LLM-based semantic extraction method supporting FAIR implementation in scholarly workflows. It introduces a context learning–driven zero-/few-shot domain adaptation mechanism that enables rapid, fine-tuning–free adaptation to new disciplines. We systematically benchmark multiple open-source and commercial LLMs on concept identification tasks and develop an interactive online prototype system. Contribution/Results: Empirical evaluation in computer science—complemented by user studies—demonstrates the method’s effectiveness in structured literature review, knowledge graph construction, and information retrieval. It significantly improves both accuracy and cross-domain generalization of concept extraction, offering a scalable technical pathway for intelligent, full-lifecycle scholarly knowledge services.
Traditional topic models assign a single topic to an entire document, which struggles to accurately represent multi-topic texts and often leads to topic mixing and reduced interpretability. This work proposes a Segment-Based Topic Assignment (SBTA) framework that, for the first time, refines the granularity of topic modeling from the document level to semantically coherent text segments. To support this approach, we construct the SemEval-STM dataset by combining large language model–based automatic segmentation with human refinement to produce high-quality segments, and introduce a segment-level word intrusion task to enable fine-grained evaluation. Experiments demonstrate that SBTA significantly improves topic clustering quality and interpretability across multiple topic models and evaluation metrics, confirming its effectiveness and scalability.
This work addresses the challenge of achieving fine-grained, high-precision topic modeling in narrow domains—particularly distinguishing semantically similar subtopics—under constraints of low cost and interpretability. The authors propose PRISM, a framework that fine-tunes a lightweight sentence encoder using sparse labels generated by a large language model (LLM) and applies threshold-based clustering to partition the embedding space into locally interpretable topic structures. PRISM employs an LLM-guided teacher–student distillation pipeline, requiring only minimal LLM queries and sparse supervision, while also investigating how active sampling strategies influence local embedding geometry. Experiments demonstrate that PRISM significantly outperforms existing local topic models across multiple corpora and even surpasses state-of-the-art large embedding models in clustering quality, offering a compelling balance of efficiency, accuracy, and deployability.
This study addresses the current lack of interdisciplinary understanding regarding the integration pathways, efficacy boundaries, and systemic risks of large language models (LLMs) across natural sciences, social sciences, and humanities. Through a systematic literature review and illustrative case analyses, it critically evaluates the deployment of LLMs throughout the research lifecycle—including hypothesis generation, literature synthesis, data analysis, and scholarly writing. The work identifies ten previously underappreciated systemic risks, such as diminished researcher autonomy, AI-induced confirmation bias, ambiguous authorship, and inequitable access to technology. It further demonstrates how LLMs, while enhancing efficiency, simultaneously introduce challenges like hallucination, irreproducibility, data bias, and model opacity. To guide responsible adoption, the study proposes an interdisciplinary governance framework and a roadmap for explainable AI research in scholarly contexts.
This work addresses the limitations of large language models in text classification, where stochastic attention mechanisms and sensitivity to noise often compromise accuracy and reproducibility. To mitigate these issues, the authors propose the wSSAS framework, which leverages signal-to-noise ratio (SNR) to identify high-value semantic features and organizes texts into a hierarchical “topic–narrative–cluster” structure. The framework incorporates a deterministic mechanism that jointly evaluates weighted syntactic and semantic contextual cues and employs a Summary-of-Summaries architecture to aggregate salient information. Empirical evaluations on multi-domain review datasets from Google, Amazon, and Goodreads demonstrate that wSSAS significantly reduces classification entropy, enhances clustering completeness, and improves classification accuracy, thereby validating its effectiveness in bolstering result stability and robustness against noise.
This work addresses the limitation of existing document analysis systems that flatten complex documents into plain text, thereby discarding critical hierarchical structures such as sections, tables, and figures, which hinders effective filtering and in-depth analysis. To overcome this, the authors propose a structure-aware document understanding framework that integrates full document hierarchy into semantic indexing and analysis. The approach constructs a hierarchical document tree by parsing the original layout, leverages large language models to generate structure-aware semantic representations, and introduces a multi-view interactive web interface enabling precise natural language–driven retrieval and question answering. Experiments demonstrate significant improvements in retrieval accuracy and question-answering performance on diverse complex documents, including academic papers, technical manuals, and financial reports. The code and a live demo system are publicly released.