conduct corpus analysis

Designs and executes analyses of text corpora to quantify lexical or linguistic feature prevalence and their temporal dynamics, including assembling longitudinal or large-scale textual datasets and computing frequency and time-series measures. Builds tests and visualizations to detect accelerations, shifts, or diachronic patterns, compare prevalence across groups (e.g., disciplines, regions, author origin), and correlate textual trends with external events or covariates.

conductcorpusanalysis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.12
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$112K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

ttta: Tools for Temporal Text Analysis

Mar 04, 2025
KL
Kai-Robin Lange
🏛️ TU Dortmund University | RWI - Leibniz Institute for Economic Research | University of Zurich

Contemporary NLP methods often treat text as temporally homogeneous, neglecting semantic evolution over time and thereby introducing temporal semantic bias; moreover, existing tools for temporal text analysis are fragmented and lack reproducibility. To address these limitations, we propose the first unified framework that systematically integrates multi-granular temporal text analysis, enabling both topic evolution modeling and lexical semantic shift detection. Implemented in Python, the framework incorporates dynamic topic modeling, temporal word embedding alignment, sliding-window LDA, and interactive visualization modules. It ensures end-to-end reproducibility and achieves significant improvements in topic trend identification accuracy on cross-year news and social media datasets. Our core contribution is the first standardized, open-source, and extensible toolkit for temporal NLP—bridging the gap between theoretical models of language evolution and empirical, large-scale diachronic analysis.

Addressing bias from time-homogeneous NLP techniques.Analyzing temporal changes in text data meaning.Providing unified tools for temporal text analysis.

This study proposes a statistical framework to identify emerging narratives in longitudinal textual corpora that reflect structural shifts in discourse, rather than superficial fluctuations in language use. By operationalizing “narrative emergence” as a sustained increase in the relative salience of latent topics, the method integrates topic modeling—based on Latent Dirichlet Allocation (LDA)—with time-series analysis and formal statistical inference. The approach is validated against external observable signals, such as the timing of Nobel Memorial Prizes in Economic Sciences. Applied to economics literature from 1970 to 2018, the framework successfully detects topics closely aligned with Nobel-recognized contributions, exhibiting significant upward trajectories that coincide with surges in citations and broader disciplinary recognition. This provides a statistically testable foundation for tracing structural semantic change in scholarly discourse over time.

economic discourseemergent narrativeslongitudinal text corpora

This work proposes a context-aware bias detection framework that identifies subtle linguistic biases in large language model outputs toward diverse social groups without relying on predefined lists of sensitive terms. The approach generates structured synthetic minimal-pair texts—narratively consistent except for the substitution of target group markers—and employs linguistic form abstraction combined with an enhanced variant of pointwise mutual information (PMI) for comparative analysis. Integrating quantitative statistics with qualitative evaluation, the framework is adaptable across multiple text genres and effectively quantifies asymmetric associations between social groups and levels of linguistic abstraction. It precisely localizes textual segments with high concentrations of bias signals, enabling domain experts to identify potentially harmful expressions within their contextual settings.

contextualized representationscontrastive analysislarge language models

This study addresses key challenges faced by social science researchers when using large language models (LLMs) for text annotation—namely, poor reproducibility, annotation errors that compromise statistical inference, and high technical barriers. To overcome these issues, the authors propose the first end-to-end LLM-based text annotation framework tailored specifically for the social sciences and humanities (SSH). The framework integrates structured prompt engineering, open-source LLM API integration, cross-validation, and error propagation modeling, with an explicit emphasis on avoiding prompt overfitting and quantifying annotation uncertainty. Implemented in both Python and R, this approach establishes a transparent, reproducible, and scalable workflow that substantially enhances the reliability, efficiency, and methodological rigor of automated text annotation in SSH research.

annotation errorlarge language modelsreproducibility

Latest Papers

What's happening recently
View more

This study investigates how dataset characteristics constrain the effectiveness of quantitative methods in detecting semantic change within historical linguistics. By comparing two diachronic approaches—quadripartite conceptual modeling applied to the EEBO-TCP corpus and SynFlow analysis on the Royal Society Corpus—the research examines differences in conceptual operationalization, underlying data assumptions, and diachronic interpretability. The findings highlight the limitations of purely lexical frequency-based methods and demonstrate that data structure, including temporal granularity and textual representativeness, fundamentally determines which types of semantic shifts can be reliably identified. This work thus offers methodological guidance for historical semantic research and advocates for the development of quantitative paradigms better aligned with the specific properties of historical linguistic data.

conceptual changedataset propertieshistorical linguistics

This study addresses the scarcity of annotated historical Italian news corpora by proposing an unsupervised method to automatically identify major socio-political turning points. Building on a diachronic corpus of approximately 600,000 articles from *La Repubblica* (1985–2000), the approach integrates natural language processing, word embeddings, semantic change modeling, complex network analysis, and tools from statistical physics. For the first time, complex systems theory is applied to diachronic media analysis in the Italian context. Without relying on any prior labels, the method successfully detects abrupt shifts in media discourse corresponding to pivotal events such as the transition from Italy’s First to Second Republic, the Gulf War, and the Kosovo War. This work offers a novel paradigm for digital humanities and computational social science by demonstrating how unlabeled textual data can reveal historically significant societal transformations.

complex systemsdiachronic corpushistorical turning points

This study investigates whether the widespread adoption of large language models (LLMs) has led undergraduate students to overrely on generative AI in statistical writing, thereby affecting their statistical reasoning and communicative competence. Analyzing over 1,600 student data analysis reports from 2021 to 2025 through text similarity metrics, stylometric analysis, and large-scale corpus comparisons, this work provides the first empirical evidence of LLMs’ differential impact across report sections—most pronounced in introductions and conclusions. Findings indicate that student writing increasingly converges toward LLM-generated styles, particularly in these segments, while simultaneously aligning more closely with expert statistical discourse. This dual trend suggests that LLMs entail both cognitive risks and pedagogical potential. The study further proposes a novel assessment framework integrating statistical thinking to better evaluate AI-mediated learning outcomes.

generative AIlarge language modelsstatistical communication

This study addresses the current lack of interdisciplinary understanding regarding the integration pathways, efficacy boundaries, and systemic risks of large language models (LLMs) across natural sciences, social sciences, and humanities. Through a systematic literature review and illustrative case analyses, it critically evaluates the deployment of LLMs throughout the research lifecycle—including hypothesis generation, literature synthesis, data analysis, and scholarly writing. The work identifies ten previously underappreciated systemic risks, such as diminished researcher autonomy, AI-induced confirmation bias, ambiguous authorship, and inequitable access to technology. It further demonstrates how LLMs, while enhancing efficiency, simultaneously introduce challenges like hallucination, irreproducibility, data bias, and model opacity. To guide responsible adoption, the study proposes an interdisciplinary governance framework and a roadmap for explainable AI research in scholarly contexts.

AI ethicsinterdisciplinary integrationLarge Language Models

Hot Scholars

KM

Kyle Mahowald

UT Austin
computational linguisticspsycholinguisticsnatural language processingcognitive science
BP

Barbara Plank

Professor, LMU Munich, Visiting Prof ITU Copenhagen
Natural Language ProcessingComputational LinguisticsMachine LearningTransfer Learning
KR

Karolina Rudnicka

University of Gdańsk (Poland)
variation and changecorpus linguisticsapplied linguisticsEnglish language
CZ

Chengzhi Zhang

Nanjing University of Science and Technology
Text MiningNatural Language ProcessingScience of Science
CB

Claire Bonial

Computational Linguist, ARL
Natural Language Processing