natural language processing

Designs, implements, and evaluates algorithms and systems that process human natural language in text or speech to analyze, interpret, generate, or transform linguistic content. Work includes building and testing tokenization and parsing pipelines, semantic and discourse representations, language models, information extraction, summarization, translation, question-answering and dialogue components, and speech recognition or synthesis modules, along with their evaluation metrics and error analyses.

naturallanguageprocessing

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-3.25
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$217K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Text Production and Comprehension by Human and Artificial Intelligence: Interdisciplinary Workshop Report

Jun 27, 2025
ED
Emily Dux Speltz
🏛️ Embry-Riddle Aeronautical University

This study addresses the fundamental scientific question of similarities and differences between large language models (LLMs) and human language cognition. Adopting an interdisciplinary approach integrating cognitive psychology, linguistics, and NLP, we design controlled human–machine behavioral comparison experiments to systematically evaluate how well LLMs simulate human linguistic behavior in text understanding and generation tasks. Our key contributions are threefold: (1) We establish LLMs as novel computational tools for investigating human language cognition; (2) We demonstrate that human feedback–based fine-tuning significantly enhances behavioral fidelity; and (3) We propose a new “human–AI co-enhanced language capability” paradigm. The work delineates the potentials and limitations of LLMs in cognitive modeling, providing both theoretical foundations and practical pathways for their trustworthy deployment in psychological assessment, language acquisition research, and educational interventions.

Assessing ethical implications of human-AI collaboration in linguisticsExploring AI's role in augmenting human language capabilitiesUnderstanding AI-human cognitive links in text processing

This study investigates systematic linguistic disparities between AI-generated and human-written texts across multiple hierarchical levels. Method: Using human-authored argumentative essays as a benchmark, we compare length-matched ChatGPT outputs and automatically extract phonological (e.g., consonant types), morphological, syntactic (e.g., adjective/prepositional modifiers), and lexical features (e.g., nouns, adjectives, pronouns, low-frequency words) via the Open Brain AI platform, followed by comparative statistical analysis. Contribution/Results: We introduce the first automated, multi-level linguistic assessment paradigm grounded in publicly accessible computational tools. Results reveal significant deviations in AI text—particularly in part-of-speech distributions, modifier structural complexity, and usage of low-frequency vocabulary—relative to human norms. The paradigm demonstrates high efficacy for authorship attribution and provides reproducible, scalable empirical foundations for both AI-text detection and generative model refinement.

Assessing AI's ability to emulate human writing across multiple featuresDifferentiating human-written and AI-generated texts using linguistic featuresQuantifying linguistic differences in phonological, morphological, syntactic components

Natural Language Processing RELIES on Linguistics

May 09, 2024
JO
Juri Opitz
🏛️ University of Zurich | Georgetown University

Recent large language models (LLMs) exhibit a superficial “de-linguistification” trend, marginalizing linguistics despite its foundational relevance to natural language processing (NLP). Method: This paper systematically reasserts linguistics’ indispensable structural role in NLP through an original six-dimensional RELIES framework—encompassing Resources, Evaluation, Low-resource settings, Interpretability, Explanation, and Study of language—and integrates linguistic insights via conceptual analysis, interdisciplinary synthesis, and empirical case studies. Contribution/Results: The work challenges the purely data-driven paradigm by demonstrating how linguistic theory methodologically anchors model architecture design, evaluation criteria, and ethical governance. It establishes linguistics not as auxiliary but as constitutive to NLP’s scientific rigor and human-centered grounding, offering a systematic, theory-informed roadmap for developing linguistically principled, interpretable, and equitable NLP systems.

Examines NLP's reliance on linguistics for grammar and semantics.Explores linguistic contributions to NLP in low-resource settings.Highlights linguistics' role in NLP interpretability and explanation.

Effect-driven interpretation: Functors for natural language composition

Apr 01, 2025
DB
Dylan Bumford
🏛️ University of California, Los Angeles | Yale University

This paper addresses the challenge of jointly modeling pure semantic values and contextual effects—such as reference, tense, interrogation, and negation—in compositional natural language semantics. We propose an effect-driven semantic framework inspired by denotational semantics in programming languages. Methodologically, we systematically introduce functorial structures from category theory to formally characterize the hierarchical interaction between value propagation and side-effectful processes; we integrate type-logical syntax with functional semantic composition to build an extensible interpretation system. Our contributions include a unified treatment of tense, questions, negation, and discourse coherence, significantly enhancing compositional productivity, interpretability, and theoretical unity in semantic parsing. The framework establishes a new paradigm for computational semantics that balances formal rigor with broad linguistic coverage.

Applying denotational techniques to natural language analysisExploring functors for linguistic composition interpretationModeling human language with pure and impure components

This study addresses critical challenges in multilingual NLP—model bias, insufficient robustness, and difficulty in ethical alignment—by proposing a fine-tuning and deployment framework for large language models (LLMs) targeting low bias and high robustness. Methodologically, it integrates the Hugging Face ecosystem with Transformer architectures, incorporating multilingual tokenization, domain-aware data cleaning and augmentation, and a progressive fine-tuning strategy that jointly optimizes fairness and task performance. Contributions include: (1) a lightweight, cross-lingual fine-tuning paradigm resilient to bias-induced interference; (2) empirical validation across high-stakes domains (e.g., healthcare and finance), demonstrating significant improvements in generalization and fairness for classification and named entity recognition; and (3) an interpretable, auditable, and production-ready LLM deployment pipeline that advances the practical implementation of ethically aligned AI.

Addressing data preprocessing and transformer model implementationExploring NLP and LLMs in machine learning intersectionSolving multilingual data handling and AI bias reduction

Latest Papers

What's happening recently
View more

This study examines whether large language model (LLM)-driven dialogue systems have genuinely advanced our understanding of human linguistic competence. By systematically tracing the evolution of natural language processing—from early rule-based systems to contemporary LLMs—and juxtaposing this trajectory with theoretical frameworks from linguistics and cognitive science concerning the mental mechanisms underlying human language, the paper reveals that despite remarkable progress in generative capabilities, current technologies have not substantively deepened our grasp of the nature of human language. The work’s key contribution lies in establishing an integrative, interdisciplinary analytical framework that explicitly identifies a profound disconnect between artificial language modeling and human language comprehension, thereby charting a path for future research that balances technical performance with scientific explanatory power.

Cognitive ScienceConversational AgentsHuman Language Capacity

This study examines the applicability of NLP to qualitative social science text analysis, focusing on strategic signaling themes in U.S. Presidential Directives (PDs). Adopting a hybrid paradigm that integrates expert annotation with multiple NLP approaches—including LDA, BERTopic, and supervised classification—the study systematically compares human and algorithmic performance across thematic consistency, semantic sensitivity, and interpretive validity. Results show that NLP methods efficiently detect high-frequency strategic themes (e.g., “ally coordination,” “deterrence escalation”) but exhibit significant limitations in capturing implicit intent, context-dependent rhetoric, and institutionally constrained formulations. The work introduces the first domain-specific annotation framework for strategic signaling in political discourse and proposes the “Social Science Readiness” metric—a multidimensional assessment framework evaluating AI tools’ suitability, reliability, and human–AI collaboration pathways in qualitative social research. This provides both theoretical grounding and empirical benchmarks for integrating NLP into rigorous, interpretive social science inquiry.

Assessing discrepancies between NLP results and human analysisEvaluating NLP's ability to extract topics from large text corporaIdentifying strategic signaling patterns in presidential directives

Automating the Analysis of Parsing Algorithms (and other Dynamic Programs)

Dec 29, 2025
TV
Tim Vieira
🏛️ Johns Hopkins University | ETH Zürich

This paper addresses the challenge of establishing performance guarantees for dynamic programming (DP) parsing algorithms in natural language processing. We present the first automated analysis system that unifies program analysis and complexity inference within a DP framework. Our approach integrates static analysis, type inference, abstract interpretation, and dependency graph modeling to enable formal verification and synthesis of efficient data structures. Key contributions include: (1) a unified formal model capturing DP control flow, data flow, and recurrence structure; (2) automatic inference of precise types, detection of dead code, and identification of redundant computations; and (3) generation of tight, parameterized upper bounds on time and space complexity. We evaluate our system on canonical parsing algorithms—including CKY, Earley, and Neural PCFG—demonstrating substantial improvements in both the automation level and precision of complexity analysis.

Automating analysis of parsing algorithms and dynamic programsInferring types, dead code, and verifying algorithm propertiesProviding guarantees on runtime and space complexity bounds

This work addresses the prevailing limitation in large language model (LLM) development, wherein human values are typically incorporated only post-training, lacking systematic integration across the model’s entire lifecycle. To bridge this gap, the paper introduces the Human-Centric Large Language Model (HCLLM) framework, which for the first time deeply integrates natural language processing, human-computer interaction, and responsible AI methodologies throughout all stages—from system design and data collection to training, evaluation, and deployment. The framework harmonizes ethical, economic, and technical objectives, offering developers actionable, principle-based guidance. Its forward-looking applicability and practical utility are demonstrated through a case study situated in future workplace scenarios, thereby advancing LLM development toward a genuinely human-centered paradigm.

Ethical AIHuman-Centered AIHuman-Computer Interaction

This study addresses the limitations of prevailing culturally sensitive approaches in natural language processing (NLP), which often remain confined to surface-level representations and overlook deeper sociocultural contexts and power structures. The paper proposes a sociotechnical framework grounded in pluralistic epistemologies that moves beyond monolithic cultural adaptation paradigms by integrating local knowledge systems into NLP design. Through a five-layer model of technical activity, the authors systematically examine how culture is operationalized in NLP systems and expose critical gaps in current methods concerning governance, power dynamics, and sociocultural embeddedness. This approach reframes culture from a static object of representation to a dynamic, actionable construct, thereby advancing more reflexive and inclusive forms of cultural alignment in NLP.

cultureepistemologiesnatural language processing

Hot Scholars

AG

Agam Goyal

CS PhD Student, University of Illinois Urbana-Champaign
Natural Language ProcessingHuman-AI InteractionSocial ComputingComputational Social Science
EC

Eshwar Chandrasekharan

Assistant Professor, University of Illinois Urbana-Champaign
Social ComputingHCIHuman-Centered AIOnline Governance
KS

Koustuv Saha

University of Illinois Urbana-Champaign
Computational Social ScienceSocial ComputingHuman-Centered Machine LearningWellbeing
CZ

Chengzhi Zhang

Nanjing University of Science and Technology
Text MiningNatural Language ProcessingScience of Science