syntactic parsing

Designs, builds, and evaluates systems that convert natural language (or other structured inputs) into syntactic structures—such as constituency trees or dependency graphs—using rule-based, statistical, or neural methods, and performs syntactic analysis to measure accuracy and error modes. Extends this work to produce and integrate meaning-bearing outputs (semantic parsing, logical forms, or scene-level structured representations) and to connect syntactic outputs into downstream NLP pipelines.

syntacticparsing

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.51
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

A Fundamental Algorithm for Dependency Parsing (With Corrections)

Oct 22, 2025
MA
Michael A. Covington
🏛️ The University of Georgia

This paper addresses online dependency parsing by proposing a word-by-word, real-time parsing algorithm aligned with human linguistic cognition. The method dynamically constructs a dependency tree incrementally, assigning each input token its head immediately upon arrival—without backtracking or global reanalysis. Built upon a dynamic programming framework, the algorithm has a theoretical worst-case time complexity of O(n³); however, empirical evaluation shows this bound is attained only for extremely short sentences (n ≤ 15), while for typical sentence lengths (n > 20), practical runtime scales nearly linearly. Its primary contribution lies in unifying cognitive plausibility—specifically, incremental attachment and zero backtracking—with provably polynomial time complexity within a single algorithmic framework. This design substantially improves both accuracy and robustness over conventional greedy online parsers. Moreover, it furnishes a linguistically interpretable and formally verifiable syntactic foundation for neuro-symbolic language processing models.

Achieves cubic worst-case complexity optimized for human languageDevelops dependency parsing algorithm for natural language sentencesOperates incrementally by attaching words individually during parsing

A Systematic Comparison of Syntactic Representations of Dependency Parsing

May 29, 2017
GW
Guillaume Wisniewski
🏛️ Univ. Paris-Sud | Université Paris-Saclay | University of Copenhagen

This study systematically investigates how dependency annotation schemes affect the performance of transition-based parsers. Method: Addressing language-specific non-canonical structures in Universal Dependencies (UD) treebanks, we design standardization transformation rules and comparatively evaluate parser performance—measured by LAS and UAS—under both original and standardized annotations within a unified, multilingual evaluation framework. Contribution/Results: We empirically demonstrate, for the first time, that annotation standardization does not universally improve parsing accuracy. Crucially, we reveal that linguistic typological features significantly moderate the effectiveness of annotation schemes: for certain languages, the original non-standard annotations yield higher accuracy than standardized ones. This finding challenges the implicit assumption that standardization is inherently optimal and underscores the necessity of considering language-specific syntactic properties when selecting or designing syntactic representations.

Compare parser performance across annotation schemes.Convert syntactic constructions to standard representations.Evaluate parsing performance across multiple languages.

Effect-driven interpretation: Functors for natural language composition

Apr 01, 2025
DB
Dylan Bumford
🏛️ University of California, Los Angeles | Yale University

This paper addresses the challenge of jointly modeling pure semantic values and contextual effects—such as reference, tense, interrogation, and negation—in compositional natural language semantics. We propose an effect-driven semantic framework inspired by denotational semantics in programming languages. Methodologically, we systematically introduce functorial structures from category theory to formally characterize the hierarchical interaction between value propagation and side-effectful processes; we integrate type-logical syntax with functional semantic composition to build an extensible interpretation system. Our contributions include a unified treatment of tense, questions, negation, and discourse coherence, significantly enhancing compositional productivity, interpretability, and theoretical unity in semantic parsing. The framework establishes a new paradigm for computational semantics that balances formal rigor with broad linguistic coverage.

Applying denotational techniques to natural language analysisExploring functors for linguistic composition interpretationModeling human language with pure and impure components

This study systematically investigates whether and how Transformer-based language models acquire syntactic knowledge. Through a large-scale, systematic literature review synthesizing findings from 337 studies and over 3,000 data points, the work presents the first integrated quantitative assessment of syntactic capabilities across multiple languages and model architectures by combining behavioral experiments, representation probing, and mechanistic interpretability methods. The analysis reveals that Transformers possess substantial syntactic knowledge, yet exhibit limitations in phenomena at the syntax–semantics interface and in low-resource languages. It also highlights a pronounced research bias toward English and BERT-family models, with insufficient coverage of linguistic and architectural diversity. This work provides comprehensive empirical evidence and new directions for understanding the mechanisms and boundaries of syntactic generalization in neural language models.

cross-lingualinterpretabilitysyntactic knowledge

Natural Language Processing RELIES on Linguistics

May 09, 2024
JO
Juri Opitz
🏛️ University of Zurich | Georgetown University

Recent large language models (LLMs) exhibit a superficial “de-linguistification” trend, marginalizing linguistics despite its foundational relevance to natural language processing (NLP). Method: This paper systematically reasserts linguistics’ indispensable structural role in NLP through an original six-dimensional RELIES framework—encompassing Resources, Evaluation, Low-resource settings, Interpretability, Explanation, and Study of language—and integrates linguistic insights via conceptual analysis, interdisciplinary synthesis, and empirical case studies. Contribution/Results: The work challenges the purely data-driven paradigm by demonstrating how linguistic theory methodologically anchors model architecture design, evaluation criteria, and ethical governance. It establishes linguistics not as auxiliary but as constitutive to NLP’s scientific rigor and human-centered grounding, offering a systematic, theory-informed roadmap for developing linguistically principled, interpretable, and equitable NLP systems.

Examines NLP's reliance on linguistics for grammar and semantics.Explores linguistic contributions to NLP in low-resource settings.Highlights linguistics' role in NLP interpretability and explanation.

Latest Papers

What's happening recently
View more

This work proposes CYKNN, a novel neural architecture that directly embeds the Cocke–Younger–Kasami (CYK) context-free grammar parsing algorithm into a recurrent neural network. By formulating the CYK algorithm through differentiable matrix-vector operations, CYKNN enables end-to-end trainable encoding of symbolic parsing within a neural framework, thereby achieving a deep integration of symbolic reasoning and neural computation. Empirical results demonstrate that CYKNN substantially outperforms both large language models with over 20 billion parameters and LoRA-finetuned variants of the Qwen model family on simple grammatical tasks. This approach establishes a promising new direction for neuro-symbolic systems by combining the expressivity of formal grammars with the learning capabilities of neural networks.

context-free grammarCYK algorithmneural architecture

Existing approaches to automatic formalization often overlook the hierarchical logical structure inherent in mathematical statements. This work proposes the DSR framework, which achieves modular formalization by decomposing statements, constructing operator trees, and iteratively refining and repairing subtrees. It introduces, for the first time, the topological structure of operator trees to guide error localization and correction, and presents PRIME, a high-quality benchmark of formalized theorems. By integrating neural-symbolic systems, large language models, and formal verification, DSR significantly outperforms current methods under the same computational budget, establishing a new state-of-the-art in automatic formalization.

autoformalizationformal languagehierarchical logic

This work addresses the tendency of large language models (LLMs) to generate outputs lacking verifiable syntactic structure, often resulting in structural errors and hallucinations. The authors propose a neurosymbolic framework that, for the first time, dynamically aligns the incremental derivation mechanism of Combinatory Categorial Grammar (CCG) with the prefix-driven generation process of LLMs. By leveraging the Curry–Howard isomorphism, the approach lifts model outputs into typed compositional derivations. This enables unified structural reconstruction across both natural language and formal languages—including SQL, Solidity, and OWL—and incorporates a two-tier verification mechanism to enforce structural consistency and enable early detection of factual inaccuracies in generated content.

Combinatory Categorial GrammarCompositionalityHallucination

This study investigates the efficient development of formal grammatical resources for Cantonese and Irish while preserving cross-linguistic consistency, and evaluates the potential of multilingual large language models (LLMs) to assist in grammar engineering for low-resource languages. Building upon the ParGram framework, we present the first systematic application of multilingual LLMs—such as gpt-oss-120b—to parallel treebank construction, leveraging model-assisted translation and syntactic structure generation to maintain alignment at the level of abstract functional representations. Experimental results indicate that model-generated translations show limited efficacy and are unaffected by the choice of prompt language; although syntactic generation captures predicate–argument relations to some extent, it underperforms on cross-linguistically abstract tasks, necessitating expert intervention. This work establishes a novel paradigm and empirical foundation for formal syntactic modeling of low-resource languages.

Cantonesegrammar engineeringIrish

Hot Scholars

ZJ

Zhi Jin

Sun Yat-Sen University, Associate Professor
YS

Yo-Sub Han

School of Computing, Yonsei University
automata theoryformal languagesalgorithm designinformation retrieval
HY

Hitomi Yanaka

The University of Tokyo, RIKEN
Natural Language ProcessingSemantics
NS

Nathan Schneider

Associate Professor, Linguistics and Computer Science • Georgetown University
natural language processingcomputational linguisticscognitive linguistics