Scoring Both Directions: LLMs realize the MRS they cannot reliably parse

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether the ability of large language models (LLMs) to generate high-scoring outputs equates to genuine comprehension of formal semantics. Leveraging the English Resource Grammar (ERG), the ACE processor, and BLEU/F1 metrics, it systematically evaluates the bidirectional conversion capabilities of the Claude model family between natural language text and Minimal Recursion Semantics (MRS). The findings challenge the prevailing assumption that generation implies understanding. Although the Opus model demonstrates strong generation performance (76.3 BLEU), its parsing F1 score reaches only 65.5 with an exact match rate of approximately 1%, significantly underperforming traditional grammar-based systems. This work establishes that generation metrics alone are insufficient to demonstrate that LLMs possess authentic formal semantic parsing capabilities, thereby exposing a critical misconception in evaluating their linguistic competence.
📝 Abstract
The English Resource Grammar (ERG) is a hand-written computational grammar of English. Given a sentence, its processor, ACE, produces a formal meaning representation called Minimal Recursion Semantics (MRS): a graph of the sentence's predicates and their arguments. The grammar is bidirectional and can also turn an MRS back into an English sentence. \citet{hajdik2019} used the ERG's treebank to build a benchmark for that generation task, MRS to text, and trained sequence-to-sequence models to solve it. The parsing task, text to MRS, can be tested on the same sentences. We reconstruct their 10K-sentence test split, and score two large language models, Claude Sonnet~4.5 and Claude Opus~5, in both directions against their trained systems and against ACE, with no task-specific training. Given an MRS and three examples, Opus writes the sentence at 76.3 BLEU, ten points above their system trained on 72k pairs (66.1 BLEU), and comparable to their system trained on a million extra pairs (77.2 BLEU). Sonnet scores 65.7 BLEU, and letting it choose among ACE's own candidate sentences lifts it to 69.6, while a pooled judge that keeps Opus's own sentence among the candidates adds 0.6 points (77.0 BLEU). In the parsing direction, however, the models fall far behind ACE: asked for the MRS of the same sentences, they reach 57.2 (Sonnet) and 65.5 (Opus) F$_1$ on the graph's predicates and arguments against 91.0 for ACE, and exact-match the gold on about 1\% of sentences. We characterize the failure modes for the parsing tasks, and conclude that a generation score alone does not show that models understand formal semantic representations.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Minimal Recursion Semantics
Semantic Parsing
Text Generation
English Resource Grammar
Innovation

Methods, ideas, or system contributions that make the work stand out.

Large Language Models
Minimal Recursion Semantics
English Resource Grammar
Bidirectional Evaluation
Semantic Parsing