Score
Designs, builds, and evaluates models, pipelines, and evaluation methods that extract, represent, and reason about the meaning of human natural language input (text or speech). This work covers creating representations and components for tasks such as syntactic and semantic parsing, named-entity and relation extraction, intent and slot detection, coreference and discourse resolution, semantic role labeling, and grounding language in conversational or situational context.
本文探讨了基于变换器的语言模型如何学习和表示语言意义,并介绍了三种主要的语义理论,比较了这些模型与人类学习语言方式的区别。
This work proposes a fine-grained paradigm for modeling semantic equivalence by decomposing paraphrasing into specific linguistic operations—such as lexical substitution and syntactic transformation—to construct structured, cognitively plausible representations of meaning. Rather than relying on coarse binary classification or single rewrites, the approach explicitly captures the nuanced mechanisms underlying paraphrase generation. A deep semantic model trained on annotated data achieves 89.6% accuracy on a Wikipedia plagiarism detection task and 66.5% on an arXiv dataset, substantially outperforming human baselines. Furthermore, the method demonstrates consistent performance gains on downstream tasks such as Quora duplicate question detection, validating its effectiveness in both paraphrase understanding and controllable generation.
This paper addresses the insufficient modeling of textual implicit semantics by proposing a novel paradigm that explicitly transforms “subtext” into a verifiable set of propositions. Methodologically, it systematically leverages large language models to generate implicit inference propositions from text, which are then validated for plausibility by human annotators; the resulting proposition semantics are subsequently fused into the original text representations. Key contributions include: (1) introducing the first annotated framework for implicit propositions tailored to social science tasks; and (2) demonstrating that this explicit modeling significantly outperforms literal-only representations across three distinct tasks—argument similarity assessment, public opinion interpretation, and legislative behavior simulation—with average improvements of 12.7% in F1 or accuracy. Results substantiate the effectiveness and generalizability of structured implicit semantic modeling for enhancing human-like semantic understanding.
This paper addresses the challenge of jointly modeling pure semantic values and contextual effects—such as reference, tense, interrogation, and negation—in compositional natural language semantics. We propose an effect-driven semantic framework inspired by denotational semantics in programming languages. Methodologically, we systematically introduce functorial structures from category theory to formally characterize the hierarchical interaction between value propagation and side-effectful processes; we integrate type-logical syntax with functional semantic composition to build an extensible interpretation system. Our contributions include a unified treatment of tense, questions, negation, and discourse coherence, significantly enhancing compositional productivity, interpretability, and theoretical unity in semantic parsing. The framework establishes a new paradigm for computational semantics that balances formal rigor with broad linguistic coverage.
This work addresses the challenge of precisely controlling large language models’ (LLMs) reasoning capabilities without fine-tuning. We propose a representation-level intervention method that identifies task-relevant activations within the residual stream, constructs task-specific control vectors from them, and directly modifies representations during inference to enhance inductive, deductive, and mathematical reasoning. To our knowledge, this is the first systematic application of representation engineering to LLM reasoning control—relying solely on forward-pass activation extraction and residual-stream intervention, thereby revealing the intrinsic decomposability of reasoning abilities. Evaluated on Mistral-7B-Instruct and Pythia models, our approach consistently improves accuracy across diverse reasoning benchmarks, stabilizes logit distributions, and its mechanistic validity is confirmed via KL divergence and entropy analysis. All code and analytical tools are publicly released to support reproducible representation intervention research.
Current large language models (LLMs) exhibit significant limitations in interpreting non-literal semantics, particularly idioms and metaphors. To address this, we introduce IdiomMetaphorBank—the first large-scale, multi-category, human-annotated dataset for idiom and metaphor understanding. Our methodology employs a scalable data framework integrating corpus-driven automatic extraction with expert-level manual annotation, augmented by context-aware, model-agnostic post-processing to support both slot-filling and sequence labeling tasks. Empirically, IdiomMetaphorBank improves F1 scores of mainstream pretrained models on idiom identification by 12.3%. Moreover, it enables, for the first time, fine-grained evaluation of implicit semantic detection—e.g., underlying conceptual mappings and figurative intent—thereby establishing a benchmark resource and methodological foundation for non-literal semantic modeling.
This paper investigates whether large language models (LLMs) possess genuine semantic understanding at lexical and sentential levels, focusing on core semantic phenomena—namely, reference and propositional attitude. Method: It introduces, for the first time, the Frege–Russell tradition of formal semantics to construct an integrated theoretical framework bridging philosophical semantics and computational analysis; combining Transformer-based modeling, semantic probing, Concept Activation Vectors (CAVs), and formal semantic modeling to empirically examine the internal semantic structure of LLMs’ linguistic representations. Contribution/Results: Results indicate that while LLMs exhibit semantically plausible behavior in specific tasks, they lack stable referential mechanisms and robust propositional attitude representations. Their “understanding” is fundamentally statistical association rather than meaning-based comprehension. The study establishes a cross-disciplinary methodology and theoretical criteria for evaluating semantic competence in AI systems.
This paper systematically evaluates the genuine metaphor comprehension capabilities of large language models (LLMs), investigating whether their performance on metaphor interpretation tasks stems from deep semantic understanding or reliance on superficial cues—such as lexical overlap, sentence length, or syntactic patterns—in natural language inference (NLI) and question answering (QA). Method: Leveraging multiple public metaphor datasets, the authors design diverse prompting strategies and controlled ablation experiments to quantify correlations between model performance and surface-level linguistic features. Contribution/Results: Results demonstrate that LLMs’ metaphor interpretation is predominantly driven by surface features, in-context learning, and pretraining-induced linguistic priors—not by robust semantic reasoning. Consequently, standard evaluation protocols yield overly optimistic estimates of metaphor understanding. The paper proposes a more rigorous, bias-mitigated evaluation paradigm for metaphor comprehension and open-sources all data, code, and analysis frameworks to enable reproducible research and establish a new benchmark for trustworthy metaphor modeling.
This study investigates the similarities and differences between large language models (LLMs) and humans in their cognitive mechanisms underlying linguistic knowledge representation and real-world reasoning. Integrating methods from cognitive science and artificial intelligence, the research systematically compares the two in terms of representational structure, learning efficiency, and generalization capabilities through behavioral experiments and model analyses. Findings reveal that, despite LLMs’ fluent performance on linguistic tasks, their internal representations and processing mechanisms differ substantially from those of humans. Moreover, LLMs exhibit markedly inferior learning efficiency and generalization in real-world reasoning tasks. This work challenges the prevailing paradigm of equating task performance with human-likeness and offers a novel perspective on the fundamental distinctions between artificial and human intelligence.
Current language models lack systematic evaluation on their ability to understand multiword expressions—such as idioms, noun compounds, and verb constructions—that involve deep semantic processing. This work proposes SemanticQA, a benchmark suite that, for the first time, unifies disparate multiword expression resources into a structured semantic reasoning framework encompassing four task types: extraction, classification, interpretation, and composition. Designed to support comprehensive evaluation across diverse model architectures and scales, SemanticQA enables rigorous assessment of semantic competence. Experimental results reveal significant deficiencies in existing models’ capacity to handle non-literal meanings and complex syntactic-semantic structures, thereby offering both empirical evidence and a foundational benchmark for advancing language models’ semantic reasoning capabilities.