Attention on Multiword Expressions: A Multilingual Study of BERT-based Models with Regard to Idiomaticity and Microsyntax

📅 2025-05-09
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates how fine-tuning objective—semantic versus syntactic—affects the attention allocation of multilingual BERT toward multiword expressions (MWEs). Experiments span six Indo-European languages (English, German, Dutch, Polish, Russian, Ukrainian), targeting two MWE types: idioms (semantically non-compositional) and micro-syntactic units (MSUs; syntactically irregular). Methodologically, we employ task-specific fine-tuning, construct a cross-lingual MWE-annotated dataset, and conduct layer-wise attention score analysis and visualization. Results reveal a systematic, task-driven stratification: semantic fine-tuning markedly enhances uniform attention to idioms in upper layers (especially layers 11–12), whereas syntactic fine-tuning selectively strengthens focused attention to MSUs in lower layers (layers 2–4). This work provides the first empirical evidence of a universal, cross-linguistically consistent layerwise attention shift in BERT induced by fine-tuning objective, offering novel insights into the structural–functional mapping of pretrained language models.

Technology Category

Natural Language Processing: Lexical Semantics and MorphologyMachine Learning: Large Multimodal Models (LMMs)Computer Vision: Language and Vision

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsWeb Mining and Content Analysis: Large pretrained models with web dataSearch and Retrieval-Augmented AI: Multilingual and cross-lingual Web search
📝 Abstract
This study analyzes the attention patterns of fine-tuned encoder-only models based on the BERT architecture (BERT-based models) towards two distinct types of Multiword Expressions (MWEs): idioms and microsyntactic units (MSUs). Idioms present challenges in semantic non-compositionality, whereas MSUs demonstrate unconventional syntactic behavior that does not conform to standard grammatical categorizations. We aim to understand whether fine-tuning BERT-based models on specific tasks influences their attention to MWEs, and how this attention differs between semantic and syntactic tasks. We examine attention scores to MWEs in both pre-trained and fine-tuned BERT-based models. We utilize monolingual models and datasets in six Indo-European languages - English, German, Dutch, Polish, Russian, and Ukrainian. Our results show that fine-tuning significantly influences how models allocate attention to MWEs. Specifically, models fine-tuned on semantic tasks tend to distribute attention to idiomatic expressions more evenly across layers. Models fine-tuned on syntactic tasks show an increase in attention to MSUs in the lower layers, corresponding with syntactic processing requirements.
Problem

Research questions and friction points this paper is trying to address.

Analyzes BERT attention patterns for idioms and microsyntactic units
Examines impact of task-specific fine-tuning on MWE attention allocation
Compares semantic vs syntactic task effects across six languages
Innovation

Methods, ideas, or system contributions that make the work stand out.

Analyzes BERT attention on idioms and MSUs
Fine-tunes models for semantic and syntactic tasks
Compares attention patterns across six languages
🔎 Similar Papers
No similar papers found.