🤖 AI Summary
This study investigates how fine-tuning objective—semantic versus syntactic—affects the attention allocation of multilingual BERT toward multiword expressions (MWEs). Experiments span six Indo-European languages (English, German, Dutch, Polish, Russian, Ukrainian), targeting two MWE types: idioms (semantically non-compositional) and micro-syntactic units (MSUs; syntactically irregular). Methodologically, we employ task-specific fine-tuning, construct a cross-lingual MWE-annotated dataset, and conduct layer-wise attention score analysis and visualization. Results reveal a systematic, task-driven stratification: semantic fine-tuning markedly enhances uniform attention to idioms in upper layers (especially layers 11–12), whereas syntactic fine-tuning selectively strengthens focused attention to MSUs in lower layers (layers 2–4). This work provides the first empirical evidence of a universal, cross-linguistically consistent layerwise attention shift in BERT induced by fine-tuning objective, offering novel insights into the structural–functional mapping of pretrained language models.
📝 Abstract
This study analyzes the attention patterns of fine-tuned encoder-only models based on the BERT architecture (BERT-based models) towards two distinct types of Multiword Expressions (MWEs): idioms and microsyntactic units (MSUs). Idioms present challenges in semantic non-compositionality, whereas MSUs demonstrate unconventional syntactic behavior that does not conform to standard grammatical categorizations. We aim to understand whether fine-tuning BERT-based models on specific tasks influences their attention to MWEs, and how this attention differs between semantic and syntactic tasks. We examine attention scores to MWEs in both pre-trained and fine-tuned BERT-based models. We utilize monolingual models and datasets in six Indo-European languages - English, German, Dutch, Polish, Russian, and Ukrainian. Our results show that fine-tuning significantly influences how models allocate attention to MWEs. Specifically, models fine-tuned on semantic tasks tend to distribute attention to idiomatic expressions more evenly across layers. Models fine-tuned on syntactic tasks show an increase in attention to MSUs in the lower layers, corresponding with syntactic processing requirements.