Supplementary Features of BiLSTM for Enhanced Sequence Labeling

📅 2023-05-31
📈 Citations: 6
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address BiLSTM’s limitation in capturing global sentence-level semantics for sequence labeling, this paper observes that its initial and final hidden states naturally encode holistic sentence representations. Leveraging this insight, we propose a lightweight, plug-and-play global context gating mechanism that dynamically injects sentence-level information into token-level representations at each time step. The method requires no architectural modification to the underlying RNN and is fully compatible with pretrained embeddings (e.g., BERT). Compared to various RNN variants, it achieves faster inference and training, and greater integration flexibility. Extensive experiments across nine benchmark datasets—including NER, POS tagging, and end-to-end aspect-based sentiment analysis—demonstrate consistent improvements: average F1-score and accuracy gains of 1.2–2.8 percentage points. These results validate the effectiveness and generalizability of synergistic global–local modeling for sequence labeling.
📝 Abstract
Sequence labeling tasks require the computation of sentence representations for each word within a given sentence. A prevalent method incorporates a Bi-directional Long Short-Term Memory (BiLSTM) layer to enhance the sequence structure information. However, empirical evidence Li (2020) suggests that the capacity of BiLSTM to produce sentence representations for sequence labeling tasks is inherently limited. This limitation primarily results from the integration of fragments from past and future sentence representations to formulate a complete sentence representation. In this study, we observed that the entire sentence representation, found in both the first and last cells of BiLSTM, can supplement each the individual sentence representation of each cell. Accordingly, we devised a global context mechanism to integrate entire future and past sentence representations into each cell's sentence representation within the BiLSTM framework. By incorporating the BERT model within BiLSTM as a demonstration, and conducting exhaustive experiments on nine datasets for sequence labeling tasks, including named entity recognition (NER), part of speech (POS) tagging, and End-to-End Aspect-Based sentiment analysis (E2E-ABSA). We noted significant improvements in F1 scores and accuracy across all examined datasets.
Problem

Research questions and friction points this paper is trying to address.

Enhancing global context capture in sequence labeling tasks
Improving efficiency for BiLSTM and transformer-based models
Enabling easy integration into existing architectures without speed loss
Innovation

Methods, ideas, or system contributions that make the work stand out.

Efficient global context for BiLSTM and transformers
Minimal speed degradation in training and inference
Pluggable into existing architectures easily
🔎 Similar Papers
No similar papers found.
Aalborg University | Rizhao Polytechnic | Northeast Normal University
C
Conglei Xu
Department of Computer Science, Aalborg University, Aalborg East, 9220, Denmark
K
Kun Shen
Department of Electronic information Engineering, Rizhao Polytechnic, Rizhao, 276800, China
Hongguang Sun
Hongguang Sun
School of Information Science and Technology, Northeast Normal University, Changchun, 130117, China