ExPosST: Explicit Positioning with Adaptive Masking for LLM-Based Simultaneous Machine Translation

๐Ÿ“… 2026-03-16
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the challenge in decoder-only large language models for simultaneous machine translation, where positional misalignment hinders the simultaneous achievement of high efficiency and output consistency. To resolve this, the authors propose ExPosST, a general framework that explicitly allocates fixed positional slots for source tokens, enabling efficient decoding under various positional encodings through adaptive masking and KV cache optimization. Additionally, ExPosST introduces strategy-consistent fine-tuningโ€”the first of its kindโ€”to align model behavior between training and inference. Experimental results demonstrate that ExPosST consistently delivers high-quality, low-latency simultaneous translation across multiple language pairs and translation strategies, significantly enhancing both compatibility and practicality.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: (Large) Language ModelsPlanning, Routing, and Scheduling: Planning with Language Models

Application Category

Search and Retrieval-Augmented AI: Large language models for searchSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendation
๐Ÿ“ Abstract
Large language models (LLMs) have recently demonstrated promising performance in simultaneous machine translation (SimulMT). However, applying decoder-only LLMs to SimulMT introduces a positional mismatch, which leads to a dilemma between decoding efficiency and positional consistency. Existing approaches often rely on specific positional encodings or carefully designed prompting schemes, and thus fail to simultaneously achieve inference efficiency, positional consistency, and broad model compatibility. In this work, we propose ExPosST, a general framework that resolves this dilemma through explicit position allocation. ExPosST reserves fixed positional slots for incoming source tokens, enabling efficient decoding with KV cache across different positional encoding methods. To further bridge the gap between fine-tuning and inference, we introduce a policy-consistent fine-tuning strategy that aligns training with inference-time decoding behavior. Experiments across multiple language pairs demonstrate that ExPosST effectively supports simultaneous translation under diverse policies.
Problem

Research questions and friction points this paper is trying to address.

simultaneous machine translation
large language models
positional mismatch
decoding efficiency
positional consistency
Innovation

Methods, ideas, or system contributions that make the work stand out.

explicit positioning
adaptive masking
simultaneous machine translation
KV cache efficiency
policy-consistent fine-tuning
๐Ÿ”Ž Similar Papers
2024-02-16International Workshop on Spoken Language TranslationCitations: 4