๐ค AI Summary
This work addresses the challenge in decoder-only large language models for simultaneous machine translation, where positional misalignment hinders the simultaneous achievement of high efficiency and output consistency. To resolve this, the authors propose ExPosST, a general framework that explicitly allocates fixed positional slots for source tokens, enabling efficient decoding under various positional encodings through adaptive masking and KV cache optimization. Additionally, ExPosST introduces strategy-consistent fine-tuningโthe first of its kindโto align model behavior between training and inference. Experimental results demonstrate that ExPosST consistently delivers high-quality, low-latency simultaneous translation across multiple language pairs and translation strategies, significantly enhancing both compatibility and practicality.
๐ Abstract
Large language models (LLMs) have recently demonstrated promising performance in simultaneous machine translation (SimulMT). However, applying decoder-only LLMs to SimulMT introduces a positional mismatch, which leads to a dilemma between decoding efficiency and positional consistency. Existing approaches often rely on specific positional encodings or carefully designed prompting schemes, and thus fail to simultaneously achieve inference efficiency, positional consistency, and broad model compatibility. In this work, we propose ExPosST, a general framework that resolves this dilemma through explicit position allocation. ExPosST reserves fixed positional slots for incoming source tokens, enabling efficient decoding with KV cache across different positional encoding methods. To further bridge the gap between fine-tuning and inference, we introduce a policy-consistent fine-tuning strategy that aligns training with inference-time decoding behavior. Experiments across multiple language pairs demonstrate that ExPosST effectively supports simultaneous translation under diverse policies.