train sequence-to-sequence translation models

Designs, implements, and evaluates training pipelines for sequence-to-sequence translation models that map input sequences in one language to output sequences in another. This work encompasses preparing parallel corpora and tokenization, selecting and configuring encoder–decoder architectures (e.g., attention or transformer variants), defining loss functions and optimization schedules, managing batching, regularization and checkpointing, performing hyperparameter tuning, and measuring translation quality with sequence-level evaluation metrics.

trainsequence-to-sequencetranslationmodels

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.33
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

A U-Net and Transformer Pipeline for Multilingual Image Translation

Oct 27, 2025
SS
Siddharth Sahay
🏛️ B.M.S. College of Engineering

This work addresses end-to-end multilingual image-to-text translation. Methodologically, it proposes a lightweight, fully customizable pipeline: (1) a custom U-Net architecture for robust text region detection—enhanced via synthetic data augmentation to improve generalization; (2) Tesseract-based OCR for text recognition; and (3) a from-scratch, multilingual Transformer model trained for neural machine translation across Chinese, English, Japanese, Korean, and French. Crucially, the framework eliminates reliance on large pretrained models, instead adopting a modular, plug-and-play design that enhances adaptability and deployment flexibility in resource-constrained environments. Experimental results demonstrate strong performance across text detection accuracy, OCR quality, and BLEU scores, validating the effectiveness and feasibility of an entirely self-contained, non-pretrained modeling approach for multimodal translation tasks.

Detecting text regions in images using custom U-Net modelExtracting multilingual text via Tesseract OCR engineTranslating extracted text with custom Seq2Seq Transformer model

Parallel Token Prediction for Language Models

Dec 24, 2025
FD
Felix Draxler
🏛️ University of California, Irvine | Chan-Zuckerberg Initiative | Pyramidal AI | AI in Science Institute

Autoregressive decoding in large language models incurs high latency, and existing multi-token prediction methods rely on strong independence assumptions that limit modeling fidelity. Method: We propose Parallel Token Prediction (PTP), the first framework to internalize the sampling process into the model architecture, enabling joint generation of multiple semantically coherent tokens in a single Transformer forward pass while strictly preserving expressivity over any autoregressive distribution—thereby eliminating restrictive independence assumptions. PTP integrates inverse autoregressive training with explicit sampling modeling and supports both teacher-free and distillation-based training. Results: Evaluated on Vicuna-7B, PTP achieves 4.12 average accepted tokens per speculative step on Spec-Bench, maintains full modeling capability for long-sequence generation, and attains state-of-the-art performance in speculative decoding.

Avoids restrictive independence assumptions in token predictionEnables parallel generation without loss of modeling powerReduces latency bottleneck in autoregressive decoding

The impact of quantization on multilingual machine translation—particularly for low-resource languages—remains poorly understood. Method: We systematically evaluate four post-training quantization methods—AWQ, BitsAndBytes, GGUF, and AutoRound—across 55 languages at 4-bit and 2-bit precision. Contribution/Results: Our study is the first to reveal that quantization error intensifies with decreasing language resource availability and varies across language families: 2-bit quantization severely degrades translation quality for low-resource languages, whereas 4-bit quantization preserves performance for high-resource languages. GGUF demonstrates superior robustness under 2-bit quantization. Furthermore, we validate that language-matched calibration effectively mitigates low-bit degradation. These findings provide empirical evidence and practical guidance for lightweight deployment of multilingual LLMs in resource-constrained multilingual settings.

Assessing quality degradation in low-resource languages under quantizationComparing quantization techniques and calibration strategies for optimal deploymentEvaluating post-training quantization impact on multilingual machine translation

A Recipe of Parallel Corpora Exploitation for Multilingual Large Language Models

Jun 29, 2024
PL
Peiqin Lin
🏛️ LMU Munich | Instituto Superior Técnico | Universidade de Lisboa | Instituto de Telecomunicações | Unbabel

This study systematically investigates optimal strategies for leveraging parallel corpora to enhance multilingual large language models (MLLMs), focusing on the impacts of corpus quality and scale, training objectives, and model parameter count on both bilingual tasks (e.g., machine translation) and general cross-lingual tasks (e.g., text classification). Method: We propose a noise-filtering–based parallel corpus selection mechanism—bypassing error-prone language identification preprocessing—and employ supervised fine-tuning with a pure machine translation (MT) objective, integrated with multilingual pretraining. Contribution/Results: We find that merely ~10K high-quality parallel sentence pairs achieve performance comparable to large-scale corpora; the MT-only objective substantially outperforms multitask mixed objectives; and larger models benefit more markedly from parallel data. Experiments demonstrate consistent improvements across 12 languages and five cross-lingual task categories, establishing a reusable, efficient paradigm for parallel corpus utilization in MLLM development.

Enhancing large language models with parallel dataImproving performance across diverse tasksOptimizing parallel corpora for multilingual models

Training Ultra Long Context Language Model with Fully Pipelined Distributed Transformer

Aug 30, 2024
JY
Jinghan Yao
🏛️ The Ohio State University | Microsoft Inc.

To address the high GPU memory consumption and hardware requirements in training large language models (LLMs) with ultra-long contexts, this paper proposes a fully pipelined distributed Transformer architecture. Its core innovation is a novel sequence-chunking pipeline mechanism that extends maximum sequence length by 16× without modifying the model architecture, while maintaining full compatibility with existing training techniques—including pipeline parallelism, FP16/BF16 mixed-precision training, and computational graph reordering. We successfully train an 8B-parameter model on just four GPUs to handle sequences up to 2 million tokens, achieving a sustained model FLOPs utilization (MFU) of over 55%, significantly outperforming state-of-the-art approaches. This method substantially reduces hardware dependency and resource costs for long-context LLM training, establishing a general, scalable, and efficient paradigm for ultra-long-context LLM training.

Efficiently scaling sequence length for LLMs without hardware expansionExisting adaptation methods impose significant design limitationsTraining long-context LLMs requires excessive GPU resources and memory

Latest Papers

What's happening recently
View more

ParaDySe: A Parallel-Strategy Switching Framework for Dynamic Sequence Lengths in Transformer

Nov 17, 2025
ZO
Zhixin Ou
🏛️ National University of Defense Technology

Static parallelism strategies in large Transformer training struggle to adapt to dynamic sequence lengths: short sequences trigger Communication Parallelism Cancellation (CPC), while long sequences cause out-of-memory (OOM) errors. To address this, we propose a sequence-aware dynamic parallelism switching framework. Our method introduces a modular parallel library built upon a unified tensor layout, enabling fine-grained, hot-swappable inter-layer parallelism selection driven by sequence length. We further develop a lightweight hybrid memory and time cost model and integrate a heuristic algorithm for real-time optimal parallelism decisions. Evaluated on 624K-length sequence training, our framework completely eliminates both OOM and CPC bottlenecks, significantly improving training stability and GPU resource utilization. Key contributions include: (1) the first sequence-length–driven dynamic parallelism switching mechanism; (2) a unified, modular parallel library supporting heterogeneous layer-wise strategies; and (3) an efficient hybrid cost model with real-time decision capability.

Addressing memory and communication inefficiencies in large language modelsEnabling adaptive strategy switching for varying input sequence lengthsOptimizing parallel strategies for dynamic sequence lengths in Transformer training

This work addresses the inherent trade-off between low latency and high translation quality in conventional simultaneous speech translation systems, which typically rely on offline models coupled with handcrafted read/write policies. To overcome this limitation, the authors propose Hikari, an end-to-end model that unifies simultaneous speech translation and streaming transcription by modeling read/write decisions as probabilistic WAIT tokens, thereby eliminating the need for explicit policy design. The approach integrates a causal alignment architecture, decoder time dilation, and delay-robustness-oriented supervised fine-tuning. Evaluated on English-to-Japanese, German, and Russian tasks, Hikari achieves new state-of-the-art BLEU scores across both low- and high-latency scenarios, significantly outperforming existing baselines.

End-to-End ModelingLow LatencySimultaneous Machine Translation

This study systematically evaluates the practical benefits of pretraining in DNA language models for downstream genomic tasks and investigates the effectiveness of Byte Pair Encoding (BPE) tokenization. By comparing Transformer-based architectures (e.g., DNABERT2) with convolutional models (e.g., ConvNova) across a range of genomic fine-tuning benchmarks, the work provides the first empirical analysis of whether pretraining is necessary and how BPE compares to traditional k-mer representations. The findings indicate that the performance gains conferred by pretraining are limited and that BPE does not consistently outperform k-mer tokenization across all tasks. These results offer critical empirical insights for the design of foundational models in genomics, challenging prevailing assumptions about the universal advantages of large-scale pretraining and subword tokenization in this domain.

BPE tokenizationDNA language modelsfine-tuning

Automated Snippet-Alignment Data Augmentation for Code Translation

Oct 15, 2025
ZZ
Zhiming Zhang
🏛️ Harbin Institute of Technology

Existing code translation research primarily focuses on program-level alignment (PA), neglecting finer-grained snippet-level alignment (SA), thereby limiting models’ capacity to capture local semantic and structural mappings. This work proposes, for the first time, a large language model–based approach to automatically generate high-quality SA parallel data. We further design a two-stage fine-tuning framework that jointly leverages PA and SA: Stage I performs pretraining on program-level data, while Stage II refines the model using snippet-level aligned data. This integration significantly improves the model’s ability to preserve semantic consistency and structural fidelity across cross-lingual code snippets. Evaluated on the TransCoder-test benchmark, our method achieves up to a 3.78% absolute improvement in pass@k, demonstrating its effectiveness, generalizability, and robustness.

Automatically generates snippet-alignment data for code translationEnhances model training with two-stage strategy using PA and SA dataImproves code translation accuracy by augmenting fine-grained parallel corpora

This work addresses the challenge of representation conflict in low-resource multilingual speech translation caused by uniform cross-lingual parameter sharing. To mitigate this limitation, the authors propose a fine-grained sharing strategy guided by training gradient analysis, which automatically determines language-specific sharing patterns across model layers through a three-tier mechanism: first, languages are clustered based on gradient distances; second, model capacity is dynamically allocated according to intra- and inter-task gradient divergence; and third, subspace alignment is achieved via joint factorization coupled with canonical correlation analysis. Evaluated on the SeamlessM4T-Medium architecture across four language pairs, the approach yields significant improvements in translation quality, demonstrating the effectiveness and generalizability of gradient-driven parameter sharing in multilingual speech translation.

architectural sharingconvergencelow-resource