๐ค AI Summary
This study addresses context fragmentation and terminology inconsistency in long-form subtitle translation by proposing SMART, a self-evolving multi-agent system. Without retraining large language models, the method constructs persistent memory via test-time training and integrates dynamic routing, a mixture-of-experts architecture, and a critic-based self-improvement algorithm to autonomously optimize prompts and strategies. Furthermore, it introduces Subtitle Arena, a benchmark spanning 14 genres and 15 languages, alongside the SubMQM evaluation framework. Experimental results demonstrate that SMART achieves state-of-the-art MQM scores across 15 translation directions, reducing average penalties by 6.9%, and attains the best model and human evaluation results on the MuSC benchmark with a score of 4.50 out of 5.
๐ Abstract
Long-form subtitle translation requires reasoning over discourse and cultural context spanning episodes or entire series, while maintaining consistent terminology and style. Existing single-LLM methods are largely sentence-level, and multi-agent systems often use static workflows that do not adapt to scene complexity or production context. We propose SMART, a Self-evolving Multi-Agent system for long-foRm subtitle Translation. During test-time training, SMART builds persistent series-level memory and translates a subset of sentences through a dynamic router and Mixture-of-Agents layer with tools for terminology verification, subtitle constraint validation, and contextual retrieval. A judge-refiner loop scores candidates and uses textual critiques to update agent prompts and routing policies without retraining the underlying LLMs. During test-time inference, the evolved configuration translates the remaining series. We also introduce Subtitle Arena, covering 14 genres, 2--198 episodes per series, production years 1959--2023, and 15 target locales, together with SubMQM, a subtitle-adapted MQM framework with seven dimensions and 19 error categories. SMART achieves the best overall MQM score in all 15 Subtitle Arena directions, reducing average penalty by 6.9% over the strongest competing agent system. On the public MuSC benchmark, SMART obtains the best model result across all four language pairs and also achieves the best human-evaluation result, with an overall score of 4.50/5.