Diffusion Drafts, AR Verifies: Accelerating Document OCR with Self-Speculative Decoding
为解决OCR模型解码速度慢的问题,提出GravityOCR模型,结合并行生成与自回归验证,提高了解码效率和速度。
为解决OCR模型解码速度慢的问题,提出GravityOCR模型,结合并行生成与自回归验证,提高了解码效率和速度。
This work addresses the challenge of developing efficient Korean-centric multilingual large language models (LLMs) under resource constraints. We propose Trillion-7B—the first trillion-parameter, Korean-hubbed multilingual LLM optimized for token efficiency. Methodologically, we introduce cross-lingual document attention (XLDA), integrated with language-aware data mixing, multilingual filtering, and a customized tokenizer, enabling efficient English knowledge transfer using only 2T training tokens—of which just 10% are multilingual (Korean, Japanese, Chinese). Experiments demonstrate state-of-the-art or highly competitive performance across 27 English, Korean, Japanese, and Chinese benchmarks, with significantly improved cross-lingual consistency. Full training requires only 59.4K H100 GPU-hours (≈$1.48M), achieving the highest token efficiency among existing Korean-centric multilingual LLMs.
为解决OCR模型解码速度慢的问题,提出GravityOCR模型,结合并行生成与自回归验证,提高了解码效率和速度。
This work addresses the challenge of developing efficient Korean-centric multilingual large language models (LLMs) under resource constraints. We propose Trillion-7B—the first trillion-parameter, Korean-hubbed multilingual LLM optimized for token efficiency. Methodologically, we introduce cross-lingual document attention (XLDA), integrated with language-aware data mixing, multilingual filtering, and a customized tokenizer, enabling efficient English knowledge transfer using only 2T training tokens—of which just 10% are multilingual (Korean, Japanese, Chinese). Experiments demonstrate state-of-the-art or highly competitive performance across 27 English, Korean, Japanese, and Chinese benchmarks, with significantly improved cross-lingual consistency. Full training requires only 59.4K H100 GPU-hours (≈$1.48M), achieving the highest token efficiency among existing Korean-centric multilingual LLMs.