Cross-Lingual Alignment for Decoder-Only Models using MoE Routers
This study addresses the challenge of cross-lingual representation alignment in decoder-only large language models, which arises from tokenization discrepancies across languages. To tackle this issue, this work proposes a novel contrastive learning paradigm that leverages Mixture-of-Experts (MoE) router outputs as alignment anchors. Departing from conventional auxiliary losses applied to hidden states, the method constructs robust sequence-level alignment objectives through pooling and optimizes them via controlled continual pre-training. Experimental results demonstrate that the proposed approach effectively aligns lower-layer hidden representations, substantially enhancing the multilingual performance of open-source MoE models. Overall, this research offers a promising new direction for achieving cross-lingual alignment in large language models.