Dual-Vocabulary Language Model for Cross-Tokenizer Distillation

πŸ“… 2026-09-27
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the alignment challenges of input tokenization and output logits in cross-tokenizer knowledge distillation by proposing a dual-vocabulary language model framework. Methodologically, it introduces Parallel Token Sequences (PTS) and Hybrid Prefix Attention (HPA) mechanisms to achieve distributional alignment at the input side between teacher and student models. Furthermore, this work pioneers replacing the teacher’s language modeling head with a projection head to obtain full-dimensional student logits, thereby ensuring output-side alignment while preserving the integrity of the native tokenization paradigm. Experimental results demonstrate that the proposed approach yields loss convergence consistent with the teacher model and significantly enhances student performance across six reasoning tasks.
πŸ“ Abstract
On-policy distillation (OPD) bridges teacher supervision and student behavior, but different teacher-student tokenizers introduce misalignment in both input tokenization (#1) and output logits (#2). Existing approaches address the former by matching same-text spans or converting tokens to bytes, often losing fine-grained token information or disrupting the native-token paradigm, while for the latter, strategies such as ranking, padding, or key-token selection retain only shared logit dimensions, resulting in much distribution loss. In this paper, we propose Dual-Vocabulary Language Model (DVLM), which replaces the teacher's LM head with a new student-vocabulary projection head and obtains full-dimensional student logits (for #2). To support student tokens (for #1), it takes a Parallel-Tokenized Sequence (PTS) as input, which concatenates the original teacher-tokenized sequence and a re-tokenized sequence formed by independently converting each student token into a teacher-token group. To avoid inference inconsistency with the original teacher tokens, the Hybrid-Prefix Attention (HPA) further restricts re-tokenized groups to their corresponding teacher prefix and uses its last state as the aggregation of the original student-token representation for projection into the student vocabulary space. Similarly, via the combined use of PTS and HPA, the DVLM teacher can provide distribution-aligned supervision with the student's input-tokenization and output-logit during OPD. Experimental results demonstrate that our DVLM teacher has a similar converged loss as the original teacher model and enables student models to improve performance across six reasoning tasks.
Problem

Research questions and friction points this paper is trying to address.

cross-tokenizer distillation
on-policy distillation
tokenizer misalignment
logit distribution loss
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dual-Vocabulary Language Model
Cross-Tokenizer Distillation
Parallel-Tokenized Sequence
Hybrid-Prefix Attention
On-policy Distillation
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
K
Kedi Chen
East China Normal University
C
Chen Lin
East China Normal University
Yutao Sun
Yutao Sun
Tsinghua University
Natural Language ProcessingMachine Learning
W
Wei Zhang
East China Normal University