🤖 AI Summary
This study addresses the limited transferability of existing protein structure tokenization methods, which typically rely on complex training objectives and large-scale data. We propose ProFiT, a lightweight tokenizer that naturally learns structural representations without manual semantic annotations, thereby significantly reducing engineering complexity. Furthermore, Flow Matching is introduced to optimize codebook utilization, enabling efficient training and semantic alignment. Experimental results demonstrate that ProFiT surpasses existing large-scale models in both reconstruction quality and generalization performance. Owing to its plug-and-play design, ProFiT offers a concise and effective solution for protein structure representation.
📝 Abstract
As the bridge between protein modality and discrete modeling, protein structure tokenization still largely relies on heavily engineered training objectives tailored to specific downstream tasks and large training datasets, which hinders its transfer to broader application scenarios. To address this issue, we propose ProFiT, a lightweight flow matching tokenizer. With simple training strategies that encourage healthy codebook utilization, ProFiT can be trained efficiently and naturally learns semantically meaningful representations without any manual semantic alignment, while achieving reconstruction quality and generalization that match or surpass those of substantially larger tokenizers. We conduct extensive evaluations across a wide range of settings and demonstrate that ProFiT is a plug-and-play tokenizer adaptable to diverse downstream tasks. This study further reveals the significant potential of the flow matching tokenizer paradigm. Our code is publicly available at https://github.com/QDKStorm/ProFiT.