Flow-Matching-Based Protein Structure Tokenizer Made Efficient and Easy

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited transferability of existing protein structure tokenization methods, which typically rely on complex training objectives and large-scale data. We propose ProFiT, a lightweight tokenizer that naturally learns structural representations without manual semantic annotations, thereby significantly reducing engineering complexity. Furthermore, Flow Matching is introduced to optimize codebook utilization, enabling efficient training and semantic alignment. Experimental results demonstrate that ProFiT surpasses existing large-scale models in both reconstruction quality and generalization performance. Owing to its plug-and-play design, ProFiT offers a concise and effective solution for protein structure representation.
📝 Abstract
As the bridge between protein modality and discrete modeling, protein structure tokenization still largely relies on heavily engineered training objectives tailored to specific downstream tasks and large training datasets, which hinders its transfer to broader application scenarios. To address this issue, we propose ProFiT, a lightweight flow matching tokenizer. With simple training strategies that encourage healthy codebook utilization, ProFiT can be trained efficiently and naturally learns semantically meaningful representations without any manual semantic alignment, while achieving reconstruction quality and generalization that match or surpass those of substantially larger tokenizers. We conduct extensive evaluations across a wide range of settings and demonstrate that ProFiT is a plug-and-play tokenizer adaptable to diverse downstream tasks. This study further reveals the significant potential of the flow matching tokenizer paradigm. Our code is publicly available at https://github.com/QDKStorm/ProFiT.
Problem

Research questions and friction points this paper is trying to address.

protein structure tokenization
flow matching
downstream tasks
generalization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Flow Matching
Protein Structure Tokenizer
Lightweight Model
Codebook Utilization
Plug-and-Play
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Z
Zhe Zhang
Institute for AI Industry Research (AIR), Tsinghua University; Department of Computer Science and Technology, Tsinghua University
Yikai Zhang
Yikai Zhang
Fudan university
Natural Language ProcessingAutonomous Agent
J
Jiangtao Feng
Institute for AI Industry Research (AIR), Tsinghua University
Y
Ya-Qin Zhang
Institute for AI Industry Research (AIR), Tsinghua University
Wei-Ying Ma
Wei-Ying Ma
Tsinghua University
Generative AI and Large Language Models (LLMs) for Science
H
Hao Zhou
Institute for AI Industry Research (AIR), Tsinghua University