🤖 AI Summary
This study addresses the limitation of existing enzymatic reaction models that overlook the joint dependency between catalytic function and molecular structure. To this end, it proposes VenusRX, a T5-architecture pre-trained model, along with VenusRX-Bench, the first unified benchmark for this domain. Methodologically, the approach employs a two-stage pre-training strategy, an EC-conditioned embedding mechanism, and molecule-library-constrained decoding to jointly model forward prediction, retrosynthesis, and EC number inference tasks, thereby bridging the domain gap between chemical and enzymatic reactions. Experimental results demonstrate that VenusRX outperforms representative baselines across most evaluations, confirming that incorporating EC information significantly enhances predictive accuracy and enables comprehensive modeling of the enzymatic reaction space.
📝 Abstract
Existing reaction models primarily learn molecular transformations, whereas enzy- matic reactions depend jointly on molecular structure and catalytic function. We formulate this problem as learning an enzymatic reaction space linking reactants, products, and Enzyme Commission (EC) annotations. To characterize this space, we introduce VenusRX-Bench, a unified benchmark for forward reaction prediction, single-step retrosynthesis, and EC-number prediction. VenusRX-Bench integrates reactions from multiple biochemical databases with standardized curation, leakage- controlled splits, and consistent evaluation. Benchmarking representative chemical and enzymatic models reveals a clear chemical-to-enzymatic domain gap, driven by limited domain data, catalytic-context dependency, and the difficulty of modeling large biomolecular structures. To bridge this gap, we develop VenusRX, a unified T5-style sequence-to-sequence model for enzymatic reactions. VenusRX jointly learns forward prediction, ret- rosynthesis, and reaction reconstruction, with two-stage training on millions of template-expanded reactions followed by real biochemical reactions. In addition, optional EC conditioning incorporates catalytic context, while Molecule Library- Constrained Decoding improves the generation of complex biomolecules. Across benchmark tasks and challenging generalization splits, VenusRX achieves the best or competitive performance on most evaluated settings over representative chem- ical and enzymatic baselines. Moreover, EC information consistently improves reaction prediction, while learned reaction representations support accurate EC prediction, revealing a bidirectional relationship between reaction structure and catalytic function. Together, VenusRX-Bench and VenusRX provide a unified framework for elucidating and modeling enzymatic reaction space