A standard transformer and attention with linear biases for molecular conformer generation

📅 2025-06-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Non-equivariant models for molecular conformation generation suffer from weak spatial modeling capability and rely heavily on large parameter counts. Method: This paper proposes a lightweight, efficient standard Transformer architecture featuring (i) shortest-path distance–based relative positional encoding between atoms, and (ii) an ALiBi-inspired linear attention bias mechanism that assigns distinct slopes to individual attention heads, explicitly encoding 3D spatial relationships and mitigating the lack of geometric equivariance. Contribution/Results: On the GEOM-DRUGS benchmark, our model achieves state-of-the-art performance with only 25 million parameters—significantly outperforming the prior best non-equivariant model (64 million parameters). This demonstrates superior generalization and parameter efficiency, establishing a new paradigm for conformation prediction under resource-constrained settings.

Technology Category

Machine Learning: Deep Neural Architectures and Foundation ModelsComputer Vision: Diffusion Models for VisionKnowledge Representation and Reasoning: Geometric, Spatial, and Temporal Reasoning

Application Category

Graph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsUser Modeling, Personalization and Recommendation: Attacks and countermeasures in recommendation systemsSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
Sampling low-energy molecular conformations, spatial arrangements of atoms in a molecule, is a critical task for many different calculations performed in the drug discovery and optimization process. Numerous specialized equivariant networks have been designed to generate molecular conformations from 2D molecular graphs. Recently, non-equivariant transformer models have emerged as a viable alternative due to their capability to scale to improve generalization. However, the concern has been that non-equivariant models require a large model size to compensate the lack of equivariant bias. In this paper, we demonstrate that a well-chosen positional encoding effectively addresses these size limitations. A standard transformer model incorporating relative positional encoding for molecular graphs when scaled to 25 million parameters surpasses the current state-of-the-art non-equivariant base model with 64 million parameters on the GEOM-DRUGS benchmark. We implemented relative positional encoding as a negative attention bias that linearly increases with the shortest path distances between graph nodes at varying slopes for different attention heads, similar to ALiBi, a widely adopted relative positional encoding technique in the NLP domain. This architecture has the potential to serve as a foundation for a novel class of generative models for molecular conformations.
Problem

Research questions and friction points this paper is trying to address.

Generating low-energy molecular conformations efficiently
Overcoming size limitations in non-equivariant transformer models
Improving molecular conformation accuracy with positional encoding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Standard transformer with linear biases
Relative positional encoding for graphs
Scalable non-equivariant model outperforms equivariant
💼 Related Jobs
No related jobs found.
IBM Research
Viatcheslav Gurev
Viatcheslav Gurev
IBM
T
Timothy Rumbell
IBM Research