Using large language models to probe the limits of atom-centered structural descriptors

📅 2026-07-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses a fundamental limitation in existing atom-centered structural descriptors: their susceptibility to degeneracy, wherein distinct three-dimensional atomic configurations—despite lacking symmetry equivalence—yield identical descriptors due to insufficient neighborhood resolution. For the first time, large language models are leveraged to mine and reason across fragmented scientific knowledge, enabling the construction of structural degeneracy examples that remain indistinguishable even when descriptors are extended to seventh-order neighborhoods. This work not only establishes the most challenging degeneracy cases reported to date but also demonstrates the pivotal role of large language models in synthesizing disparate scientific insights to drive interdisciplinary discovery.
📝 Abstract
Mapping an atomic structure to a compact set of geometric descriptors is an essential step in any machine-learning application to atomic-scale modeling. A powerful and widely-used approach can be understood as a discretization of the histogram of pair distances, triangles, etc., that results in a hierarchy of symmetry-invariant atom-centered descriptors. Unfortunately, the lower rungs on this hierarchy (two, three, four-neighbor clusters) were found to be incomplete, with symmetry-unrelated pairs of structures having exactly the same descriptors. However, all the ``descriptor degeneracies'' reported so far are resolved by considering larger clusters of neighbors to build the descriptors. We report examples of 3D structures that are indistinguishable even if one considers clusters of up to seven neighbors, and to arbitrary order when considering a practical level of discretization of the descriptors, discovered with the assistance of large language models. The key ingredients in their construction can be traced to results that have been known for decades in different communities; the model was able to find the references and recognize their significance for the problem at hand. We believe this experiment exposes an extremely fruitful usage pattern for AI in science: translating results between different communities and application domains, accelerating the process by which serendipitous discoveries in a field become paradigm-shifting breakthroughs in another.
Problem

Research questions and friction points this paper is trying to address.

structural descriptors
descriptor degeneracy
symmetry-invariant
atomic-scale modeling
machine learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

large language models
atom-centered descriptors
descriptor degeneracy
symmetry-invariant representations
cross-disciplinary knowledge transfer