R-GroundBench: A Diagnostic Benchmark for R-Group Groundingin Markush Molecular Editing

πŸ“… 2026-09-30
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limitation of existing benchmarks in evaluating AI reasoning for patent molecular editing, particularly regarding R-group grounding in Markush structures. We construct the first diagnostic benchmark targeting incomplete chemical representations derived from real patents. By employing multimodal visual question answering and open-ended generation tasks with a controlled difficulty stratification strategy, this work systematically evaluates models’ cross-modal comprehension and execution capabilities for variable R-groups. Our approach effectively distinguishes superficial recognition from deep reasoning, revealing critical limitations in domain-pretrained models. Experimental results demonstrate that model accuracy degrades significantly in complex scenarios, with extremely low generation success rates, confirming that current AI systems lack reliable Markush editing capabilities.
πŸ“ Abstract
Recent advances in AI for scientific discovery enable molecular understandingand design, yet reasoning over incomplete chemical representations remainsunclear.Markush structures, which encode molecular families through variable R-groupplaceholders (\textit{R\textsubscript{1}}, \textit{R\textsubscript{2}}, \textit{X}, etc.), are ubiquitous in pharmaceutical patents and requiregrounding across molecular, textual, and chemical information.However, existing molecule-language benchmarks focus on fully specifiedmolecules, leaving R-group grounding largely unevaluated.We introduce R-GroundBench:, a diagnostic benchmark built from real patent Markushstructures, featuring a Multiple-Choice (VQA) track with controlled difficultyand modality splits, and an open-ended Generation track.Our results reveal a substantial gap between recognition andmolecular grounding.While models achieve over 90\% accuracy on Easy VQA, performance drops to56--66\% on Hard VQA when shortcuts are controlled.Chemical-domain VLMs also remain unreliable, achieving only 25.7--46.2\% on HardVQA despite domain-specific pretraining.Moreover, Generation Exact Match remains below 20\% for most models and below8\% when visual input is required.These findings reveal that current AI systems lack reliable grounding andexecution for Markush editing, highlighting challenges for AI-drivenscientific discovery.
Problem

Research questions and friction points this paper is trying to address.

Markush structures
R-group grounding
molecular editing
benchmark evaluation
AI for science
Innovation

Methods, ideas, or system contributions that make the work stand out.

Markush structures
R-group grounding
diagnostic benchmark
visual question answering
molecular editing
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
X
Xin Wang
Westlake University
Z
Zichuan Ying
Westlake University; The University of Hong Kong; Shanghai Innovation Institute
X
Xinna Lin
Westlake University; Shanghai Innovation Institute; Zhejiang University
J
Junqi Zhang
Zhejiang University
H
Hanyi Xiong
Westlake University; The University of Hong Kong
Tianyu Gao
Tianyu Gao
Princeton University
Natural Language Processing
H
Hairong Zhang
Shanghai Artificial Intelligence Laboratory
Q
Qixiang Hua
The Hong Kong University of Science and Technology (Guangzhou)
Botian Shi
Botian Shi
Shanghai Artificial Intelligence Laboratory
VLMsDocument UnderstandingAutonomous Driving
Zhenhailong Wang
Zhenhailong Wang
University of Illinois at Urbana-Champaign
Natural Language ProcessingComputer VisionDeep Learning
Kaicheng Yu
Kaicheng Yu
Assistant Professor, Westlake University, PI of Autonomous Intelligence Lab
computer vision3D understandingautonomous perceptionautomatic machine learning