🤖 AI Summary
This study addresses the limited mechanistic reasoning capabilities of scientific language models caused by text fragmentation. We propose MS³ (Material–Sensor–Signal–System), an architecture that constructs an evidence-driven mechanistic knowledge base. Leveraging knowledge graphs and evidence unit extraction, our approach maps unstructured literature into structured mechanistic graphs comprising traceable evidence, role-based entities, and directed paths, supported by a source-repair workflow to ensure data consistency. A multi-model evaluation across 13,689 papers demonstrates that this method significantly outperforms conventional retrieval-augmented generation (RAG) baselines in scientific correctness, citation entailment, and answer completeness.
📝 Abstract
Scientific language models often access literature through untyped text chunks, which fragment the functional and evidential structure required for mechanism-rich questions. We introduce an evidence-grounded mechanism knowledge substrate that organizes scientific literature into provenance-linked evidence units, role-typed entities, and directed mechanism paths. We instantiate it as MS$^3$, a Material-Sensor-Signal-System schema for conductive-fiber flexible sensors, over 13,689 papers, 131,083 evidence items, and 26,648 mechanism objects. On in-domain and coverage-shift question-answering benchmarks, we compare closed-book generation, Web search, Raw-PDF RAG, and MS$^3$ retrieval across ten language models. MS$^3$ improves macro-averaged scientific correctness. It also improves citation entailment and answer completeness. These results support mechanism substrates as a reliable representation layer for scientific language models and motivate a source-repair workflow in which insufficient MS$^3$ evidence triggers targeted retrieval from its linked papers rather than assuming that a user has already supplied the correct PDFs.