🤖 AI Summary
Current AI-generated hypotheses in materials science, while linguistically fluent, lack verifiable scientific grounding in their internal reasoning mechanisms. This work proposes a visual diagnostic framework integrating semantic backtracking, graph-structure perturbation, and activation recovery metrics to scrutinize the “graph-to-answer” hypothesis generation pathway—the first such approach to focus mechanism interpretability on this specific process. Leveraging the Graph-PrefLexOR-8B model, residual stream scanning and layer–token grid visualizations across 100 open-ended materials science questions reveal that mechanistic recovery is highly concentrated in late-stage synthesis layers—particularly around layers 30 and 36—and that the final answers align most closely with the model’s own synthesized representations, thereby identifying critical loci where graph-based reasoning mechanisms are predominantly encoded.
📝 Abstract
AI co-scientists can generate fluent materials-science hypotheses, but fluency does not show that an answer preserves a scientifically meaningful mechanism. We present a graph-to-answer mechanism-tracing case study for Graph-PRefLexOR-8B, a Qwen3-8B model adapted to expose distinct stages for brainstorming, graph construction, pattern extraction, and synthesis. We organize semantic backtracking, graph corruption, activation-based recovery measurements, and layer-by-token-region grids into a visual diagnostic workflow for inspecting this pathway. Across 100 open-ended materials-science questions, final answers remain closest to the model's own structured stages, especially synthesis. Under graph corruption, a full sweep over 37 residual-stream checkpoints, the embedding output and 36 transformer blocks, shows little mechanism recovery in the earlier transition region at layers 7--10, recovery instead concentrates in late synthesis and answer-start regions around layers 30 and 36. The workflow is intended to help scientists and model developers identify where a generated hypothesis loses or regains mechanism support before it is passed to downstream experimental planning.