🤖 AI Summary
This work addresses the unreliable reliance on task-critical evidence in multimodal large language models (MLLMs) for power grid diagnosis, despite their ability to integrate topological, measurement, and textual data. To enhance trustworthiness, the study proposes the first MLLM audit framework tailored for grid diagnostics, which detects evidence usage biases through pre-registered engineering importance specifications, self-reported dependency analysis, and modality ablation interventions. The framework further incorporates evidence-gated generation and an independent re-audit mechanism to correct identified biases and establish a verifiable reasoning loop. Experiments on IEEE 39- and 118-bus systems demonstrate that the approach effectively identifies and rectifies task-conditioned credibility failures across MLLMs of varying scales without compromising diagnostic accuracy.
📝 Abstract
Multimodal large language models (LLMs) can combine topology, measurements, and incident text for grid diagnosis, yet answer accuracy does not establish that task-appropriate evidence was used. This letter proposes a general framework in order to conduct task-conditional faithfulness audit. It compares self-reported reliance, intervention-derived behavioral reliance, and preregistered engineering importance. The framework first registers task-specific evidence requirements and compares them with self-reported reliance and behavioral changes under controlled modality ablations. To resolve detected discrepancies, we design an evidence-gated correction and re-audit mechanism that regenerates failed responses under evidence constraints and independently re-ablates them to verify improved grounding without performance loss. Case studies evaluate three differently scaled LLMs on IEEE 39- and 118-bus scenarios. These results validate the framework ability to detect, diagnose, and correct task-conditional faithfulness failures.