🤖 AI Summary
This study addresses the limitations of large language models in deep contextual understanding and factual consistency during question answering, as well as the inadequacy of conventional evaluation metrics. To tackle these challenges, this work proposes a knowledge graph-based fine-grained evaluation framework. The core innovation lies in the design of S3KG, a hybrid similarity metric that integrates semantic and structural information to assess deep logical consistency, enabling precise attribution and diagnosis of reasoning errors at the triplet level. Extensive experiments across nine benchmarks demonstrate that the proposed method improves F1 scores by up to 7.6 points and achieves an AUROC of 0.973, significantly outperforming existing evaluation approaches.
📝 Abstract
While large language models (LLMs) have achieved remarkable linguistic capabilities, a profound question lingers at their core: do these models truly comprehend context or simply excel at pattern matching on an unprecedented scale? Contextual understanding in LLMs refers to the ability to correctly extract relevant information from a given context, integrate it into a coherent internal representation, and reason over it to produce factually consistent and contextually grounded responses. However, traditional methods such as BiLingual Evaluation Understudy (BLEU) and perplexity simply measure surface-level performance. This reveals a critical gap in question answering (QA), where responses must be contextually grounded rather than simply being memorized associations. To fill this void, we propose a novel knowledge graph (KG) based evaluation framework for LLM contextual understanding in QA. Central to this is Semantic Structural Similarity for KGs (S3KG), a hybrid similarity measure combining structural and semantic signals into a single score. In addition, a diagnostic analysis framework is developed to identify and categorize reasoning errors at the triplet level, enabling fine-grained analysis of model failures. Together, across nine benchmarks, S3KG achieves F1 gains of up to $+7.6$ points over the strongest baseline and AUROC up to $0.973$.