Do LLMs Understand Context? A Knowledge Graph-Based Evaluation Framework

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of large language models in deep contextual understanding and factual consistency during question answering, as well as the inadequacy of conventional evaluation metrics. To tackle these challenges, this work proposes a knowledge graph-based fine-grained evaluation framework. The core innovation lies in the design of S3KG, a hybrid similarity metric that integrates semantic and structural information to assess deep logical consistency, enabling precise attribution and diagnosis of reasoning errors at the triplet level. Extensive experiments across nine benchmarks demonstrate that the proposed method improves F1 scores by up to 7.6 points and achieves an AUROC of 0.973, significantly outperforming existing evaluation approaches.
📝 Abstract
While large language models (LLMs) have achieved remarkable linguistic capabilities, a profound question lingers at their core: do these models truly comprehend context or simply excel at pattern matching on an unprecedented scale? Contextual understanding in LLMs refers to the ability to correctly extract relevant information from a given context, integrate it into a coherent internal representation, and reason over it to produce factually consistent and contextually grounded responses. However, traditional methods such as BiLingual Evaluation Understudy (BLEU) and perplexity simply measure surface-level performance. This reveals a critical gap in question answering (QA), where responses must be contextually grounded rather than simply being memorized associations. To fill this void, we propose a novel knowledge graph (KG) based evaluation framework for LLM contextual understanding in QA. Central to this is Semantic Structural Similarity for KGs (S3KG), a hybrid similarity measure combining structural and semantic signals into a single score. In addition, a diagnostic analysis framework is developed to identify and categorize reasoning errors at the triplet level, enabling fine-grained analysis of model failures. Together, across nine benchmarks, S3KG achieves F1 gains of up to $+7.6$ points over the strongest baseline and AUROC up to $0.973$.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Contextual Understanding
Evaluation Framework
Question Answering
Knowledge Graph
Innovation

Methods, ideas, or system contributions that make the work stand out.

Knowledge Graph
Contextual Understanding
Evaluation Framework
S3KG
Diagnostic Analysis
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Subavarshana Arumugam
Department of Computer Science & Engineering, University of Moratuwa, Sri Lanka
M
Mamta Nallaretnam
Department of Computer Science & Engineering, University of Moratuwa, Sri Lanka
K
Kithuni Wickramasinghe
Department of Computer Science & Engineering, University of Moratuwa, Sri Lanka
C
Chamath Gunapala
Department of Computer Science & Engineering, University of Moratuwa, Sri Lanka
P
Pragatheeswaran Vipulanandan
Department of Electrical and Computer Engineering, University of Miami, USA
Kamal Premaratne
Kamal Premaratne
Professor, Electrical and Computer Engineering, University of Miami, Coral Gables, Florida, USA
graph theoretic methodsknowledge discovery from uncertain datainterval-valued probabilitiesDempster-Shafer (DS) theory
Uthayasanker Thayasivam
Uthayasanker Thayasivam
Senior Lecturer Department of Computer Science and Engineering, University of Moratuwa
nlpmldata science