context divergence scoring

Design and implement a quantitative metric and scoring pipeline that measures pairwise divergence between contexts or knowledge states, producing a lightweight scalar context divergence score (CDS) and optionally multi-dimensional discrepancy components; and build the comparison, thresholding, and flagging logic used to rank or identify high-divergence conditions and analyze patterns of knowledge-state mismatch.

contextdivergencescoring

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.31
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the inconsistency in model rankings caused by commonly used ranking metrics—such as MRR, Hits@k, and Mean Rank—in knowledge graph completion (KGC) evaluation, which hinders fair comparison and reproducibility. For the first time, KGC evaluation is framed as a multi-criteria decision-making problem, and seven aggregators are systematically assessed across five dimensions: consistency, cross-dataset stability, metric independence, noise robustness, and generalization capability. Through leave-one-model-out (LOMO) and leave-one-group-out (LOGO) cross-validation, Pareto optimality analysis, and multidimensional sensitivity tests, the Z-score aggregator emerges as the most balanced overall—favoring DualE for tail entity prediction and FMS for relation prediction. The experiments further reveal that consistency and stability are insensitive to removal strategies, whereas generalization and independence exhibit the highest sensitivity.

Evaluation MetricsKnowledge Graph CompletionMetric Disagreement

This work addresses the discrepancy between language models’ behavior in safety evaluations and real-world deployment, which undermines the external validity of current safety benchmarks. The authors propose a paired-prompt protocol that controls for rewrite variation, benchmark familiarity, and evaluator sensitivity to formally define and quantify “evaluation–context divergence.” Using this framework, they systematically measure behavioral differences among open-source large language models across evaluation, deployment, and neutral contexts. Their experiments reveal that OLMo-3-Instruct exhibits evaluation-cautious behavior, whereas most other mainstream models are deployment-cautious. They further identify alignment training as the critical phase driving this behavioral reversal and demonstrate that findings are highly sensitive to the choice of safety evaluator. This study thus uncovers the heterogeneous impact of alignment procedures on model caution across contexts.

contextual framingdeployment alignmentevaluation-context divergence

Addressing challenges in data engineering—including schema drift, difficulty handling heterogeneous data types, and insufficient interpretability in file, database, and query-result diffing—this paper introduces the first unified, scalable differential analysis framework. Our method integrates schema-aware mapping, type-specific comparators, and an LLM-enhanced retrieval-constrained multi-label explanation generator to enable high-accuracy comparison and root-cause localization across structured and semi-structured data. Evaluated on million-row datasets, the framework achieves >95% precision and recall, outperforms baselines by 30–40% in throughput, reduces memory consumption by 30–50%, and shortens root-cause analysis time from 10 hours to 12 minutes. These advances significantly improve reliability and interpretability in data migration validation, regression testing, and regulatory compliance auditing.

Handling schema drift and heterogeneous data typesImproving efficiency and accuracy in data comparisonProviding explainable differences in large-scale data

This study investigates whether fidelity metrics commonly used in large language model quantization—such as per-token KL divergence—reliably predict downstream task performance. Through a systematic analysis of KL divergence and its variants (including perplexity and Top-1 consistency) against downstream benchmarks, including LiveCodeBench, the authors find that while KL divergence exhibits a strong overall negative correlation with performance (ρ = –0.72 to –0.86), it fails within a “silent zone” near baseline performance levels. Crucially, they demonstrate for the first time that KL divergence primarily captures the magnitude of distributional shift rather than its directionally relevant impact on task outcomes. Consequently, it proves ineffective both as a failure predictor and as a cross-model router, achieving only 42.3%–49.4% accuracy, thereby challenging prevailing assumptions in quantization evaluation.

benchmark correlationfidelity metricsKL divergence

How Scale Breaks "Normalized Stress" and KL Divergence: Rethinking Quality Metrics

Oct 09, 2025
KS
Kiran Smelser
🏛️ University of Arizona | Technical University of Munich

Existing visualization quality metrics—such as normalized stress and KL divergence—are highly sensitive to uniform scaling of projections, despite such transformations preserving structural fidelity and thus inducing evaluation distortion. This work is the first to systematically characterize this scale dependence and proposes a theoretically grounded, scale-invariant correction: normalizing the pairwise distance matrix prior to metric computation, ensuring that assessments reflect only relative structural relationships. Experiments across multiple standard benchmark datasets demonstrate that the corrected metrics achieve significantly improved stability and discriminative power—accurately distinguishing high- from low-quality projections in controlled evaluations while exhibiting complete robustness to arbitrary uniform scaling. Crucially, the modification preserves the original computational efficiency and interpretability of the base metrics. This advancement establishes a more perceptually aligned, fair, and reliable foundation for evaluating dimensionality reduction algorithms.

KL divergence in t-SNE suffers from scale sensitivity issuesNormalized stress metric is sensitive to uniform scalingQuality metrics need scale-invariance for accurate projection evaluation

Latest Papers

What's happening recently
View more

Multiple Token Divergence: Measuring and Steering In-Context Computation Density

Dec 28, 2025
VH
Vincent Herrmann
🏛️ The Swiss AI Lab IDSIA | USI | SUPSI

Existing metrics—such as negative log-likelihood or latent state compressibility—fail to accurately quantify the implicit computational effort exerted by language models during contextual reasoning. Method: We propose Multiple Token Divergence (MTD), a lightweight, training-free metric based on the KL divergence across multi-head prediction distributions. MTD quantifies implicit reasoning intensity via head-wise output divergence and enables Divergence Steering—a non-intrusive, plug-and-play decoding-time mechanism for adaptive inference depth control. Our approach integrates entropy-aware adaptive sampling with zero-shot evaluation. Contribution/Results: MTD significantly outperforms compressibility-based baselines. In mathematical reasoning tasks, MTD values exhibit a strong positive correlation with problem difficulty and a robust negative correlation with answer accuracy, effectively stratifying inference load across reasoning-depth tiers.

Introducing decoding to control computational character of textMeasuring in-context computational effort of language modelsProposing a lightweight method to analyze computational dynamics

This work addresses the lack of a unified evaluation framework for knowledge graph integration pipelines, which hinders systematic comparison and selection of methods. To bridge this gap, the paper introduces KGI-Bench, the first comprehensive benchmark specifically designed for evaluating knowledge graph data integration. KGI-Bench assesses integration performance across three key dimensions—coverage, correctness, and consistency—when incorporating heterogeneous input data (structured, semi-structured, and unstructured) into a target knowledge graph. Using a curated dataset in the movie domain, the benchmark evaluates twelve representative integration pipelines, revealing significant performance variations attributable to input data types and architectural choices. The results demonstrate the effectiveness and practical utility of KGI-Bench in enabling rigorous, reproducible evaluation of knowledge graph integration approaches.

benchmarkdata integrationknowledge graph

This work addresses the challenge in federated learning where client data heterogeneity and anomalous behaviors often lead to unstable model updates, complicating the distinction between benign distribution shifts and harmful outliers. To this end, the paper proposes a lightweight, permutation-invariant geometric divergence metric that leverages a shared probing set to analyze discrepancies in how local and global models partition the input space at the representation level. By focusing on functional behavior in the representation space rather than model parameters or gradients, the method accurately quantifies each client’s functional deviation and effectively discriminates between stably heterogeneous clients and truly anomalous ones, thereby providing a reliable basis for risk-aware aggregation.

Anomalous InputsAtypical ClientsClient Heterogeneity

This study addresses the lack of systematic and neutral comparisons among similarity measures for categorical datasets. It presents the first comprehensive evaluation of several prominent methods—including edge-count tests, constrained minimum distance, graph-based tests, Classifier Two-Sample Tests (C2ST), and the Maximum Mean Discrepancy with Categorical Metrics (MMCM)—assessing their ability to detect distributional differences and their computational costs in both two-sample and multi-sample settings. The results demonstrate that the Friedman–Rafsky test achieves the best overall performance in two-sample tasks, while MMCM excels in multi-sample scenarios by offering both high statistical power and computational efficiency. This work provides empirical evidence and practical guidance for selecting appropriate similarity measures when analyzing categorical data.

categorical datadataset similaritymulti-sample comparison

This study addresses the threat to research credibility posed by code drift in qualitative analysis over time. To mitigate this issue, the authors propose an AI-augmented qualitative coding platform that delivers real-time, evidence-based consistency feedback through a three-stage auditing workflow, seamlessly integrated into researchers’ existing practices. The core innovation lies in the first-time integration of deterministic embedding-based consistency metrics with large language model (LLM) reasoning, where the former constrains the latter’s outputs—maintaining error margins within ±0.15—and leverages historical coding patterns to automatically generate code definitions. This synergy establishes a trustworthy real-time auditing signal and feedback loop. Experimental results demonstrate that the approach effectively detects and mitigates coding drift, confirming that deterministic metrics substantially enhance the reliability of LLMs in qualitative analysis.

CAQDAScoding consistencyqualitative coding

Hot Scholars

OW

Oliver Weißl

TUM, fortiss GmbH
GenAIEvolutionary RoboticsComputer VisionDeep Learning
AS

Andrea Stocco

Technical University of Munich
Software EngineeringSoftware TestingTest AutomationDeep Learning Testing
CZ

Chengzhi Zhang

Nanjing University of Science and Technology
Text MiningNatural Language ProcessingScience of Science