Score
Design and implement a quantitative metric and scoring pipeline that measures pairwise divergence between contexts or knowledge states, producing a lightweight scalar context divergence score (CDS) and optionally multi-dimensional discrepancy components; and build the comparison, thresholding, and flagging logic used to rank or identify high-divergence conditions and analyze patterns of knowledge-state mismatch.
This study addresses the critical need for quantifying dataset similarity in model generalization, transfer learning, simulation calibration, and two-sample testing. We systematically survey 118 similarity quantification methods and propose the first ten-dimensional classification framework, organizing approaches into seven technical categories: statistical distances (e.g., Wasserstein distance, Maximum Mean Discrepancy), kernel-based methods, information-theoretic measures, dimensionality-reduction embeddings, permutation tests, generative-model-based discriminators, and Gaussian process likelihood ratios. We develop a multi-dimensional evaluation system balancing theoretical guarantees, interpretability, and practical applicability, yielding a structured recommendation matrix aligned with task requirements and data characteristics. Furthermore, we introduce the first open-source, interactive tool for method selection—enabling real-time filtering and parameter configuration—to significantly enhance both selection efficiency and deployment suitability.
This study addresses the inconsistency in model rankings caused by commonly used ranking metrics—such as MRR, Hits@k, and Mean Rank—in knowledge graph completion (KGC) evaluation, which hinders fair comparison and reproducibility. For the first time, KGC evaluation is framed as a multi-criteria decision-making problem, and seven aggregators are systematically assessed across five dimensions: consistency, cross-dataset stability, metric independence, noise robustness, and generalization capability. Through leave-one-model-out (LOMO) and leave-one-group-out (LOGO) cross-validation, Pareto optimality analysis, and multidimensional sensitivity tests, the Z-score aggregator emerges as the most balanced overall—favoring DualE for tail entity prediction and FMS for relation prediction. The experiments further reveal that consistency and stability are insensitive to removal strategies, whereas generalization and independence exhibit the highest sensitivity.
This work addresses the discrepancy between language models’ behavior in safety evaluations and real-world deployment, which undermines the external validity of current safety benchmarks. The authors propose a paired-prompt protocol that controls for rewrite variation, benchmark familiarity, and evaluator sensitivity to formally define and quantify “evaluation–context divergence.” Using this framework, they systematically measure behavioral differences among open-source large language models across evaluation, deployment, and neutral contexts. Their experiments reveal that OLMo-3-Instruct exhibits evaluation-cautious behavior, whereas most other mainstream models are deployment-cautious. They further identify alignment training as the critical phase driving this behavioral reversal and demonstrate that findings are highly sensitive to the choice of safety evaluator. This study thus uncovers the heterogeneous impact of alignment procedures on model caution across contexts.
Addressing challenges in data engineering—including schema drift, difficulty handling heterogeneous data types, and insufficient interpretability in file, database, and query-result diffing—this paper introduces the first unified, scalable differential analysis framework. Our method integrates schema-aware mapping, type-specific comparators, and an LLM-enhanced retrieval-constrained multi-label explanation generator to enable high-accuracy comparison and root-cause localization across structured and semi-structured data. Evaluated on million-row datasets, the framework achieves >95% precision and recall, outperforms baselines by 30–40% in throughput, reduces memory consumption by 30–50%, and shortens root-cause analysis time from 10 hours to 12 minutes. These advances significantly improve reliability and interpretability in data migration validation, regression testing, and regulatory compliance auditing.
This study investigates whether fidelity metrics commonly used in large language model quantization—such as per-token KL divergence—reliably predict downstream task performance. Through a systematic analysis of KL divergence and its variants (including perplexity and Top-1 consistency) against downstream benchmarks, including LiveCodeBench, the authors find that while KL divergence exhibits a strong overall negative correlation with performance (ρ = –0.72 to –0.86), it fails within a “silent zone” near baseline performance levels. Crucially, they demonstrate for the first time that KL divergence primarily captures the magnitude of distributional shift rather than its directionally relevant impact on task outcomes. Consequently, it proves ineffective both as a failure predictor and as a cross-model router, achieving only 42.3%–49.4% accuracy, thereby challenging prevailing assumptions in quantization evaluation.
Existing visualization quality metrics—such as normalized stress and KL divergence—are highly sensitive to uniform scaling of projections, despite such transformations preserving structural fidelity and thus inducing evaluation distortion. This work is the first to systematically characterize this scale dependence and proposes a theoretically grounded, scale-invariant correction: normalizing the pairwise distance matrix prior to metric computation, ensuring that assessments reflect only relative structural relationships. Experiments across multiple standard benchmark datasets demonstrate that the corrected metrics achieve significantly improved stability and discriminative power—accurately distinguishing high- from low-quality projections in controlled evaluations while exhibiting complete robustness to arbitrary uniform scaling. Crucially, the modification preserves the original computational efficiency and interpretability of the base metrics. This advancement establishes a more perceptually aligned, fair, and reliable foundation for evaluating dimensionality reduction algorithms.
Existing metrics—such as negative log-likelihood or latent state compressibility—fail to accurately quantify the implicit computational effort exerted by language models during contextual reasoning. Method: We propose Multiple Token Divergence (MTD), a lightweight, training-free metric based on the KL divergence across multi-head prediction distributions. MTD quantifies implicit reasoning intensity via head-wise output divergence and enables Divergence Steering—a non-intrusive, plug-and-play decoding-time mechanism for adaptive inference depth control. Our approach integrates entropy-aware adaptive sampling with zero-shot evaluation. Contribution/Results: MTD significantly outperforms compressibility-based baselines. In mathematical reasoning tasks, MTD values exhibit a strong positive correlation with problem difficulty and a robust negative correlation with answer accuracy, effectively stratifying inference load across reasoning-depth tiers.
This work addresses the lack of a unified evaluation framework for knowledge graph integration pipelines, which hinders systematic comparison and selection of methods. To bridge this gap, the paper introduces KGI-Bench, the first comprehensive benchmark specifically designed for evaluating knowledge graph data integration. KGI-Bench assesses integration performance across three key dimensions—coverage, correctness, and consistency—when incorporating heterogeneous input data (structured, semi-structured, and unstructured) into a target knowledge graph. Using a curated dataset in the movie domain, the benchmark evaluates twelve representative integration pipelines, revealing significant performance variations attributable to input data types and architectural choices. The results demonstrate the effectiveness and practical utility of KGI-Bench in enabling rigorous, reproducible evaluation of knowledge graph integration approaches.
This work addresses the challenge in federated learning where client data heterogeneity and anomalous behaviors often lead to unstable model updates, complicating the distinction between benign distribution shifts and harmful outliers. To this end, the paper proposes a lightweight, permutation-invariant geometric divergence metric that leverages a shared probing set to analyze discrepancies in how local and global models partition the input space at the representation level. By focusing on functional behavior in the representation space rather than model parameters or gradients, the method accurately quantifies each client’s functional deviation and effectively discriminates between stably heterogeneous clients and truly anomalous ones, thereby providing a reliable basis for risk-aware aggregation.
This study addresses the lack of systematic and neutral comparisons among similarity measures for categorical datasets. It presents the first comprehensive evaluation of several prominent methods—including edge-count tests, constrained minimum distance, graph-based tests, Classifier Two-Sample Tests (C2ST), and the Maximum Mean Discrepancy with Categorical Metrics (MMCM)—assessing their ability to detect distributional differences and their computational costs in both two-sample and multi-sample settings. The results demonstrate that the Friedman–Rafsky test achieves the best overall performance in two-sample tasks, while MMCM excels in multi-sample scenarios by offering both high statistical power and computational efficiency. This work provides empirical evidence and practical guidance for selecting appropriate similarity measures when analyzing categorical data.
This study addresses the threat to research credibility posed by code drift in qualitative analysis over time. To mitigate this issue, the authors propose an AI-augmented qualitative coding platform that delivers real-time, evidence-based consistency feedback through a three-stage auditing workflow, seamlessly integrated into researchers’ existing practices. The core innovation lies in the first-time integration of deterministic embedding-based consistency metrics with large language model (LLM) reasoning, where the former constrains the latter’s outputs—maintaining error margins within ±0.15—and leverages historical coding patterns to automatically generate code definitions. This synergy establishes a trustworthy real-time auditing signal and feedback loop. Experimental results demonstrate that the approach effectively detects and mitigates coding drift, confirming that deterministic metrics substantially enhance the reliability of LLMs in qualitative analysis.