Score
Designs and applies diagnostic methods and metrics that identify when a model’s representations or outputs have lost semantic distinctions—producing consistently incorrect interpretations, ghost errors, or irrecoverable corruption. Builds detectors, evaluation protocols, and analyses to quantify the prevalence and severity of such semantic collapse and to measure its impact on benchmarks and evaluation metrics.
Current machine learning evaluation practices predominantly rely on surface-level performance metrics, often neglecting the internal mechanisms of models. This work proposes trustworthy interpretability as a central evaluation paradigm and, for the first time, systematically demonstrates that it satisfies core criteria from the philosophy of science—namely falsifiability, reproducibility, and predictive power. By constructing an evaluation framework that integrates causal analysis with mechanistic probing, the study delineates three functional pathways through which interpretability enables the identification of behavioral origins, detection of latent flaws, and prediction of potential failure modes. This approach advances model assessment beyond performance-oriented benchmarks toward a deeper understanding of underlying mechanisms.
This study addresses the fundamental question of whether semantic similarity measures genuinely comprehend semantic relationships. We propose the first evaluation framework based on controlled, small-scale semantic transformations to systematically assess the semantic discrimination capability of 18 state-of-the-art methods—including bag-of-words, embedding-based, LLM-based, and structure-aware models—on software engineering texts and code. Experiments reveal that mainstream embedding methods exhibit up to 99.9% misclassification rates in semantic opposition scenarios, exposing their reliance on superficial surface patterns. Substituting cosine similarity for Euclidean distance improves performance by 24–66%. LLM-based methods demonstrate superior fine-grained semantic distinction. Critically, our framework uncovers a foundational limitation in existing measures: their failure to capture semantic essence. It establishes the first reproducible, scalable benchmark paradigm for trustworthy semantic computation in software engineering contexts.
This work identifies “model collapse”—a phenomenon wherein models trained on unfiltered text embeddings (TEs) degenerate to predicting a single class, yielding spurious high accuracy and false correlations in downstream tasks. To quantify collapse severity, we propose the first dedicated metric and demonstrate that TE quality serves as an effective proxy for data cleaning. Through controlled experiments comparing model performance on original tabular data versus their TE representations, we show that uncurated TEs consistently induce collapse, severely impairing out-of-distribution generalization. Our study is the first to systematically reveal the critical impact of TE quality on learning robustness. It establishes a reproducible evaluation framework and provides principled data filtering criteria for embedding-driven modeling—bridging a key gap between embedding usage and reliable machine learning practice.
This study addresses the challenge of “silent failures” in AI systems—errors that occur during training or deployment without manifesting as anomalies in loss functions, thereby evading conventional evaluation mechanisms. The work introduces the concept of “evaluation blind spots” and presents the first unified framework modeling silent failures across both training and deployment phases. It formally defines detectability predicates, establishes a six-category taxonomy of such failures, and proposes a risk-based failure budget framework. Through retrospective analysis of 50 real-world incidents, case studies, formal verification, and an open-source implementation featuring gradient tracing, log auditing, and contamination detection, the study reveals that 53% of publicly reported AI failures are silent in nature, identifies genuine gradient errors in the TRL library, and demonstrates that one of the six production failure categories is structurally silent in 100% of observed cases.
This work addresses the challenge of behavioral inconsistency in automatically generated BPMN models due to semantic ambiguity in natural language process descriptions. It proposes the first closed-loop diagnosis and repair framework that operates without requiring ground-truth BPMN annotations. By analyzing the distribution of key performance indicators (KPIs) across multiple model generations, the approach identifies behavioral variations and employs model-based diagnostic techniques to pinpoint gateway logic discrepancies. These discrepancies are traced back to specific source text fragments, which are then refined through an evidence-driven textual revision process. Evaluated on clinical guidelines for diabetic kidney disease management, the method significantly reduces behavioral variability in regenerated models and enhances the semantic stability of executable process models, establishing an end-to-end mapping from behavioral inconsistency to targeted textual correction.
This study investigates the “blind compliance” behavior of code large language models (LLMs), which execute erroneous instructions even when they recognize the errors, leading to irreversible “ghost errors” and semantic collapse. Through four sets of experiments on algorithmic Python problems from the RunBugRun dataset with deterministic test cases, the work systematically evaluates model responses in both single-attempt and multi-round repair scenarios. It formally defines and reveals the blind compliance phenomenon, quantifies the proportion of irrecoverable semantic failures, and demonstrates that extended reasoning fails to effectively mitigate such errors. The findings expose critical limitations of conventional pass-rate metrics and raise significant concerns about the reliability of code LLMs in production environments.
This study addresses the fundamental challenge posed by AI-generated code to the long-standing assumption in software engineering that authorship implies understanding. The authors demonstrate, for the first time, that AI-assisted programming systematically undermines the validity of authorship as a proxy for knowledge, thereby rendering traditional knowledge metrics—such as the truck factor—ineffective. By integrating principles from software engineering measurement theory, knowledge modeling, and logical reasoning, the work reinterprets the semantics of version control data in the context of AI collaboration. The research reveals that existing measures of knowledge concentration no longer reflect actual comprehension in AI-augmented development environments and advocates for a paradigm shift toward metrics grounded in verifiable evidence of understanding. It further identifies the construction of system-level understanding metrics as a critical open problem for the field.
This study addresses the frequent operationalization failures that arise when large language models generate analytical workflows, stemming from a semantic gap between user intent and system-executable actions. Through cross-domain empirical analysis across finance, human resources, and public safety, the authors manually examined 236 analytical intents and their automatically generated workflows, systematically identifying and categorizing five distinct semantic-level failure patterns: comparative anchoring, procedural reasoning, quantitative reasoning, role confusion, and policy anchoring. The findings reveal fundamental limitations in the semantic expressiveness of current data systems and provide both theoretical grounding and practical guidance for improving the reliability of agent-generated analytical workflows.
This work addresses a common yet critical issue in machine learning code: semantic errors arising from mismatches between data properties and model assumptions—such as applying scale-sensitive models to unnormalized data—which traditional debugging approaches can only detect after training, resulting in inefficiency. To enable early and automatic error detection, the authors propose a novel data-aware static analysis method that integrates dataflow and control-flow analysis with API specifications, thereby incorporating data semantics directly into the static analysis framework for the first time. Evaluation on real-world machine learning notebooks demonstrates that the approach effectively identifies subtle semantic bugs that conventional techniques fail to catch, highlighting its practical utility and methodological innovation.
Current automated formalization evaluations lack interpretable diagnostics for semantic errors, hindering both system optimization and human understanding. This work proposes FormalRx, a novel framework that introduces the first fine-grained, hierarchical taxonomy of Semantic Correctness Issues (SCI) comprising 28 error categories, and develops an end-to-end diagnostic model, FormalRx-8B, capable of aligning, classifying, localizing, and correcting formalization errors. Trained on 56,287 fine-grained annotated samples, the model achieves strong performance across four diagnostic tasks—0.88 F1 for alignment, 0.71 F1 for classification, 0.75 accuracy for localization, and 0.73 accuracy for correction—significantly outperforming both general-purpose large language models and specialized baselines. The study also releases FormalRx-Test, the first fine-grained diagnostic benchmark, thereby establishing a closed-loop pipeline from opaque evaluation to actionable feedback.