detect semantic collapse

Designs and applies diagnostic methods and metrics that identify when a model’s representations or outputs have lost semantic distinctions—producing consistently incorrect interpretations, ghost errors, or irrecoverable corruption. Builds detectors, evaluation protocols, and analyses to quantify the prevalence and severity of such semantic collapse and to measure its impact on benchmarks and evaluation metrics.

detectsemanticcollapse

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.58
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Current machine learning evaluation practices predominantly rely on surface-level performance metrics, often neglecting the internal mechanisms of models. This work proposes trustworthy interpretability as a central evaluation paradigm and, for the first time, systematically demonstrates that it satisfies core criteria from the philosophy of science—namely falsifiability, reproducibility, and predictive power. By constructing an evaluation framework that integrates causal analysis with mechanistic probing, the study delineates three functional pathways through which interpretability enables the identification of behavioral origins, detection of latent flaws, and prediction of potential failure modes. This approach advances model assessment beyond performance-oriented benchmarks toward a deeper understanding of underlying mechanisms.

behavioral metricsinterpretabilitymachine learning

How Small Transformation Expose the Weakness of Semantic Similarity Measures

Sep 08, 2025
SL
Serge Lionel Nikiema
🏛️ University of Luxembourg

This study addresses the fundamental question of whether semantic similarity measures genuinely comprehend semantic relationships. We propose the first evaluation framework based on controlled, small-scale semantic transformations to systematically assess the semantic discrimination capability of 18 state-of-the-art methods—including bag-of-words, embedding-based, LLM-based, and structure-aware models—on software engineering texts and code. Experiments reveal that mainstream embedding methods exhibit up to 99.9% misclassification rates in semantic opposition scenarios, exposing their reliance on superficial surface patterns. Substituting cosine similarity for Euclidean distance improves performance by 24–66%. LLM-based methods demonstrate superior fine-grained semantic distinction. Critically, our framework uncovers a foundational limitation in existing measures: their failure to capture semantic essence. It establishes the first reproducible, scalable benchmark paradigm for trustworthy semantic computation in software engineering contexts.

Evaluating semantic similarity measures for software engineering tasksIdentifying flaws where methods confuse opposites and synonymsTesting 18 methods including embeddings and LLMs on semantic understanding

This work identifies “model collapse”—a phenomenon wherein models trained on unfiltered text embeddings (TEs) degenerate to predicting a single class, yielding spurious high accuracy and false correlations in downstream tasks. To quantify collapse severity, we propose the first dedicated metric and demonstrate that TE quality serves as an effective proxy for data cleaning. Through controlled experiments comparing model performance on original tabular data versus their TE representations, we show that uncurated TEs consistently induce collapse, severely impairing out-of-distribution generalization. Our study is the first to systematically reveal the critical impact of TE quality on learning robustness. It establishes a reproducible evaluation framework and provides principled data filtering criteria for embedding-driven modeling—bridging a key gap between embedding usage and reliable machine learning practice.

Model collapse occurs when training on uncurated text embeddingsText embedding quality significantly impacts downstream learning outcomesUncurated text embeddings lead to spurious performance prediction

This study addresses the challenge of “silent failures” in AI systems—errors that occur during training or deployment without manifesting as anomalies in loss functions, thereby evading conventional evaluation mechanisms. The work introduces the concept of “evaluation blind spots” and presents the first unified framework modeling silent failures across both training and deployment phases. It formally defines detectability predicates, establishes a six-category taxonomy of such failures, and proposes a risk-based failure budget framework. Through retrospective analysis of 50 real-world incidents, case studies, formal verification, and an open-source implementation featuring gradient tracing, log auditing, and contamination detection, the study reveals that 53% of publicly reported AI failures are silent in nature, identifies genuine gradient errors in the TRL library, and demonstrates that one of the six production failure categories is structurally silent in 100% of observed cases.

AI system reliabilitydetectabilityevaluation blindness

This work addresses the challenge of behavioral inconsistency in automatically generated BPMN models due to semantic ambiguity in natural language process descriptions. It proposes the first closed-loop diagnosis and repair framework that operates without requiring ground-truth BPMN annotations. By analyzing the distribution of key performance indicators (KPIs) across multiple model generations, the approach identifies behavioral variations and employs model-based diagnostic techniques to pinpoint gateway logic discrepancies. These discrepancies are traced back to specific source text fragments, which are then refined through an evidence-driven textual revision process. Evaluated on clinical guidelines for diabetic kidney disease management, the method significantly reduces behavioral variability in regenerated models and enhances the semantic stability of executable process models, establishing an end-to-end mapping from behavioral inconsistency to targeted textual correction.

ambiguity detectionbehavioral inconsistencyBPMN

Latest Papers

What's happening recently
View more

This study investigates the “blind compliance” behavior of code large language models (LLMs), which execute erroneous instructions even when they recognize the errors, leading to irreversible “ghost errors” and semantic collapse. Through four sets of experiments on algorithmic Python problems from the RunBugRun dataset with deterministic test cases, the work systematically evaluates model responses in both single-attempt and multi-round repair scenarios. It formally defines and reveals the blind compliance phenomenon, quantifies the proportion of irrecoverable semantic failures, and demonstrates that extended reasoning fails to effectively mitigate such errors. The findings expose critical limitations of conventional pass-rate metrics and raise significant concerns about the reliability of code LLMs in production environments.

Blind ObedienceCode LLMsIncorrect Instructions

This study addresses the fundamental challenge posed by AI-generated code to the long-standing assumption in software engineering that authorship implies understanding. The authors demonstrate, for the first time, that AI-assisted programming systematically undermines the validity of authorship as a proxy for knowledge, thereby rendering traditional knowledge metrics—such as the truck factor—ineffective. By integrating principles from software engineering measurement theory, knowledge modeling, and logical reasoning, the work reinterprets the semantics of version control data in the context of AI collaboration. The research reveals that existing measures of knowledge concentration no longer reflect actual comprehension in AI-augmented development environments and advocates for a paradigm shift toward metrics grounded in verifiable evidence of understanding. It further identifies the construction of system-level understanding metrics as a critical open problem for the field.

AI code generationauthorship-based metricscomprehension

This study addresses the frequent operationalization failures that arise when large language models generate analytical workflows, stemming from a semantic gap between user intent and system-executable actions. Through cross-domain empirical analysis across finance, human resources, and public safety, the authors manually examined 236 analytical intents and their automatically generated workflows, systematically identifying and categorizing five distinct semantic-level failure patterns: comparative anchoring, procedural reasoning, quantitative reasoning, role confusion, and policy anchoring. The findings reveal fundamental limitations in the semantic expressiveness of current data systems and provide both theoretical grounding and practical guidance for improving the reliability of agent-generated analytical workflows.

agentic data systemsanalytical workflowslarge language models

This work addresses a common yet critical issue in machine learning code: semantic errors arising from mismatches between data properties and model assumptions—such as applying scale-sensitive models to unnormalized data—which traditional debugging approaches can only detect after training, resulting in inefficiency. To enable early and automatic error detection, the authors propose a novel data-aware static analysis method that integrates dataflow and control-flow analysis with API specifications, thereby incorporating data semantics directly into the static analysis framework for the first time. Evaluation on real-world machine learning notebooks demonstrates that the approach effectively identifies subtle semantic bugs that conventional techniques fail to catch, highlighting its practical utility and methodological innovation.

data-aware analysismachine learning codescale-sensitive models

Current automated formalization evaluations lack interpretable diagnostics for semantic errors, hindering both system optimization and human understanding. This work proposes FormalRx, a novel framework that introduces the first fine-grained, hierarchical taxonomy of Semantic Correctness Issues (SCI) comprising 28 error categories, and develops an end-to-end diagnostic model, FormalRx-8B, capable of aligning, classifying, localizing, and correcting formalization errors. Trained on 56,287 fine-grained annotated samples, the model achieves strong performance across four diagnostic tasks—0.88 F1 for alignment, 0.71 F1 for classification, 0.75 accuracy for localization, and 0.73 accuracy for correction—significantly outperforming both general-purpose large language models and specialized baselines. The study also releases FormalRx-Test, the first fine-grained diagnostic benchmark, thereby establishing a closed-loop pipeline from opaque evaluation to actionable feedback.

autoformalizationerror diagnosisevaluation framework

Hot Scholars

XW

Xilin Wei

Phd, Fudan University
MLLMefficient reasoningvideo understanding
IB

Ioana Boureanu

Professor, University of Surrey
Provable SecurityFormal VerificationApplied Cryptography
JJ

Junhao Jia

Hangzhou Dianzi University
Explainable AI (XAI)Interpretable Computer VisionMedical Image Analysis
AL

Alexander Lerchner

Senior Staff Scientist, Google DeepMind
NeuroscienceArtificial IntelligenceRepresentation Learning
YZ

Yefeng Zheng

Professor, Westlake University, Hangzhou, China, IEEE Fellow, AIMBE Fellow
AI in HealthMedical ImagingComputer VisionNatural Language Processing