Score
Designs and applies quantitative attribution and analysis methods to localize toxic content representations inside a model’s internals, producing layer- and neuron-level rankings, activation-differential maps, and sets of candidate “toxic” neurons. Builds tools and procedures to identify which layers and individual neurons disproportionately encode toxicity signals and to support targeted probing, ablation, or other interventions.
This work investigates the causal relationship between internal representations of language models (LMs) and their toxic output behavior. We propose the first alignment probing framework that enables fine-grained, dynamic alignment between output toxicity and multilayer internal representations, systematically dissecting toxicity encoding mechanisms across 20+ mainstream models (e.g., OLMo, Llama, Mistral). Our analysis reveals that: (i) lower-layer representations strongly encode input-level toxicity, and causal interventions at these layers significantly reduce output toxicity; (ii) toxicity exhibits intrinsic heterogeneity across semantic dimensions—particularly in threat-related contexts. Methodologically, we integrate probe-based interpretability, inter-layer feature attribution, cross-model comparison, causal intervention, and multi-scenario case studies. Key contributions include: (1) a reproducible, quantitative metric for toxicity attribution; and (2) empirical validation of the framework’s practical utility in model debugging, safety alignment, and training-time monitoring.
Large language models (LLMs) often generate harmful content, posing serious risks to AI safety and public trust. Existing neuron-level intervention methods suffer from poor stability, strong context dependency, and unintended degradation of linguistic capabilities. To address these limitations, we propose EigenShift—a fine-tuning-free, low-overhead layer-wise intervention framework. EigenShift decouples the toxicity generation mechanism via inter-layer feature aggregation and output-layer feature decomposition. It innovatively separates toxicity detection and generation into distinct expert modules, enabling structured intervention through eigenvalue decomposition and generation-alignment analysis. Evaluated on the Jigsaw and ToxiCN benchmarks, EigenShift achieves significant and robust suppression of toxic outputs while preserving language modeling performance. The method demonstrates strong interpretability, cross-dataset generalizability, and deployment efficiency—offering a practical, principled solution for safe LLM inference.
This work addresses the propensity of large language models to generate toxic content, a challenge exacerbated by limited understanding of their internal mechanisms. By analyzing activation differences between toxic and neutral prompts, the study identifies—without any model retraining—that toxicity primarily originates from specific neurons in early MLP layers. Leveraging the Meow2X and TRNE frameworks, the authors employ activation difference analysis, inference-time scaling, and rank-one weight editing to precisely suppress toxic outputs. Evaluated across five models, two benchmarks, and 90 configurations, the approach significantly reduces toxicity while preserving language modeling performance. Furthermore, the research demonstrates that reliance on a single toxicity evaluator systematically underestimates risk, thereby underscoring the necessity of multi-evaluator safety assessments.
The prevailing hypothesis that DPO reduces language model toxicity by merely suppressing a small set of “toxic neurons” is overly simplistic and lacks mechanistic grounding. Method: We employ activation patching, toxicity probe projection, hierarchical clustering, and causal attribution analysis to dissect DPO’s internal mechanisms across model layers. Contribution/Results: We reveal that DPO induces distributed, progressive activation shifts across numerous MLP neurons in multiple layers—not localized suppression. We identify four functionally distinct neuron groups—two detoxifying and two anti-toxicity-promoting—whose cumulative activation shifts account for 95.1% of the observed toxicity reduction; conventional “toxic neuron suppression” contributes only 4.9%. Full activation patching of these four groups fully restores DPO’s toxicity mitigation effect. This work establishes DPO as a distributed, multi-stage regulatory process, challenging localized attribution models and providing a new interpretability framework for alignment.
Large language models are prone to generating toxic content, and existing alignment methods struggle to fully eliminate toxicity embedded in model parameters while remaining vulnerable to adversarial attacks. This work proposes GLOSS, a novel approach that introduces the concept of a “global toxicity subspace” for the first time. Leveraging mechanistic interpretability, GLOSS precisely identifies and removes this subspace within feedforward network (FFN) layers, overcoming limitations of local methods that are susceptible to reconstruction or noise interference. Requiring no extensive retraining and incurring minimal computational overhead, GLOSS achieves state-of-the-art detoxification performance on large models such as Qwen3 while effectively preserving their general capabilities.
This work addresses the lack of concept-level interpretability and imbalanced concept attribution—leading to misclassifications—in toxic language detection. Methodologically: (1) it treats semantic subtypes (e.g., insult, threat, identity attack) as interpretable concepts, constructs a target lexicon, and proposes a Word–Concept Alignment (WCA) score to quantify each token’s contribution to misclassification via concept gradients (CG); (2) it introduces, for the first time, a delexicalized generative data augmentation strategy to assess model reliance on abstract toxic patterns rather than surface lexical cues. Experiments demonstrate that CG precisely identifies critical toxic tokens and reveal that models over-attribute toxicity to conceptual features even when explicit toxic words are absent—exposing systematic generalization biases toward deep semantic patterns and implicit dependency mechanisms. This establishes a novel paradigm for interpretable toxic language detection grounded in concept-level attribution and causal probing.
This work addresses the challenge of detecting implicit toxicity in multimodal data, where harmful semantics emerge only through cross-modal fusion and thus evade existing detection methods. To tackle this issue, the authors propose Toxicity Association Graphs (TAGs) to model semantic relationships between seemingly benign entities and latent toxic content, and introduce a Multimodal Toxicity Covertness (MTC) metric to quantify the degree of implicit toxicity. Leveraging this framework, they construct Covert Toxic Dataset—the first benchmark dataset focused on high-covert multimodal toxicity—and integrate graph neural networks with explainable AI techniques to enable interpretable and auditable toxicity detection. Experimental results demonstrate that the proposed approach consistently outperforms state-of-the-art methods across both high- and low-covertness scenarios, significantly advancing the field of explainable multimodal toxicity detection.
This study addresses the challenges in preclinical drug toxicity assessment posed by limited expert resources and the scarcity of rare pathological samples. The authors propose an AI-driven anomaly detection framework that leverages whole-slide images to segment healthy and known pathological regions in rodent liver tissue while effectively identifying rare anomalies absent from the training data. The method integrates the DINOv2 vision transformer with LoRA fine-tuning to achieve high-precision tissue segmentation and introduces a novel class-aware Mahalanobis distance combined with class-specific thresholding to distinguish between known pathologies and unknown anomalies. Evaluated on mouse liver data, the approach demonstrates exceptional performance, misclassifying only 0.16% of lesions as healthy and 0.35% of healthy tissue as pathological, thereby significantly enhancing generalization and detection accuracy.
Current neurotoxicity assessments rely on subjective manual scoring, which struggles to efficiently quantify multiscale degenerative lesions from C. elegans neural imagery or predict associated behavioral phenotypes. This work proposes a scale-adaptive masked image modeling approach to build a self-supervised visual foundation model tailored for dopaminergic neurons in C. elegans, enabling joint learning of multiresolution features under fixed computational budgets while overcoming conventional grid constraints. For the first time, this framework achieves end-to-end extraction of structural semantics from sparse, multiscale images. Evaluated on CeNeuMorph—a newly established multimodal confocal imaging benchmark—the model outperforms general-purpose and biomedical foundation models across classification, segmentation, and detection tasks. By integrating morphological and visual features, it effectively predicts dopamine-related behavioral deficits (R² = 0.498) and successfully screens 180 agrochemicals, identifying benzimidazole as a novel dopaminergic neurotoxicity determinant.
This study addresses the limitations of existing speech toxicity detection methods, which predominantly rely on textual content while neglecting paralinguistic cues such as emotion and prosody, and suffer from a scarcity of large-scale annotated data. To bridge this gap, the authors introduce ToxiAlert-Bench, a novel dataset comprising over 30,000 audio clips, systematically annotated with seven coarse-grained toxicity categories and twenty fine-grained labels, explicitly distinguishing toxicity sources—textual versus paralinguistic. They propose a source-aware dual-task detection framework employing a two-headed neural network and a multi-stage training strategy (independent pretraining followed by joint fine-tuning), enhanced with class-balanced sampling and a weighted loss function to effectively disentangle textual and paralinguistic toxicity signals. Experiments demonstrate that the proposed approach achieves a relative improvement of 21.1% in Macro-F1 and 13.0% in accuracy, significantly outperforming current baselines.