Score
Designs and implements pipelines and artifacts that extract, engineer, and validate interpretable linguistic features and lexica from text—e.g., keyword/lexicon sets, syntactic and lexical cues, and scalar linguistic metrics—to serve as inputs for feature-based detectors or classifiers. Analyzes and audits those lexica and detectors to diagnose failures (polysemy, syntactic blindness, categorical absence, measurement inversions), produce lexicon/feature failure reports, and refine feature-based, interpretable detection models.
This study addresses the limited understanding of generalizable linguistic features that reliably distinguish human-written from large language model–generated text. Through large-scale empirical analysis, the authors evaluate the robustness of 284 interpretable linguistic features across 27 large language models and 10 textual domains, and develop a classifier driven solely by linguistic features. They identify lexical richness as the most robust discriminative signal across both models and domains—a finding reported for the first time—and expose the limitations of many context-dependent features. The results demonstrate that AI-generated text can be effectively detected using only linguistic cues, and pinpoint core features with strong generalization capabilities, thereby laying the groundwork for interpretable AI text detection.
This work addresses the limited generalization of current AI-generated text detectors in real-world scenarios, where it remains unclear whether models learn universal machine-authorship signatures or dataset-specific stylistic artifacts. To investigate this, the authors propose a detection framework integrating linguistic feature engineering, machine learning, and SHAP-based interpretability analysis. Their approach systematically reveals a fundamental contradiction: linguistic features effective within a domain often fail to generalize across domains. Comprehensive cross-generator and cross-domain evaluations demonstrate that prevailing methods heavily rely on dataset-specific cues rather than stable generative signals. While the model achieves a strong in-domain F1 score of 0.9734 on PAN CLEF 2025 and COLING 2025 benchmarks, its performance degrades significantly under domain shift. The authors release an open-source Python toolkit supporting both prediction and instance-level explanations.
The dynamic acquisition of linguistic knowledge during language model training remains poorly understood; existing interpretability tools are largely post-hoc, rely on scalar metrics, or involve complex integrations—hindering deployment and reproducibility. This paper introduces a modular interpretability analysis framework enabling fine-grained tracking of linguistic and representational signals throughout both training and inference in Transformer models. Methodologically, it integrates (1) ABSynth, a controllable synthetic corpus generator that uncovers evolution patterns—such as early syntactic emergence, delayed semantic acquisition, and representation compression—overlooked by conventional metrics; and (2) a multi-faceted diagnostic suite combining feature probing, intrinsic dimension estimation, Hessian curvature analysis, and output diagnostics to support layer-wise interpretation, convergence-driven early stopping, and structural error detection. Lightweight and fully reproducible, the framework significantly enhances the operationality and systematicity of interpretability research.
Existing automatic explanation methods predominantly adopt an input-centric paradigm, limiting their ability to characterize causal feature effects on model outputs and handle “dead features” in large language models. This work proposes an output-centric feature attribution paradigm, anchoring explanations at the model’s final output. It enables lightweight, automated causal attribution via steering interventions, vocabulary-decoupled head projections, and token-weight analysis. Crucially, it shifts the explanatory focus from input activations to output effects, explicitly modeling features’ causal contributions to generation outcomes, while supporting dead-feature reactivation and diagnosis. Experiments demonstrate that our method significantly outperforms baselines in output-behavior explanation accuracy. When integrated with both input- and output-centric perspectives, it achieves state-of-the-art performance on both input-relevance and output-faithfulness evaluation metrics.
Existing exemplar-based interpretation methods for Transformer hidden states rely on subjective visual inspection, limiting accurate semantic characterization of features. Method: We propose a rule-based paradigm for feature interpretation, modeling attention-layer feature behavior as three interpretable rule classes—skip-grams, omission rules, and counting rules—capturing cross-position dependencies, critical token omission effects, and numerical counting logic, respectively. Using sparse autoencoders (SAEs) to extract attention features from GPT-2 small and automated pattern mining, we systematically uncover structured input–output mapping regularities. Contribution/Results: Our approach reveals that over 25% of features exhibit early-layer omission rules; multiple counting rules are successfully identified; and the majority of features receive precise, empirically verifiable rule descriptions. This significantly improves interpretability accuracy and generalizability across layers and tasks, establishing the first systematic, rule-driven framework for attention feature explanation.
Lexical data in linguistic documentation frequently contain transcription errors and unannotated loanwords, introducing bias into phonological analysis. This paper addresses these challenges for the low-resource language Kokborok by proposing an unsupervised anomaly detection method that innovatively integrates character-level and syllable-aware phonological features to identify both transcription errors and covert loanwords in lexical inventories. Evaluated on a Kokborok–Bengali multilingual dataset, the method significantly outperforms a character-only baseline, achieving high recall while maintaining practical applicability and systematic rigor. Although precision is constrained by the subtle, linguistically embedded nature of certain anomalies, the method’s strong recall enables field linguists to perform actionable data quality diagnostics. It thus establishes a novel paradigm for automated cleaning and annotation of lexical data in low-resource language documentation.
Conventional quantitative reliability analysis struggles to extract deep semantic information from unstructured maintenance logs of wind turbines, while existing machine learning methods are limited to shallow classification tasks. Method: This paper proposes a reproducible large language model (LLM)-driven framework that employs an LLM as a “collaborative pilot” to perform four types of deep semantic analysis: failure mode identification, causal chain inference, site-to-site comparative analysis, and data quality auditing. The framework integrates industrial-scale operational logs with the LLM’s natural language understanding and reasoning capabilities to construct a semantic reliability–oriented analytical pipeline. Contribution/Results: Evaluated on real-world wind power datasets, the framework generates actionable, expert-level diagnostic hypotheses—systematically unlocking implicit knowledge embedded in unstructured logs for the first time—and significantly advances the intelligence level of wind farm operations.
This work addresses the lack of traceability in existing domain-specific fine-tuning approaches, which often leads to blind and inefficient data augmentation. The authors propose a “programming with data” paradigm that treats structured knowledge representations as a unified foundation for both training and evaluation, drawing an analogy to software development: training data serve as source code, model training as compilation, evaluation as unit testing, and data refinement as debugging. This framework enables precise, concept- and reasoning-chain–oriented model repair through structured knowledge extraction, test-driven data engineering, concept-level gap analysis, and diagnosis of broken reasoning chains. Validated across 16 disciplines, the approach significantly enhances model performance without compromising general capabilities, and the authors release an open-source knowledge base, evaluation suite, and training corpora to support reproducibility and further research.
Intermittent failures in continuous integration (CI) pipelines are notoriously difficult to diagnose, leading to wasted resources and reduced development efficiency. This work proposes FlaXifyer, a few-shot learning approach that integrates the interpretable AI technique LogSift to fine-tune pretrained language models on pipeline logs using only 12 labeled examples per failure class. The method simultaneously predicts failure categories and pinpoints critical log entries indicative of root causes. Evaluated on 2,458 real-world CI failures, FlaXifyer achieves a Macro F1 score of 84.3% and a Top-2 accuracy of 92.0%, reducing the required log inspection effort by 74.4%. Furthermore, it successfully identifies the underlying fault in 87% of cases, demonstrating its effectiveness in accelerating failure diagnosis with minimal labeled data.