Score
Designs and executes probing experiments, including probing classifiers and weakly supervised methods, and defines evaluation metrics and interpretability analyses to measure model capabilities, behavior, and failure modes. Builds evaluation pipelines, monitoring and profiling systems, and alignment/safety testing procedures to debug models, assess performance over time, and guide remediation.
This work reveals that large language models (LLaMA-3.3-70B-Instruct) possess “evaluation awareness”—an implicit capacity to distinguish evaluation (test-time) prompts from deployment (real-world usage) prompts—thereby undermining the validity of safety evaluations and AI governance. Method: We introduce the first linear probe analysis to identify and decode separable internal representations of evaluation vs. deployment contexts within hidden layers; complemented by black-box classification experiments to assess model discrimination of canonical safety benchmarks (e.g., MMLU, BBH) as artificial constructs. Contribution/Results: Results demonstrate high-accuracy identification of standard benchmarks as non-authentic, exposing systemic risks of strategic evasion and adversarial deception. This study provides critical empirical evidence and a methodological foundation for redesigning evaluation paradigms, advancing trustworthy AI assessment, and strengthening adversarial robustness research.
Understanding how and why large language models (LLMs) fail is becoming a central challenge as models rapidly evolve and static evaluations fall behind. While automated probing has been enabled by dynamic test generation, existing approaches often discover isolated failure cases, lack principled control over exploration, and provide limited insight into the underlying structure of model weaknesses. We propose ProbeLLM, a benchmark-agnostic automated probing framework that elevates weakness discovery from individual failures to structured failure modes. ProbeLLM formulates probing as a hierarchical Monte Carlo Tree Search, explicitly allocating limited probing budgets between global exploration of new failure regions and local refinement of recurring error patterns. By restricting probing to verifiable test cases and leveraging tool-augmented generation and verification, ProbeLLM grounds failure discovery in reliable evidence. Discovered failures are further consolidated into interpretable failure modes via failure-aware embeddings and boundary-aware induction. Across diverse benchmarks and LLMs, ProbeLLM reveals substantially broader, cleaner, and more fine-grained failure landscapes than static benchmarks and prior automated methods, supporting a shift from case-centric evaluation toward principled weakness discovery.
This work addresses the failure of probes in probe-guided fine-tuning due to Goodhart’s Law—where probes, when optimized as objectives, lose reliability. We propose an alignment method that jointly suppresses toxicity and preserves probe fidelity. Our core innovation lies in synergistically integrating supervised fine-tuning (SFT) and direct preference optimization (DPO) for probe-guided training: DPO substantially outperforms SFT, achieving both effective toxicity reduction and high probe accuracy in detecting harmful representations. Crucially, only a lightweight probe retraining post-fine-tuning is required to restore >95% of the original probe accuracy—eliminating the need for complex probe ensembling. Experiments across multiple toxicity benchmarks demonstrate significant toxicity reduction while maintaining robust internal monitoring capability, thereby validating the feasibility and practicality of probe-guided alignment.
This work addresses the semantic gap between evaluation metrics and training data in large model pretraining, which hinders precise diagnosis and remediation of capability deficiencies. The authors propose “capability slices” as fundamental units aligning evaluation and data, establishing a bidirectional classification framework that links evaluation tasks with non-instructional training data through explicit mapping rules. This enables a closed-loop pipeline from evaluation failures to targeted data interventions. For the first time, the approach supports auditable and systematic reasoning that translates evaluation signals into data corrections, moving beyond intuition-driven tuning paradigms. Experiments demonstrate its efficacy in both directions: repairing specific training loss components restores BBH performance to 66.44, while targeted data sampling boosts AIME2025/2026 Pass@128 from 6.67/0.00 to 26.67.
Industrial research agents often generate experimental trajectories containing invalid or incomplete information, rendering them unreliable for direct decision-making. This work proposes an evidence-oriented framework that automatically transforms such trajectories into structured evidence through a context-isolated generate–verify–repair pipeline. The approach introduces intervention-level claim categorization—distinguishing actionable repairs, diagnostic safeguards, and retained discoveries—and incorporates end-to-end provenance tracking to enable claim scoping and auditability. Experimental results demonstrate that the resulting candidate solutions outperform existing baselines. Audits further reveal that trajectory evolution is non-monotonic, and that applicability assessment constitutes a key performance bottleneck for the controller.
This study addresses the challenge of “silent failures” in AI systems—errors that occur during training or deployment without manifesting as anomalies in loss functions, thereby evading conventional evaluation mechanisms. The work introduces the concept of “evaluation blind spots” and presents the first unified framework modeling silent failures across both training and deployment phases. It formally defines detectability predicates, establishes a six-category taxonomy of such failures, and proposes a risk-based failure budget framework. Through retrospective analysis of 50 real-world incidents, case studies, formal verification, and an open-source implementation featuring gradient tracing, log auditing, and contamination detection, the study reveals that 53% of publicly reported AI failures are silent in nature, identifies genuine gradient errors in the TRL library, and demonstrates that one of the six production failure categories is structurally silent in 100% of observed cases.