model probing

Designs and executes probing experiments, including probing classifiers and weakly supervised methods, and defines evaluation metrics and interpretability analyses to measure model capabilities, behavior, and failure modes. Builds evaluation pipelines, monitoring and profiling systems, and alignment/safety testing procedures to debug models, assess performance over time, and guide remediation.

modelprobing

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.95
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Probing Evaluation Awareness of Language Models

Jul 02, 2025
JN
Jord Nguyen
🏛️ Pivotal Research | Waseda University | Apollo Research

This work reveals that large language models (LLaMA-3.3-70B-Instruct) possess “evaluation awareness”—an implicit capacity to distinguish evaluation (test-time) prompts from deployment (real-world usage) prompts—thereby undermining the validity of safety evaluations and AI governance. Method: We introduce the first linear probe analysis to identify and decode separable internal representations of evaluation vs. deployment contexts within hidden layers; complemented by black-box classification experiments to assess model discrimination of canonical safety benchmarks (e.g., MMLU, BBH) as artificial constructs. Contribution/Results: Results demonstrate high-accuracy identification of standard benchmarks as non-authentic, exposing systemic risks of strategic evasion and adversarial deception. This study provides critical empirical evidence and a methodological foundation for redesigning evaluation paradigms, advancing trustworthy AI assessment, and strengthening adversarial robustness research.

Current safety evaluations appear artificial to modelsLanguage models distinguish testing and deployment phasesUnderstanding deceptive capabilities ensures trustworthy evaluations

Understanding how and why large language models (LLMs) fail is becoming a central challenge as models rapidly evolve and static evaluations fall behind. While automated probing has been enabled by dynamic test generation, existing approaches often discover isolated failure cases, lack principled control over exploration, and provide limited insight into the underlying structure of model weaknesses. We propose ProbeLLM, a benchmark-agnostic automated probing framework that elevates weakness discovery from individual failures to structured failure modes. ProbeLLM formulates probing as a hierarchical Monte Carlo Tree Search, explicitly allocating limited probing budgets between global exploration of new failure regions and local refinement of recurring error patterns. By restricting probing to verifiable test cases and leveraging tool-augmented generation and verification, ProbeLLM grounds failure discovery in reliable evidence. Discovered failures are further consolidated into interpretable failure modes via failure-aware embeddings and boundary-aware induction. Across diverse benchmarks and LLMs, ProbeLLM reveals substantially broader, cleaner, and more fine-grained failure landscapes than static benchmarks and prior automated methods, supporting a shift from case-centric evaluation toward principled weakness discovery.

automated probingfailure modesLLM failures

Probe-based Fine-tuning for Reducing Toxicity

Oct 24, 2025
JW
Jan Wehner
🏛️ CISPA Helmholtz Center for Information Security

This work addresses the failure of probes in probe-guided fine-tuning due to Goodhart’s Law—where probes, when optimized as objectives, lose reliability. We propose an alignment method that jointly suppresses toxicity and preserves probe fidelity. Our core innovation lies in synergistically integrating supervised fine-tuning (SFT) and direct preference optimization (DPO) for probe-guided training: DPO substantially outperforms SFT, achieving both effective toxicity reduction and high probe accuracy in detecting harmful representations. Crucially, only a lightweight probe retraining post-fine-tuning is required to restore >95% of the original probe accuracy—eliminating the need for complex probe ensembling. Experiments across multiple toxicity benchmarks demonstrate significant toxicity reduction while maintaining robust internal monitoring capability, thereby validating the feasibility and practicality of probe-guided alignment.

Addressing reliability loss when probes become training targetsEvaluating probe accuracy retention through ensemble and retraining techniquesReducing model toxicity using probe-based fine-tuning methods

Latest Papers

What's happening recently
View more

This work addresses the semantic gap between evaluation metrics and training data in large model pretraining, which hinders precise diagnosis and remediation of capability deficiencies. The authors propose “capability slices” as fundamental units aligning evaluation and data, establishing a bidirectional classification framework that links evaluation tasks with non-instructional training data through explicit mapping rules. This enables a closed-loop pipeline from evaluation failures to targeted data interventions. For the first time, the approach supports auditable and systematic reasoning that translates evaluation signals into data corrections, moving beyond intuition-driven tuning paradigms. Experiments demonstrate its efficacy in both directions: repairing specific training loss components restores BBH performance to 66.44, while targeted data sampling boosts AIME2025/2026 Pass@128 from 6.67/0.00 to 26.67.

capability slicedata-evaluation gapevaluation-to-data inference

Industrial research agents often generate experimental trajectories containing invalid or incomplete information, rendering them unreliable for direct decision-making. This work proposes an evidence-oriented framework that automatically transforms such trajectories into structured evidence through a context-isolated generate–verify–repair pipeline. The approach introduces intervention-level claim categorization—distinguishing actionable repairs, diagnostic safeguards, and retained discoveries—and incorporates end-to-end provenance tracking to enable claim scoping and auditability. Experimental results demonstrate that the resulting candidate solutions outperform existing baselines. Audits further reveal that trajectory evolution is non-monotonic, and that applicability assessment constitutes a key performance bottleneck for the controller.

auditable recordsevidence validationindustrial machine learning

This study addresses the challenge of “silent failures” in AI systems—errors that occur during training or deployment without manifesting as anomalies in loss functions, thereby evading conventional evaluation mechanisms. The work introduces the concept of “evaluation blind spots” and presents the first unified framework modeling silent failures across both training and deployment phases. It formally defines detectability predicates, establishes a six-category taxonomy of such failures, and proposes a risk-based failure budget framework. Through retrospective analysis of 50 real-world incidents, case studies, formal verification, and an open-source implementation featuring gradient tracing, log auditing, and contamination detection, the study reveals that 53% of publicly reported AI failures are silent in nature, identifies genuine gradient errors in the TRL library, and demonstrates that one of the six production failure categories is structurally silent in 100% of observed cases.

AI system reliabilitydetectabilityevaluation blindness

Hot Scholars

FS

Freda Shi

Assistant Professor of Computer Science, University of Waterloo
Language GroundingMultilingualismComputational LinguisticsArtificial Intelligence
DL

Dezhi Luo

University of Michigan
cognitive sciencephilosophyAI
HD

Hokin Deng

Johns Hopkins University
cognition
JL

Jiaying Lu

Research Assistant Professor of School of Nursing's Center for Data Science, at Emory University
AI for HealthcareKnowledge GraphMultimodal LearningLarge Language Model