layerwise prior localization

Designs and evaluates methods that identify and map where credibility priors are encoded across a model’s layers and at the level of individual features, isolating sparse, seed-stable feature encodings and measuring scale-dependent encoding strength. Builds mechanisms to steer or intervene on those localized priors—for example targeted feature interventions and layerwise modulation—to establish causal locus and control of prior-driven behavior.

layerwisepriorlocalization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.1
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge in causal discovery where external priors are often heterogeneous and of uncertain reliability—traditional approaches either blindly trust such priors, amplifying errors, or discard them entirely, forfeiting useful signals. To overcome this, we propose PRCD-MAP, which introduces, for the first time, an edge-level trust mechanism that assigns each prior edge an independent learnable weight. This enables fine-grained, adaptive integration of imperfect priors by dynamically modulating the ℓ₁/ℓ₂ mixed regularizer within the maximum a posteriori (MAP) objective. Combining empirical Bayes calibration, Laplace-approximated marginal likelihood, and MLP-based trust propagation, our method achieves substantial performance gains under theoretical safety guarantees: on the CausalTime real-world datasets, it yields significant AUROC improvements (AQI +0.123, Medical +0.043), outperforming PCMCI+ and BayesDAG, while remaining robust in high-dimensional (d=300) and latent-confounded settings.

causal discoveryexternal knowledgeheterogeneous reliability

This work addresses a critical limitation in existing interventional interpretability evaluations, which rely on point estimates and struggle to disentangle true causal effects from sampling or adaptive biases. The authors reformulate the problem as a causal estimation task and introduce, for the first time, an anytime-valid statistical certification framework that accommodates adaptive intervention sampling. By integrating Hoeffding-type confidence sequences with variance-adaptive betting strategies and bounded mixture importance weighting, the method yields both confidence intervals and dynamic confidence sequences for intervention fidelity. Empirical validation on MNIST abstractions and GPT-2 Small IOI circuit experiments demonstrates that the approach not only certifies high-fidelity interpretability claims and detects statistically insignificant differences between methods but also reduces certification costs by 10–30× compared to existing baselines.

adaptive samplingcausal claimsinterventional evaluation

This work addresses the susceptibility of vision-language models to media brand cues—such as mastheads and logos—that induce content-agnostic priors in news credibility assessment. The authors introduce the CueTrust benchmark and the Source-Override Index (SOI), leveraging sparse autoencoders, inter-layer interventions, and cross-model diagnostics to mechanistically reveal, for the first time, that this prior exhibits a dual-encoding structure, is spatially localizable, and consistently replicable across models. Experiments demonstrate that brand cues can significantly modulate credibility judgments across an 11 log-odds range (ρ = 0.88). Targeted interventions reduce the source-overriding effect by 41% and generalize effectively to unseen media outlets, highlighting both the vulnerability and mitigability of such spurious associations in multimodal trust evaluation.

credibility priormedia biasnews trustworthiness

This work addresses the challenge of detecting and regulating sycophantic behavior—excessive user flattery—in language models by proposing an iterative data generation method based on cascaded linear samples. Departing from conventional binary contrastive examples, the approach constructs sequences of samples with continuously varying behavioral intensities, revealing for the first time a linearly separable structure of sycophancy in activation space. This enables precise identification and disentanglement of the associated feature subspace. Through activation manipulation and subspace analysis, the method matches or exceeds baseline approaches such as LLM-as-a-judge and system prompting in detection accuracy, calibration, and robust controllability, while incurring lower computational overhead and substantially improving the interpretability of behavioral interventions.

activation steeringbehavior controlinterpretable features

This study addresses the growing threat posed by increasingly realistic misinformation generated by foundation models to trustworthy online information ecosystems. The authors propose a dual-axis framework, JudgeGPT and RogueGPT, which decouples “factual accuracy” from “source attribution,” and introduces the concept of the “fluency trap” to elucidate the cognitive mechanisms underlying human susceptibility to hallucinations. Leveraging a structural causal model, 918 human evaluations, and comparisons across multiple models—including GPT-4 and Llama-2—the research finds that political orientation exerts minimal influence, whereas familiarity with fake news serves as a key mediating variable (r = 0.35). Notably, GPT-4–generated content achieves a human–machine confusion rate of 0.20. The findings advocate for “prebunking” interventions centered on cognitive source monitoring, offering both empirical grounding and a novel pathway toward fostering a more reliable information ecosystem.

disinformationfoundation modelshallucinations

Latest Papers

What's happening recently
View more

This study addresses the degradation of explanation fidelity caused by data drift, wherein legacy explanations fail to effectively monitor updated model behavior. To investigate this, we employ ExIFFI to generate local explanations and integrate conditional drift detection with multi-level intervention evaluation, systematically examining how covariate shift affects explanation faithfulness. Our work provides the first empirical evidence that structural stability does not entail functional faithfulness, demonstrating that explanations must be independently monitored and bound to specific models. Furthermore, we show that although pre-retraining explanations retain relevance after model updates, their fidelity is significantly inferior to that of newly generated explanations, while no systematic widening of this discrepancy is observed.

Concept DriftData DriftExplainability

This study addresses the problem of directed feature learning for predefined concepts in language models, bridging a gap in sparse autoencoders for hypothesis-driven scenarios. By systematically evaluating combinations of signals—such as activations and gradients—with various estimators, this work proposes four novel methods for directed feature extraction and benchmarks them against contrastive mean, 1D autoencoders, and sparse autoencoders. Results reveal that concept detection and causal intervention are predominantly governed by activations and gradients, respectively, exhibiting complementary properties. Furthermore, the quality of directed features depends critically on the joint selection of signal and estimator. By delineating the optimal application scenarios for each approach, this research provides systematic guidance for controllable feature extraction in language models.

Causal InterventionFeature DetectionLanguage Model Interpretability

This study addresses the challenge of rapidly distinguishing individual sample contributions and the lack of controllability in behavioral interventions during large model training. To this end, it introduces mutual information to quantify intra-batch interference and proposes a novel Behavior Gradient Uniqueness (BGU) metric grounded in geometric interpretation, alongside the BS-Ghost algorithm. These innovations enable gradient-free batch-space shared computation and real-time weight intervention. Consequently, this work presents a low-overhead framework for training-time attribution and dynamic control. Evaluated on Qwen2.5-7B, the proposed approach incurs only an 8% increase in training time while accurately identifying causal data samples and steering model evolution toward target behaviors, thereby offering an efficient and controllable paradigm for optimizing large model training.

batch interferencebehavior controlmodel alignment

Hot Scholars

MP

Maria Pateraki

Associate Professor National Technical University of Athens, Affiliated Researcher FORTH
PhotogrammetryComputer VisionRobotic perception
DB

Djallel Bouneffouf

Unknown affiliation
Reinforcement learningMulti-armed banditsContext-aware Recommender systems
FC

Florin C. Ghesu

Digital Technology & Innovation, Siemens Healthineers
Machine LearningImage UnderstandingAlgorithmsMedical Image Analysis
AK

Andreas K. Maier

Friedrich-Alexander-Universität Erlangen-Nürnberg
pattern recognitionmachine learningspeech processingmedical speech processing
SI

Saahil Islam

Friedrich–Alexander University Erlangen–Nürnberg
Computer visionmedical imaging