Score
Designs and implements layerwise difference-in-differences analyses that compute how model outputs or internal signals change across consecutive layers by comparing treatment and control conditions (producing DID trajectories) and applying polarity-controlled contrasts. Uses those trajectories and summary metrics to detect transient mid-layer wrong preferences (dips), to quantify where preferences flip or correct across depth, and to stratify items by dip magnitude to predict or explain model failures.
This paper addresses three key challenges in difference-in-differences (DID) estimation under continuous treatment: (i) selection bias due to non-random treatment assignment, (ii) incomparability of treatment effects across varying intensities, and (iii) ambiguous causal interpretation of conventional two-way fixed-effects (TWFE) estimators. We propose a generalized parallel trends assumption and establish the first rigorous identification framework for continuous-treatment DID. We formally prove that TWFE estimators—even in a two-period setting—lack clear causal interpretation under continuous treatment. To overcome this, we develop a bias-corrected, group-weighted estimator grounded in treatment-effect decomposition and augmented with selection-bias sensitivity analysis. Empirically, our method substantially revises policy effect estimates derived from standard TWFE, mitigating systematic misattribution. The proposed approach provides a robust, interpretable tool for causal inference in settings involving graded or intensity-varying interventions.
This paper addresses how treatment-group selection threatens the parallel trends assumption in Difference-in-Differences (DiD) estimation—a critical yet under-characterized identification challenge. Method: We formally characterize the empirical content of this threat and derive necessary and sufficient conditions for parallel trends to hold under general selection mechanisms. We propose a “selection-driven bias decomposition framework” that systematically partitions DiD estimation bias into selection effects and time-varying heterogeneity effects, and develop operational benchmarking strategies—both with and without covariates—grounded in causal inference theory, selection modeling, and sensitivity analysis. Contribution/Results: Applied to the National Supported Work (NSW) experiment reanalysis, our approach quantifies and corrects selection bias, substantially improving the credibility of DiD estimates and the robustness of causal conclusions.
This paper addresses causal effect estimation under continuous, time-varying treatments (e.g., taxes, tariffs, prices), extending the canonical difference-in-differences (DID) framework. Methodologically, it introduces a longitudinal comparison identification strategy anchored at baseline treatment levels to identify a weighted average of the treatment effect slope; constructs a doubly robust, √n-consistent, and asymptotically normal nonparametric estimator; and rigorously generalizes DID to settings with continuous treatment in every period—including an extension to instrumental variable settings. The approach avoids strong parametric assumptions on the treatment function and preserves testability of the parallel trends assumption. Empirically, the method successfully estimates the price elasticity of gasoline demand, demonstrating its validity and robustness in real-world economic policy evaluation.
Existing methods struggle to align and interpret distribution shifts across heterogeneous, domain-consistent datasets—such as tabular, textual, visual, and time-series data—especially when scale and modality disparities are pronounced, resulting in poor interpretability. This paper introduces the first human-centric, cross-modal distribution discrepancy explanation framework, implemented as an interpretable dataset comparison toolbox. It integrates statistical hypothesis testing, feature importance decomposition, class activation mapping (CAM), contrastive representation learning, and interpretable generative modeling to enable fine-grained, semantically readable attribution and visualization of distributional shifts. Evaluated across diverse real-world scenarios, the framework significantly improves users’ efficiency in understanding shift causes and enhances the accuracy of intervention decisions—thereby overcoming the limitations of conventional black-box shift detection approaches.
This paper addresses bias and efficiency concerns in Difference-in-Differences (DiD) estimation of the Average Treatment Effect on the Treated (ATT) under compositional change—i.e., time-varying population structure—in repeated cross-sectional data. Methodologically, it introduces: (i) the first nonparametric Hausman-type test to detect compositional change; (ii) a rate-form doubly robust estimation framework; and (iii) a consistent stochastic expansion for a local polynomial multinomial logit estimator. Theoretically, it achieves the semiparametric efficiency bound with optimal convergence rates within a semiparametric efficiency framework and characterizes the bias–efficiency trade-off arising from ignoring compositional change. Monte Carlo simulations and empirical applications demonstrate that the proposed method substantially outperforms existing approaches in bias control, statistical efficiency, and test power.
This study addresses the identification bias in difference-in-differences (DiD) estimation arising from the use of covariates affected by treatment—commonly termed “bad controls.” When the parallel trends assumption holds only after conditioning on such variables, the paper proposes two novel approaches: first, conditioning exclusively on their pre-treatment values; second, under an unconfoundedness condition for the covariates, combining imputation with double/debiased machine learning to estimate treatment effects. The work provides the first systematic characterization of the identification conditions under which “bad controls” can serve as valid control variables and extends the framework to staggered treatment settings, offering corresponding pre-tests. Empirical application to the effect of unemployment on income demonstrates the validity and robustness of the proposed methods.
研究解决了使用差分法在受限评分尺度上评估LLM法官偏差时可能制造出虚假效应的问题,通过审计和数学推导方法揭示了这一机制。
本文研究了具有连续处理变量的交错采用差异中的差异方法,通过两个维度(水平和响应)来解决因果效应估计问题。
This work reveals a “wrong-then-correct” phenomenon in aligned language models, wherein intermediate layers transiently favor incorrect answers—a behavior termed “causal probing”—before later layers rectify the output. Through minimal pairs, layer-wise differential analysis, and Patchscope-style activation transplantation, the study provides the first evidence that this mechanism is contingent on model scale (≥3B parameters) and alignment methodology, and substantially undermines compression robustness. To mitigate this issue, the authors propose a LoRA-based training strategy that penalizes erroneous intermediate representations. This approach reduces causal probing by 67–70% while preserving task accuracy, thereby significantly enhancing the stability of model compression.
本文提出了一种针对具有任意边界形状的边界不连续设计(BDD)的操纵测试方法,通过k-均值聚类和二项平衡测试来检测处理组与对照组的均衡性。