Circuit-Diff: Factual Edit-based Intervention Method for Localizing Knowledge in Attribution Graphs

📅 2026-09-20
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为了解决神经网络中电路识别问题,提出Circuit-Diff方法,通过低秩事实编辑干预模型,识别出与编辑知识相关的特征。
📝 Abstract
Mechanistic interpretability defines features as the fundamental units of a neural network and circuits as the weighted subgraphs that carry out its computation. Because individual neurons are polysemantic, Cross-Layer Transcoders (CLTs) were introduced as a way to approximate a model's circuits by generating an attribution graph. The nodes of that graph, however, are unlabeled features: reading a graph means pruning it and then working out by hand what each surviving node means. To make CLTs easier to use for circuit discovery, we introduce Circuit-Diff, which intervenes on the model itself with a low-rank factual edit and takes the features whose role in the attribution graph changes under that edit as related to the edited knowledge. On the edits we examine, the flagged nodes are not only detectors of the object token: read off the CLT's released feature dashboards, they include features for the history, geography and associations surrounding the old and new objects. We formalize the method, measure how reliable a frozen CLT remains after a factual edit, test the selected nodes causally by patching them on up to 24 CounterFact edits, give a case study, and release an open-source implementation built on the circuit-tracer package, together with two further tools (multi-prompt aggregation and rule-based supernode labeling).
Problem

Research questions and friction points this paper is trying to address.

mechanistic interpretability
neural network circuits
Cross-Layer Transcoders (CLTs)
attribution graphs
unlabeled features
Innovation

Methods, ideas, or system contributions that make the work stand out.

Circuit-Diff
low-rank factual edit
attribution graph
Cross-Layer Transcoders (CLTs)
mechanistic interpretability