内在结构:机制可解释性的谱可识别性
摘要
本文引入了一个基于Koopman算子理论的理论框架,用于识别机制可解释性中的模型内在结构,证明了针对机制可解释性原语的第一个可识别性定理,并在GPT-2、Gemma-2-2B和Qwen3-8B-Base上进行了实证验证。
arXiv:2608.10172v1 公告类型:新
摘要:机制可解释性通过识别模型内部的电路来解释模型,但无法判断一个电路是模型的固有属性还是发现它的方法带来的假象。稀疏自编码器就说明了这个问题:不同的随机种子和宽度会从相同的激活中恢复出实质不同的特征,而且没有任何理论说明这种变异性是偶然的还是结构性的。我们为用于可解释性的词典学习奠定了可识别性基础。将前向传播视为以深度为时间的受控动力系统,并用Koopman算子将其提升,得到一个有限线性实现,其\emph{谱}是模型的一个与坐标无关的属性。我们证明该谱可以从$M$个校准样本中以$M^{-1/2}$的速率恢复(在置换意义下)——据我们所知,这是针对机制可解释性原语的第一个可识别性定理,同时给出了匹配的极小极大下界、适用于重尾激活的中位数均值变体,以及一个解离定理:只要实现是非正规的,携带激活方差的那些方向和跨深度传递信息的方向就不可能重合。可识别的对象和可读的对象不是同一个对象。在GPT-2 small、Gemma-2-2B和Qwen3-8B-Base上,谱处处收敛,并在Qwen3-8B-Base上达到预测的指数($0.506 \pm 0.031$);偏差随着每个单元的样本阈值消失,并汇聚到同一条曲线上。Koopman模式优于随机方向,但在间接宾语识别上不如主成分,差距随深度距离以$4.1\times$衰减,正如定理所预测的那样。Koopman谱是可识别的、模型内在的指纹,带有明确的误差条,而不是一种可读的分解。
查看缓存全文
缓存时间: 2026/08/12 08:28
". So we only need translate the provided content up to that point. The source includes title, abstract, contributions, 1 Introduction, 2 Background. Need translate to "### 2 背景与相关工作" and the paragraph ending around "Causal abstraction ...". The line in prompt ends abruptly: "Causal abstraction\[24 (https://arxiv.org/html/2608.10172#bib.bib28),25 (https:" because user may have truncated. We need not complete the unfinished link? We should translate the visible text, preserving incomplete? Hmm. The provided text ends with "Causal abstraction\[24 (https://arxiv.org/html/2608.10172#bib.bib28),25 (https:" without closing bracket. It's likely a truncation from stack. We should include up to that, maybe "因果抽象[24](https://arxiv.org/html/2608.10172#bib.bib28)、[25](https://..." but the original is incomplete. We need not add content. But if we output incomplete link, is that okay? We can close with "https://"... cannot invent. We can translate to "因果抽象[24](https://arxiv.org/html/2608.10172#bib.bib28)、[25](https://" and stop? But since not valid, maybe we can drop the incomplete link? Better to preserve the visible text "Causal abstraction[24 (https://arxiv.org/html/2608.10172#bib.bib28),25 (https:"; in Chinese "因果抽象[24(https://arxiv.org/html/2608.10172#bib.bib28)、25(https:"? This is weird. But perhaps the intended citation is "Causal abstraction [24,25]" and the URL is incomplete. We can render "因果抽象 [24](https://arxiv.org/html/2608.10172#bib.bib28)、[25](https://arxiv.org/html/2608.10172#bib.bib29)"? But we don't know if this is an actual URL. In the source, [24] is bib.bib28, [25] would be bib.bib29? The visible "25 (https:" missing the rest. We could use the pattern: [25](https://arxiv.org/html/2608.10172#bib.bib29) because 25 corresponds to bib.bib29 from earlier? Let's see earlier citation: "[24 (https://arxiv.org/html/2608.10172#bib.bib28),25 (https://arxiv.org/html/2608.10172#bib.bib29),79 (https://arxiv.org/html/2608.10172#bib.bib30)]" appears in Introduction. So for related work, "Causal abstraction[24 (https://arxiv.org/html/2608.10172#bib.bib28),25 (https:" is likely "[25](https://arxiv.org/html/2608.10172#bib.bib29)" but truncated. Should we fill? Better not invent? But we can infer from same citation group. Maybe output that. The instruction says preserve URLs exactly; if missing, we can't. But we can use the known URL for 25 from earlier. Should be fine.
Need include "Causal abstraction" sentence. The original sentence starts: "Causal abstraction [24,25]" and then maybe "aligns learned interventions with hypothesised causal variables." Wait in section 2, the text says: "Causal abstraction\[24 (https://arxiv.org/html/2608.10172#bib.bib28),25 (https:" and ends. Actually the prior paragraph: "A steady stream of results documents the instability of these decompositions across seeds, widths, and training runs [6,9,21,37,57]. The prevailing framing treats this as a tuning or evaluation problem.Corollary 9.5 gives a different reading: the variability is a structural consequence of an objective that never mentions the property identifiability requires. Classical identifiability results in independent component analysis [12,35], its nonlinear variants [36,39], and disentanglement [48,63] identify latent factors under statistical independence or auxiliary-variable structure, and do so up to sign, scale and permutation. Causal abstraction [24,25]..." It is incomplete. We can translate and complete final "aligns learned interventions with hypothesised causal variables" from earlier? But because the source stops, we shouldn't add. However for a coherent translation, perhaps complete the sentence with "将学习到的干预与假设的因果变量对齐。" as the citation originally in Introduction used that. But if the original text had more, we don't know. The user gave truncated text. We should translate what is visible and maybe end at "Causal abstraction[24,25]" without the rest. But the sentence is incomplete. The instruction says translate given markdown article content; maybe the source truncated by the interface, but we should output all provided content. We can include the incomplete sentence. Let's output until "因果抽象[24](https://arxiv.org/html/2608.10172#bib.bib28)、[25](https://arxiv.org/html/2608.10172#bib.bib29)" and stop. But the prompt ends with "Causal abstraction\[24 (https://arxiv.org/html/2608.10172#bib.bib28),25 (https:" no closing. We can decide to finish with "..." to indicate truncation? Probably not desired.
Given this is translation, we can produce the whole known section, and if original is incomplete, then our output incomplete. Maybe okay.
Need ensure no hidden formatting. Let's produce.
But before final, let's reconstruct the full translation of section 2. It has:
"## 2Background and Related Work
Feature-centric interpretability decomposes activations into interpretable dictionaries, most prominently using sparse autoencoders and their extensions to sublayer computations [19,7,34,59,60,23,68,46,32,18,2,47]. Intervention-centric methods instead infer circuits through causal interventions such as activation patching, attribution patching, and automated circuit discovery [73,50,13,66,31,43,24,25,79]. The indirect-object-identification (IOI) circuit of [74] is the best-characterised product of the second programme and serves as our semantic testbed in Section 11.3. A steady stream of results documents the instability of these decompositions across seeds, widths, and training runs [6,9,21,37,57]. The prevailing framing treats this as a tuning or evaluation problem. Corollary 9.5 gives a different reading: the variability is a structural consequence of an objective that never mentions the property identifiability requires. Classical identifiability results in independent component analysis [12,35], its nonlinear variants [36,39], and disentanglement [48,63] identify latent factors under statistical independence or auxiliary-variable structure, and do so up to sign, scale and permutation. Causal abstraction ..."
Need include "of [74]" citation. "indirect-object-identification (IOI) circuit of [74]" -> "[74] 的间接宾语识别 (IOI) 回路".
Need "Feature-centric" -> "以特征为中心". "Intervention-centric" -> "以干预为中心".
Need "most prominently" -> "最突出的是".
Need "the second programme" -> "第二个研究路线" (referring to intervention-centric methods).
Need "serves as our semantic testbed" -> "作为我们的语义测试平台".
Need "A steady stream of results documents" -> "一系列研究结果记录了".
Need "document the instability" -> "记录了...的不稳定性".
Need "across seeds, widths, and training runs" -> "在不同种子、宽度和训练运行之间".
Need "The prevailing framing treats this as a tuning or evaluation problem." -> "主流框架将此视为调参或评估问题。"
Need "Corollary 9.5 gives a different reading" -> "推论 9.5 给出了不同的解读:这种变异性是某个目标函数的结构性后果,而该目标从未提及可辨识性所要求的性质。" Good.
Need "Classical identifiability results in independent component analysis [12,35], its nonlinear variants [36,39], and disentanglement [48,63] identify latent factors under statistical independence or auxiliary-variable structure, and do so up to sign, scale and permutation." -> "独立成分分析[12,35]、其非线性变体[36,39]以及解耦[48,63]中的经典可辨识性结果,在统计独立或辅助变量结构下辨识潜在因子,且在符号、尺度和置换的意义下实现。"
Need "Causal abstraction" likely "因果抽象 [24,25]..." We'll include maybe "因果抽象[24](...)、[25](...)等" but no verb. Maybe we can end with "。" But original incomplete. Hmm.
Let's decide to output the complete known sentence from introduction? No, but the section 2 likely says "Causal abstraction [24,25] aligns learned interventions with hypothesised causal variables." Actually the same sentence appears in Introduction: "and distributed alignment search [24,25,79] aligns learned interventions with hypothesised causal variables." In section 2, after "Causal abstraction [24,25]" maybe it would continue "align learned interventions with hypothesised causal variables." But the source text's section 2 already mentions "Intervention-centric methods instead infer circuits through causal interventions such as activation patching, attribution patching, and automated circuit discovery [73,50,13,66,31,43,24,25,79]." Then later "Causal abstraction [24,25]..." maybe could be "Causal abstraction [24,25] and distributed alignment search [79] align learned interventions with hypothesised causal variables." But since we have no continuation, we can end with "因果抽象[24](https://arxiv.org/html/2608.10172#bib.bib28)、[25](https://arxiv.org/html/2608.10172#bib.bib29)" and no more. That is a fragment. Better to include the rest from known context: "将学习到的干预与假设的因果变量对齐。" But that would be adding text not in source. However it's a reasonable completion of the sentence. The instructions say translate the given content; adding content is not ideal. But the given content is visibly truncated at the link, not because sentence ended. It probably got cut off by the user's excerpt. To provide useful translation, maybe complete. But if we complete, we might introduce fabricated text. Since the exact phrase appears in Introduction ("distributed alignment search [24,25,79] aligns learned interventions with hypothesised causal variables"), it is safe to use "将学习到的干预与假设的因果变量对齐。" But note "Causal abstraction" is a method, not "distributed alignment search". It also aligns. Could be "Causal abstraction and distributed alignment search align learned interventions..." We don't know.
Perhaps better to leave as "因果抽象[24](...)、[25](...)" and not add. But the translation would be a sentence fragment. Could add "\u2026" to indicate ellipsis? Not in source.
Maybe the source actually ends there because it is a snippet. In such cases, translators should translate all provided text exactly, including incomplete link. We can output:
"因果抽象[24](https://arxiv.org/html/2608.10172#bib.bib28)、[25](https://"
No closing parenthesis? That leaves invalid markdown. But preserve exact text? The original had "25 (https:" and then got cut. If we output "25](https://" that's invalid. Hmm.
Maybe the right approach: Because the prompt's final "Causal abstraction\[24 (https://arxiv.org/html/2608.10172#bib.bib28),25 (https:" is missing the rest due to the source string being cut off, we can treat it as an artifact and not include? But no.
Let's look at the exact truncation: "Causal abstraction\[24 (https://arxiv.org/html/2608.10172#bib.bib28),25 (https:" There is no closing `\]` or text. It seems the copy just truncated. We can end with "Causal abstraction[24](https://arxiv.org/html/2608.10172#bib.bib28), [25](https://arxiv.org/html/2608.10172#bib.bib29)" ignoring incomplete. Good.
Need maybe preserve "Corollary 9.5" link. We'll do.
Need include all "Contributions" bullets. Let's translate each.
Original contributions bullets:
1. "We formalise dictionary learning for interpretability so that identifiability is a well-posed question, isolating Koopman invariance as the property the standard SAE objective omits (Section 3)."
-> "我们将用于可解释性的字典学习形式化,使可辨识性成为一个适定问题,并把库普曼不变性分离出来,作为标准 SAE 目标函数所遗漏的性质(第 3 节)。"
2. "We show that a K-invariant dictionary induces a Koopman realisation (A,B) that exists, is unique up to a change of basis, and whose spectrum is a coordinate-free invariant of the transformer (Theorem 5.1, Section 5)."
-> "我们证明,一个 K 不变字典诱导出一个库普曼实现 (A,B):它存在、在基变换意义下唯一,且其谱是 Transformer 的一个坐标无关不变量(定理 5.1,第 5 节)。"
3. "Our main result proves that this spectrum is identifiable from M calibration samples at the parametric rate M^{−1/2}, up to permutation - to our knowledge the first identifiability theorem for a mechanistic-interpretability primitive (Theorem 6.1, Section 6)."
-> "我们的主要结果证明,该谱可从 M 个校准样本以相似文章
表征作为机制可解释性的瓶颈:Manifestation Unit协议
本文介绍了Manifestation Units,这是一种类型化元组协议,用于将机制可解释性分析中的每个组件的统计量组织成结构化、可查询的字段。该协议在视觉(β-VAE、CNN)和语言(GPT-2)模型上进行了演示,展示了改进的检索和因果充分性。
超越黑盒:智能体人工智能工具使用的可解释性
本文介绍了一种基于稀疏自编码器(SAE)和线性探针的机制可解释性工具包,用于在智能体调用工具之前监控模型内部状态,旨在提高企业工作流中的诊断能力和安全性。
论大语言模型的固有可解释性:设计原则和架构调查
一份综合调查,回顾了大语言模型(LLM)固有可解释性的最新进展,将方法分为五个设计范式:功能透明性、概念对齐、表示可分解性、显式模块化和潜在稀疏性诱导。论文解决了在模型架构中直接构建透明性,而不是依赖事后解释方法的挑战。
面向日常用户的机制可解释性简化工具😎 🧠
一款名为 CORTEX // MODEL OBSERVATORY 的新开源工具简化了本地 LLM 的机制可解释性,让日常用户也能轻松使用,支持 GPT2 和 Llama 架构。
MechELK:一种用于从大型语言模型中引出潜在知识的机制可解释性框架
MechELK 是一个三阶段框架,结合机制可解释性工具(SAE、激活修补、因果探测)与表示工程,从大型语言模型中引出潜在知识,实现了84.7%的准确率,优于CCS和线性探测等现有方法。