causal-intervention

Tag

Cards List
#causal-intervention

Understanding Decision-Making Mechanisms in Neural Routing Solvers

arXiv cs.LG ↗ · 5d ago Cached

该论文通过行为分析、表示探查和因果干预,研究了AM、POMO和LEHD等神经组合优化(NCO)路由求解器的内部决策机制,揭示了不同架构在构造解决方案时的差异化模式,例如LEHD依赖当前节点表示进行局部决策、起始节点提供全局导航参考。

0 favorites 0 likes
#causal-intervention

FTB Graph: Determining and Validating First-token Broadcasters and Language-Identity Head Circuits in Multilingual Language Models

arXiv cs.AI ↗ · 2026-09-28 Cached

The paper presents a structural circuit analysis to identify first-token broadcasters and language-identity heads in multilingual language models, revealing consistent hub topologies and that routing circuitry is primarily established during pre-training.

0 favorites 0 likes
#causal-intervention

Knowing, and Saying It Only When Asked: LLM Endognostics and the Schizognosis of Minerva-7B

arXiv cs.CL ↗ · 2026-09-22 Cached

This paper introduces LLM endognostics, a white-box framework for auditing latent knowledge in language models, showing that behavioral evaluations can fail to reflect internal knowledge distinctions, as demonstrated on Minerva-7B-Instruct-v1.0.

0 favorites 0 likes
#causal-intervention

Encoded Early, Used Late: Where Transformers Begin to Act on an Inferred Partner's Expertise

Hugging Face Daily Papers ↗ · 2026-09-07 Cached

This paper demonstrates that transformers encode a dialogue partner's expertise in early layers but activate it causally only in later layers, revealing a gap that constrains intervention points for steering model behavior in multi-turn dialogues.

0 favorites 0 likes
#causal-intervention

A Calibrated Test of Internal Action Maps: State Signals Without Global Affine Closure

arXiv cs.AI ↗ · 2026-08-17 Cached

This paper introduces a calibrated test of internal action maps in language models, showing that state signals can be decodable and causally usable without global affine closure, using an evidence lattice framework validated on finite worlds and the Qwen3-4B model.

0 favorites 0 likes
#causal-intervention

Steering the Language Axis: From Linear Decodability to Causal Control

arXiv cs.CL ↗ · 2026-08-14 Cached

This paper investigates whether language identity in LLMs is linearly decodable and causally controllable via compact activation directions. Through steering and ablation experiments across multiple model families, the authors show that language selection is direction-dependent, layer-specific, and reverts to English when the language signal is ablated.

0 favorites 0 likes
#causal-intervention

Vision-Language Models are Fragile Multilingual Associators

arXiv cs.CL ↗ · 2026-08-14 Cached

This paper introduces M2BIND, a benchmark to evaluate whether vision-language models maintain stable visual-linguistic associations across languages. It finds that binding is not language-invariant, with cross-family and cross-script settings causing significant performance collapse and weaker internal causal binding.

0 favorites 0 likes
#causal-intervention

Decodable But Not Detachable: Training Data Granularity Determines Parametric Modularity in Large Language Models

arXiv cs.AI ↗ · 2026-08-12 Cached

This paper investigates when large language models develop domain-specific parametric shells (causally necessary neuron populations), finding that modular training data at the token level (e.g., languages, code) produces functional shells, while academic subject domains do not, despite being linearly decodable.

0 favorites 0 likes
#causal-intervention

Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence

Hugging Face Daily Papers ↗ · 2026-08-12 Cached

Introduces Mechanist, an autonomous agentic system that uses AI to discover and control the mechanisms underlying model intelligence, generating hypotheses, performing causal interventions, and improving safety and performance.

0 favorites 0 likes
#causal-intervention

CausalGate: Causal Importance Distillation for Transformer Module Pruning

arXiv cs.LG ↗ · 2026-07-28 Cached

CausalGate introduces a method that uses causal interventions to measure the importance of transformer sub-layers and distills this into static scalar gates for efficient inference without runtime overhead, outperforming existing pruning and routing methods.

0 favorites 0 likes
#causal-intervention

Attention Degradation, Function Token Anchoring, and the Limits of Attention-Based Intervention in Large Language Models

arXiv cs.AI ↗ · 2026-07-24 Cached

This paper investigates short-term attention degradation in LLMs, finding a universal exponential-then-plateau pattern and that function token anchoring is architecture-dependent. Causal tests show that increasing attention mass on function tokens does not improve retrieval, suggesting attention degradation is descriptive rather than prescriptive.

0 favorites 0 likes
#causal-intervention

Visual Access Boundaries in Vision-Language Model Reasoning

arXiv cs.AI ↗ · 2026-07-15 Cached

This paper introduces Visual Access Sweep, a causal intervention method to measure the minimal image-token access needed for Vision-Language Model reasoning, and finds that Chain-of-Thought prompting does not primarily improve performance by prolonging direct image access but by enabling extended language-side computation over visual information.

0 favorites 0 likes
#causal-intervention

Final Checkpoints Are Not Enough: Analyzing Latent Reasoning Faithfulness Along Training Trajectories

arXiv cs.LG ↗ · 2026-07-09 Cached

This paper investigates how the faithfulness of latent reasoning steps evolves during training, finding that it depends on training stage and answer format, rather than just final checkpoint performance.

0 favorites 0 likes
#causal-intervention

Attending to Multimodal Generation One Token at a Time

Hugging Face Daily Papers ↗ · 2026-07-04 Cached

This paper investigates token-level attention shifts in multimodal large language models during generation, revealing consistent patterns and proposing a simple test-time intervention that significantly improves task performance.

0 favorites 0 likes
#causal-intervention

Dismantling Pathological Shortcuts: A Causal Framework for Faithful LVLM Decoding

arXiv cs.AI ↗ · 2026-06-29 Cached

This paper reveals that hallucination in large vision-language models is caused by a dynamic structural misalignment where certain attention heads act as risky mediators, decoupling from visual evidence to lock onto language priors. The authors propose Fox, a training-free causal intervention framework that diagnoses and physically severs these pathological shortcuts, achieving state-of-the-art performance in faithful decoding.

0 favorites 0 likes
#causal-intervention

The Weight Norm Sets the Grokking Timescale: A Causal Delay Law

arXiv cs.LG ↗ · 2026-06-15 Cached

This paper demonstrates that the weight norm causally controls the timescale of grokking in neural networks, reconciling conflicting accounts. Through interventions, it shows that grokking follows an exponential delay law and that norm magnitude dominates grokking time over learning rate across architectures.

0 favorites 0 likes
#causal-intervention

Causal Evidence for Attention Head Imbalance in Modality Conflict Hallucination

arXiv cs.AI ↗ · 2026-05-20 Cached

This paper identifies imbalanced attention head groups in MLLMs that drive or resist modality-conflict hallucination, and proposes MACI, a causal intervention that suppresses hallucination-driving heads only when conflict is detected, achieving large hallucination reduction across five models.

0 favorites 0 likes
← Back to home

Submit Feedback