mechanistic-interpretability

Tag

Cards List
#mechanistic-interpretability

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training

arXiv cs.LG · 4h ago Cached

The paper introduces CD-RFT, a method to decouple the shared control bottleneck in RL post-training by regularizing a novel control coefficient, improving multi-task capability on models like Qwen2.5-7B and Llama-3.2-3B.

0 favorites 0 likes
#mechanistic-interpretability

Transformers are famously bad at arithmetic, so I set one's weights by hand (no training) and it multiplies with 100% accuracy [P]

Reddit r/MachineLearning · 15h ago

The author hand-codes transformer weights (no training) using a compiler called Torchwright to implement exact multiplication, achieving 100% accuracy on three-digit math and publishing checkpoints that handle up to 12-digit multiplication.

0 favorites 0 likes
#mechanistic-interpretability

Why Knowing Both Hops Is Not Enough: Understanding Two-Hop Generalization in Language Models

arXiv cs.CL · yesterday Cached

This paper investigates why language models fail at two-hop generalization, showing that models succeed when the second hop follows training distribution but fail when it deviates, and proposes a recurrent-style training strategy to improve out-of-distribution two-hop reasoning.

0 favorites 0 likes
#mechanistic-interpretability

Finding Usable Weight Mechanisms with Tiled SVD

arXiv cs.AI · yesterday Cached

This paper proposes extracting mechanism mounts directly from linear weight sites via column-tiled SVD, offering an alternative to proxy dictionaries like sparse autoencoders for mechanistic interpretability. Evaluated on Gemma-2-2B, the method passes all 182 site-layer checks.

0 favorites 0 likes
#mechanistic-interpretability

@Redpoint: Why did a spend management platform start its own AI research lab? @karimatiyeh describes @RampLabs as a “collection of…

X AI KOLs Following · 3d ago Cached

Ramp, a spend management platform, launched its own AI research lab called Ramp Labs a year ago. The lab has worked on projects like a production-focused coding benchmark 'Ramp SWE-Bench', integrating Claude Code into RollerCoaster Tycoon, and a mechanistic interpretability playground.

0 favorites 0 likes
#mechanistic-interpretability

Cross-Architecture Steering Transfer in Language Models: A Systematic Empirical Study

arXiv cs.CL · 4d ago Cached

A systematic empirical study showing that concept directions extracted from one language model can steer other independently trained models when sufficient scale (≥1.7B parameters) is reached, providing functional evidence for the Platonic Representation Hypothesis and highlighting scale thresholds for cross-model interpretability tools.

0 favorites 0 likes
#mechanistic-interpretability

Safe Evolution with Circuit Anchors

arXiv cs.CL · 4d ago Cached

This paper proposes Circuit-Anchored Evolution (CAE), a method that uses mechanistic interpretability to identify and anchor a tiny safety circuit in LLMs during self-evolution, preventing models from misevolving into capable but dangerous systems while preserving capability.

0 favorites 0 likes
#mechanistic-interpretability

Subliminal Learning is Non-Semantic Distillation

arXiv cs.AI · 4d ago Cached

This paper investigates subliminal learning in language models, showing that biases can transfer from teacher to student via seemingly random synthetic data. The authors find that adding Gaussian noise to weights increases transfer, and that students inherit not just the semantic bias but also the type of intervention used, with implications for training safety and data auditing.

0 favorites 0 likes
#mechanistic-interpretability

The Ignition Index: Measuring Global Workspace Dynamics in Language Models

arXiv cs.AI · 4d ago Cached

The paper introduces the Ignition Index, a metric for measuring global workspace dynamics in language models, validated across multiple architectures and tasks, showing selective detection of ignition-like representational transitions.

0 favorites 0 likes
#mechanistic-interpretability

Test, then Route: How Language Models Execute In-Context Conditional Rules Across Models and Languages

arXiv cs.CL · 5d ago Cached

This paper investigates how language models execute in-context conditional rules by probing whether testing and routing are separable mechanisms. Using activation patching across three open models and six languages, the authors find that predicate testing is modular while route representations are token-bound and non-transferable.

0 favorites 0 likes
#mechanistic-interpretability

Visualizing Graph-to-Answer Mechanism Recovery in Materials-Science Hypothesis Generation

arXiv cs.CL · 5d ago Cached

This paper presents a graph-to-answer mechanism-tracing case study for Graph-PRefLexOR-8B, a materials-science hypothesis generation model, using visualization and activation-based diagnostics to localize where mechanism support is lost or recovered during generation.

0 favorites 0 likes
#mechanistic-interpretability

DUD: Decoupled Update Dynamics for Reliable Uncertainty Quantification in Large Language Models

arXiv cs.CL · 6d ago Cached

The paper introduces DUD (Decoupled Update Dynamics), a framework that separates Feed-Forward Network and Attention contributions via causal interventions to improve uncertainty quantification and calibration in large language models, outperforming state-of-the-art baselines.

0 favorites 0 likes
#mechanistic-interpretability

LLMs Can Annotate Attribution Graphs

arXiv cs.LG · 6d ago Cached

The paper presents a simple pipeline that uses LLMs to automatically group features into supernodes in attribution graphs, matching human annotator interpretability and recovering intermediate hop supernodes in a two-hop task.

0 favorites 0 likes
#mechanistic-interpretability

Bridging the English-Arabic Medical Knowledge Gap: Targeted Low-Rank Adaptation via Causal Layer Selection

arXiv cs.CL · 2026-08-04 Cached

This paper investigates why LLMs underperform in Arabic medical tasks, showing via mechanistic analysis that knowledge exists internally but fails to surface, then proposes TLoRA, a targeted low-rank adaptation method that outperforms full-network LoRA on medical QA and introduces a new Arabic clinical dialogue benchmark.

0 favorites 0 likes
#mechanistic-interpretability

LAWFUL: Law-Aligned Witness for Faithful Use of Latents

arXiv cs.LG · 2026-08-03 Cached

This paper introduces LAWFUL, a framework for verifying whether neural networks learn and causally use physical laws over continuous variables, addressing gaps in coverage-aware causal consistency and domain-of-validity testing. It demonstrates the approach on a Mocap2Radar transformer and the Doppler frequency law.

0 favorites 0 likes
#mechanistic-interpretability

Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations

arXiv cs.CL · 2026-07-31 Cached

This paper presents Fairness Pruning, a structural intervention method that locates demographic bias in GLU-MLP layers of large language models by identifying differentially activated neurons. Zeroing a very small number of neurons disrupts bias processing while preserving reasoning and general knowledge, suggesting bias and capabilities rely on dissociable circuits.

0 favorites 0 likes
#mechanistic-interpretability

ECG-InterpBench: Benchmarking the Interpretability of ECG Foundation Models with Matched-Scale Sparse Autoencoders

arXiv cs.LG · 2026-07-31 Cached

ECG-InterpBench is a new benchmark that systematically evaluates the interpretability of ECG foundation model representations using matched-scale sparse autoencoders, covering reconstruction fidelity, clinical concept accessibility, and reproducibility across 450 cells.

0 favorites 0 likes
#mechanistic-interpretability

The Kinetics of Training: A Driven-Nucleation Rate Law for Emergence, Plasticity Loss, and Circuit Control in Language Models

arXiv cs.LG · 2026-07-31 Cached

This theoretical paper proposes a driven-nucleation rate law to explain capability emergence, plasticity loss, and circuit control in language models, supported by experiments on Pythia and a controlled gated-attention model.

0 favorites 0 likes
#mechanistic-interpretability

Probing the Origins of Reasoning Performance: Representational Quality for Mathematical Problem-Solving in RL vs. SFT Fine-Tuned Models

arXiv cs.AI · 2026-07-31 Cached

This paper investigates internal representational differences between RL and SFT fine-tuned models on mathematical reasoning, finding that RL models exhibit more linearly separable hidden states and hierarchical layer importance. Token allocation variability under repeated sampling suggests training pipeline dependence rather than RL vs SFT alone.

0 favorites 0 likes
#mechanistic-interpretability

Mechanistic interpretability streamlined for everyday users like us😎 🧠

Reddit r/LocalLLaMA · 2026-07-30

A new open-source tool called CORTEX // MODEL OBSERVATORY streamlines mechanistic interpretability for local LLMs, making it accessible to everyday users, with support for GPT2 and Llama architectures.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback