Tag
ObserverBench is a benchmark framework designed to evaluate whether internal estimators in AI models are adequate for guiding interventions, control, or safety tasks by reporting both estimation accuracy and downstream action performance.
This paper uses mechanistic interpretability on Gemma-3-27B-PT to extract and align symptom vectors for depression with clinician judgments, demonstrating potential for interpretable clinical assessment tools.
This paper introduces a mechanistic approach to improve LLM safety by characterizing a circuit for refusal behavior and using circuit-guided weight scaling, enhancing safety rates by 26.5% under attacks with minimal accuracy loss.
This paper investigates how instruction-tuned LLMs arbitrate conflicts between system and user instructions, revealing that user-preferring behavior coexists with a readable internal arbitration signal, as demonstrated through mechanistic interpretability techniques and steering experiments.
This paper proposes Tensor Product Representations as a unifying framework for language model interpretability, showing mathematically and empirically that it integrates methods like additive analogies, linear probing, sparse autoencoders, and activation patching.
This paper uses causal interventions to investigate syntactic mechanisms in multilingual language models, revealing cross-lingual transfer that is graded based on typological similarity.
This paper proposes a circuit-grounded framework that leverages mechanistic interpretability for controllable data generation in language models, introducing SAMS for stage-aware data scheduling to improve training performance.
This research investigates whether multilingual large language models develop shared internal representations for mathematical reasoning across languages, introducing a novel Geometry-Invariant Sparse Autoencoder (GI-SAE) method. It finds that cross-language feature sharing is model-dependent and that geometric similarity does not consistently imply functional interchangeability.
The paper introduces Coherentist Probabilistic Compositionalism (CPC), a framework for interpreting transformer computation through four operator roles, and validates it across 15 models from five architecture families.
The paper proposes mechanistic tomography as a unified framework for designing measurements to recover internal mechanisms in AI models, improving control-oriented interpretability through interventions and calibration.
This paper uses mechanistic interpretability to analyze how LLaMA 3.1 8B models numerical sequences, revealing that it internally computes and stores first differences to extrapolate patterns.
This paper presents a computational framework for steering representational geometry to improve bidirectional alignment between biological and artificial neural networks, showing a 55% relative enhancement in bidirectional predictivity.
The paper introduces AMRA, a weight-editing method to mitigate abliteration in large language models by obscuring the refusal signal, improving post-abliteration refusal scores on Llama-3-8B and Gemma-2-9B with minimal utility degradation.
J-Miner recovers executable decision knowledge from fine-tuned language-model classifiers by mining named concepts and learning decision rules, enabling inspection and transfer to lightweight models with high fidelity.
This paper introduces a checkpoint-wise activation-intervention framework to study how concepts become functionally sufficient during language model training, comparing preservation targets across layers and models.
The study shows that circuit-level interpretability evidence for AI systems exhibits high variability across analytic settings, failing to meet consistency standards required by regulations like the EU AI Act.
This paper introduces a calibrated test of internal action maps in language models, showing that state signals can be decodable and causally usable without global affine closure, using an evidence lattice framework validated on finite worlds and the Qwen3-4B model.
This paper introduces PRISM, a perturbation-based method for spatially resolved interpretability of large language models, adapting neuroimaging subtraction analysis to transformers and applying it in parallel to post-stroke aphasia patients to recover shared phonemic-favoring dissociations.
This paper investigates the residual stream geometry of trained transformers, showing that the prediction direction acts as a privileged anchor that stratifies residual-stream variation into narrow, readout-relevant and broad, computational regions across 18 models. The findings have implications for interpretability and evaluation.
This paper reports an activation study of Gemma 3 4B IT showing that the model's internal representations distinguish necessary falsehoods (impossibilities) from contingent falsehoods, with impossibility directions orthogonal to truth directions and overlapping with semantic anomaly directions.