Tag
This paper introduces NObSP, a framework for decomposing neural network predictions into per-feature contribution functions and interaction residuals using oblique subspace projections, improving interpretability and reducing attribution errors compared to existing methods.
This paper presents a mathematical framework for reverse-engineering transformer models to enhance mechanistic interpretability and address safety concerns in AI systems.
This paper conducts a multi-level analysis of how input perturbations propagate through decoder-only language models, assessing robustness via output behavior, hidden-state geometry, and attention-head function across models like GPT-2 and Qwen2.5.
This paper studies seed dependence in sparse autoencoders, finding that stable features carry most predictive signal while unstable features reflect reproducible low-dimensional subspaces.
This paper introduces Attention-Shifting (AS), a novel framework for selective machine unlearning in LLMs that balances effective removal of sensitive information while preventing hallucinations and preserving model utility. The method uses importance-aware attention suppression and retention enhancement to achieve up to 15% higher accuracy preservation compared to existing unlearning approaches on standard benchmarks.
OpenAI introduces Activation Atlases, a technique for visualizing and understanding the internal representations of neural networks, enabling humans to discover spurious correlations and unexpected behaviors such as fooling image classifiers by adding noodles to images.