neural-network-interpretability

Tag

Cards List
#neural-network-interpretability

NObSP: Functional Decomposition of Neural Networks via Oblique Subspace Projections

arXiv cs.LG ↗ · 2026-09-17 Cached

This paper introduces NObSP, a framework for decomposing neural network predictions into per-feature contribution functions and interaction residuals using oblique subspace projections, improving interpretability and reducing attribution errors compared to existing methods.

0 favorites 0 likes
#neural-network-interpretability

A Mathematical Framework for Transformer Circuits (2021)

Hacker News Top ↗ · 2026-09-12 Cached

This paper presents a mathematical framework for reverse-engineering transformer models to enhance mechanistic interpretability and address safety concerns in AI systems.

0 favorites 0 likes
#neural-network-interpretability

How Perturbations Propagate: A Multi-Level Analysis of Robustness in Large Language Models

arXiv cs.CL ↗ · 2026-09-04 Cached

This paper conducts a multi-level analysis of how input perturbations propagate through decoder-only language models, assessing robustness via output behavior, hidden-state geometry, and attention-head function across models like GPT-2 and Qwen2.5.

0 favorites 0 likes
#neural-network-interpretability

Unstable Features, Reproducible Subspaces: Understanding Seed Dependence in Sparse Autoencoders

Hugging Face Daily Papers ↗ · 2026-06-10 Cached

This paper studies seed dependence in sparse autoencoders, finding that stable features carry most predictive signal while unstable features reflect reproducible low-dimensional subspaces.

0 favorites 0 likes
#neural-network-interpretability

Wisdom is Knowing What not to Say: Hallucination-Free LLMs Unlearning via Attention Shifting

arXiv cs.CL ↗ · 2026-04-20 Cached

This paper introduces Attention-Shifting (AS), a novel framework for selective machine unlearning in LLMs that balances effective removal of sensitive information while preventing hallucinations and preserving model utility. The method uses importance-aware attention suppression and retention enhancement to achieve up to 15% higher accuracy preservation compared to existing unlearning approaches on standard benchmarks.

0 favorites 0 likes
#neural-network-interpretability

Introducing Activation Atlases

OpenAI Blog ↗ · 2019-03-06 Cached

OpenAI introduces Activation Atlases, a technique for visualizing and understanding the internal representations of neural networks, enabling humans to discover spurious correlations and unexpected behaviors such as fooling image classifiers by adding noodles to images.

0 favorites 0 likes
← Back to home

Submit Feedback