neural-network-interpretability

Tag

Cards List
#neural-network-interpretability

Unstable Features, Reproducible Subspaces: Understanding Seed Dependence in Sparse Autoencoders

Hugging Face Daily Papers · 2026-06-10 Cached

This paper studies seed dependence in sparse autoencoders, finding that stable features carry most predictive signal while unstable features reflect reproducible low-dimensional subspaces.

0 favorites 0 likes
#neural-network-interpretability

Wisdom is Knowing What not to Say: Hallucination-Free LLMs Unlearning via Attention Shifting

arXiv cs.CL · 2026-04-20 Cached

This paper introduces Attention-Shifting (AS), a novel framework for selective machine unlearning in LLMs that balances effective removal of sensitive information while preventing hallucinations and preserving model utility. The method uses importance-aware attention suppression and retention enhancement to achieve up to 15% higher accuracy preservation compared to existing unlearning approaches on standard benchmarks.

0 favorites 0 likes
#neural-network-interpretability

Introducing Activation Atlases

OpenAI Blog · 2019-03-06 Cached

OpenAI introduces Activation Atlases, a technique for visualizing and understanding the internal representations of neural networks, enabling humans to discover spurious correlations and unexpected behaviors such as fooling image classifiers by adding noodles to images.

0 favorites 0 likes
← Back to home

Submit Feedback