dictionary-learning

Tag

Cards List
#dictionary-learning

Driving the Wrong Way: Leveraging Interpretability in End2End Autonomous Driving Models

arXiv cs.AI · 2026-07-08 Cached

This paper introduces a concept-based interpretability framework for end-to-end autonomous driving models, using unsupervised dictionary learning via sparse autoencoders to decompose driving behavior into human-interpretable concepts. The framework enables analysis and targeted correction of model decisions, improving overall driving performance.

0 favorites 0 likes
#dictionary-learning

Size Doesn't Matter: Cosine-Scored Sparse Autoencoders

arXiv cs.LG · 2026-06-16 Cached

This paper proposes replacing the inner product scoring in sparse autoencoders with a learned combination of cosine similarity and input magnitude, showing that the resulting features are more interpretable and concept-aligned, with the optimizer consistently preferring cosine over inner product.

0 favorites 0 likes
#dictionary-learning

Diverse Dictionary Learning

Hugging Face Daily Papers · 2026-04-19 Cached

The paper introduces diverse dictionary learning, showing that key set-theoretic relationships among latent variables can be identified from observational data without strong assumptions, enabling partial or full identifiability with minimal inductive bias.

0 favorites 0 likes
← Back to home

Submit Feedback