activation-geometry

Tag

Cards List
#activation-geometry

Unsupervised Features Mining via Activation Geometry

arXiv cs.AI · 2026-07-07 Cached

This paper introduces Mining via Activation Geometry (MAG), an unsupervised framework that extracts reasoning features from LLM activations using natural-language instructions, enabling activation steering and effective training data selection for classifier probes.

0 favorites 0 likes
#activation-geometry

ICA Lens: Interpreting Language Models Without Training Another Dictionary

Hugging Face Daily Papers · 2026-06-10 Cached

ICA Lens revives independent component analysis as an efficient method for interpreting language model representations, offering a faster alternative to sparse autoencoder training while maintaining competitive performance.

0 favorites 0 likes
#activation-geometry

When Is Rank-1 Steering Cheap? Geometry, Granularity, and Budgeted Search

arXiv cs.LG · 2026-05-19 Cached

This paper investigates when rank-1 activation steering is effective and cost-efficient, proposing geometry-guided search and the concept of granularity to explain variability, and introduces the GRACE framework for efficient LLM control.

0 favorites 0 likes
← Back to home

Submit Feedback