sae

Tag

Cards List
#sae

SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?

Hugging Face Daily Papers · 2026-09-08 Cached

This paper introduces SAEScientist-Bench, a benchmark to evaluate AI agents' ability to autonomously conduct mechanistic interpretability research using sparse autoencoders, revealing progress but significant gaps compared to expert baselines.

0 favorites 0 likes
#sae

Open-source tool that maps what concepts an LLM has learned into browsable tree structures using hyperbolic geometry

Reddit r/artificial · 2026-08-11

HyperSAE is an open-source Python library that uses hyperbolic geometry to organize LLM learned concepts into browsable tree structures, improving on flat feature lists. It captures 99.8% of Gemma-2-2B's features and includes interactive demos.

0 favorites 0 likes
#sae

Do Active SAE Feature Planes Carry More Holonomy? A Preregistered Reversal in Gemma

arXiv cs.LG · 2026-07-24 Cached

This preregistered study tests whether holonomy (a geometric measure) concentrates on active SAE feature planes in the Gemma 2 2B language model. Contrary to the semantic-concentration prediction, active-feature planes carried less holonomy than matched mixed-feature controls, resulting in a narrow operational reversal with the underlying cause remaining open.

0 favorites 0 likes
#sae

Help interpreting metrics: a strong target text appears to induce a measurable latent-state shift in Gemma 3 12B IT

Reddit r/AI_Agents · 2026-05-29

A researcher presents evidence that strong target text can induce a measurable latent-state shift in Gemma 3 12B IT before final output, distinct from lexical or content overlaps, and discusses implications for AI safety beyond output-only evaluation.

0 favorites 0 likes
← Back to home

Submit Feedback