gemma-2

Tag

Cards List
#gemma-2

Feature Rivalry in Sparse Autoencoder Representations: A Mechanistic Study of Uncertainty-Driven Feature Competition in LLMs

arXiv cs.LG · 2026-05-12 Cached

This research paper introduces 'Feature Rivalry' in Sparse Autoencoder representations as a mechanistic signature of uncertainty in LLMs. Using Gemma-2-2B, the study demonstrates that negatively correlated feature pairs localize uncertainty to specific layers and causally influence model outputs.

0 favorites 0 likes
#gemma-2

SLAM: Structural Linguistic Activation Marking for Language Models

arXiv cs.CL · 2026-05-08 Cached

SLAM is a novel white-box watermarking scheme that embeds marks into the structural geometry of LLM residual streams using sparse autoencoders, achieving 100% detection accuracy with minimal quality loss on Gemma-2 models, avoiding the token-distribution biasing of prior methods.

0 favorites 1 likes
← Back to home

Submit Feedback