model-activations

Tag

Cards List
#model-activations

Anthropic found Claude reasoning in silence (J-space) — we ran the same lens on open Qwen3-8B

Reddit r/LocalLLaMA · 2026-07-12

Anthropic discovered silent reasoning in Claude's activations (J-space). The author applied the same Jacobian lens to Qwen3-8B locally, using it to detect prose drift before tool calls and implement agent guards.

0 favorites 0 likes
#model-activations

@IbrahimDagher20: So first the J-space discovery, and now a massive improvement in probing. Mech interp is clearly getting faster, in par…

X AI KOLs Timeline · 2026-07-08 Cached

GoodfireAI introduces Block-Sparse Featurizers (BSFs), a new research method for finding concepts in model activations using multidimensional blocks instead of single directions.

0 favorites 0 likes
← Back to home

Submit Feedback