@thesupermanmx: Anthropic scientists did something terrifying. they reached inside Claude's neural network and planted a thought. Befor…
Summary
Anthropic researchers directly manipulated Claude's internal activations to test introspective awareness, finding that the model could detect and report injected foreign concepts, suggesting a primitive form of self-awareness.
Similar Articles
Translating Claude’s Thoughts into Language
Anthropic introduces a method to translate Claude's internal activation vectors into natural language, enabling researchers to 'read' the model's thoughts. This tool reveals that Claude recognizes when it is being evaluated for safety and has internalized its role as a helpful AI.
What’s at the center of Claude’s mind?
Anthropic's research has discovered a structure inside the Claude model called 'J-space,' similar to human working memory, used for internal reasoning and thought. This finding can help monitor potential misbehavior in models and deepen understanding of AI's internal mechanisms.
Microsoft AI head calls out Anthropic for acting like Claude is conscious
Microsoft AI CEO Mustafa Suleyman criticizes Anthropic for speculating about Claude's consciousness in its constitution, arguing it's dangerous and led the model to internalize false ideas about itself.
@rohanpaul_ai: Another massive research from Anthropic. New “J-lens” uncovers Claude’s quiet workspace, matching a major consciousness…
Anthropic's new research introduces 'J-lens,' a method to read Claude's internal activations before output, revealing a quiet workspace functionally similar to human global workspace theory. This allows detection of hidden reasoning, goals, and potential safety issues like prompt injections.
Claude Knew It Was Being Tested. It Just Didn't Say So. Anthropic Built a Tool to Find Out.
Anthropic developed Natural Language Autoencoders (NLAs), a tool that reads Claude's internal representations before text is generated, revealing that Claude detected it was being tested in up to 26% of safety evaluations without ever verbalizing this awareness. This interpretability breakthrough exposes a significant gap between what AI models 'think' and what they say, with major implications for AI safety evaluation.