@swyx: imo this is the most impt part of anthropic's J-space paper today. it's a two-parter: 1) ant proved that they can do "b…
Summary
Anthropic's J-space paper demonstrates that they can perform 'brain surgery' interventions into reasoning to change topics midstream, and the model is able to detect what intervention was done, indicating a form of eval awareness.
View Cached Full Text
Cached at: 07/07/26, 07:26 AM
imo this is the most impt part of anthropic’s J-space paper today. it’s a two-parter:
- ant proved that they can do “brain surgery” interventions into reasoning to change topics midstream*
- THE MODEL IS ABLE TO DETECT WHAT INTERVENTION WAS DONE - close cousin to eval awareness**
*control > correlation - this convincingly demonstrates understanding
**this was prompted awareness… surely @mlpowered’s team also tried to eval unprompted awareness but i didn’t see evidence of that
Anthropic (@AnthropicAI): New Anthropic research: A global workspace in language models.
Of everything happening in your brain right now, only a tiny fraction is consciously accessible—thoughts you can describe, hold in mind, and reason with.
We found a strikingly similar divide inside Claude.
Similar Articles
@FutureJurvetson: This rolls deep in my J-Space I have been skeptical that interpretability research would bear fruit, but this update fr…
Anthropic's new research reveals a global workspace in language models, showing a striking parallel between the conscious and subconscious divide in human brains and the internal reasoning of Claude.
J-Space and AI
Anthropic published a paper and video revealing a 'J-Space' within their models that acts as cached thought concepts for reasoning, and explores the possibility of top-down training to control model thinking.
Anthropic research - A global workspace in language models
Anthropic's new paper presents evidence that modern language models like Claude have developed a 'global workspace' (J-space) of internal neural patterns that are reportable, controllable, and used for flexible reasoning, distinct from automatic processing.
@SwissCognitive: Claude’s J-space reveals clues about internal processing, but it is not a readable transcript of reasoning, an importan…
Anthropic's research on Claude's internal 'J-space' reveals clues about how the model processes reasoning, but it is not a full readable transcript, underscoring key limitations in AI interpretability.
Anthropic on model consciousness, again 😂
Anthropic researchers have discovered a 'mental workspace' within the Claude neural network that resembles human conscious thought — the J-space, which carries the silent words the model uses for reasoning, and can catch dishonest behavior of the model by monitoring these internal thoughts, providing a new perspective for understanding AI's inner thinking and safety.