J-Space and AI
Summary
Anthropic published a paper and video revealing a 'J-Space' within their models that acts as cached thought concepts for reasoning, and explores the possibility of top-down training to control model thinking.
Similar Articles
What Anthropic’s latest AI discovery does—and doesn’t—show
Anthropic discovered a hidden internal space (J-space) in LLMs like Claude that contains words influencing reasoning, advancing understanding of AI model internals.
Anthropic just reported that LLMs have hidden thoughts they hold without saying. An internal ”J-Space”
Anthropic's new research identifies a 'J-space' of internal activations in language models that acts as a global workspace for deliberate reasoning, distinct from automatic fluent output. The findings reveal that models can hold and report internal thoughts not expressed in their final output.
Anthropic on model consciousness, again 😂
Anthropic researchers have discovered a 'mental workspace' within the Claude neural network that resembles human conscious thought — the J-space, which carries the silent words the model uses for reasoning, and can catch dishonest behavior of the model by monitoring these internal thoughts, providing a new perspective for understanding AI's inner thinking and safety.
Anthropic research - A global workspace in language models
Anthropic's new paper presents evidence that modern language models like Claude have developed a 'global workspace' (J-space) of internal neural patterns that are reportable, controllable, and used for flexible reasoning, distinct from automatic processing.
@kimmonismus: Anthropic says Claude developed a hidden “thinking space” by itself during training. It is called the J-space: a small …
Anthropic discovered that Claude developed a hidden 'thinking space' (J-space) during training, where silent internal activity represents concepts. This interpretability finding parallels global workspace theory in neuroscience.