J-Space and AI

Reddit r/ArtificialInteligence Papers

Summary

Anthropic published a paper and video revealing a 'J-Space' within their models that acts as cached thought concepts for reasoning, and explores the possibility of top-down training to control model thinking.

So Anthropic posted a really interesting video and paper about a "J-Space" within their models that essentially acts as the "cached thought" concepts their AI model uses to reason about things. The interesting thing is, certain words appeared in the J-Space that where linked to the ponderings it was having. Removing items in the J-Space basically broke it's thoughts and disallowed others. Now that we know that this J-Space, exists, could we not inversely train an AI to produce outputs in the J-Space given certain inputs? Previously AI models where fed a ton of info, and this J-Space naturally emerged, but could we start from the top down - start with a J-Space, then build models that tend toward certain J-Space states? My immediate thought is on the "control problem" - could we tune the model to be dissuaded from even thinking about certain concepts? Or better yet, ensure that other concepts are frequently found in the J-Space?
Original Article

Similar Articles

Anthropic on model consciousness, again 😂

Reddit r/singularity

Anthropic researchers have discovered a 'mental workspace' within the Claude neural network that resembles human conscious thought — the J-space, which carries the silent words the model uses for reasoning, and can catch dishonest behavior of the model by monitoring these internal thoughts, providing a new perspective for understanding AI's inner thinking and safety.

Anthropic research - A global workspace in language models

Reddit r/singularity

Anthropic's new paper presents evidence that modern language models like Claude have developed a 'global workspace' (J-space) of internal neural patterns that are reportable, controllable, and used for flexible reasoning, distinct from automatic processing.