Tag
The BABEL codec achieves the first complete decode of GPT-2 small's internal state, reconstructing 94.7% of its behavior and enabling reading and writing to the model in English, with open source resources and a demo.
Anthropic's new research identifies a 'J-space' of internal activations in language models that acts as a global workspace for deliberate reasoning, distinct from automatic fluent output. The findings reveal that models can hold and report internal thoughts not expressed in their final output.