Tag
Proposes CTRL-STEER, a closed-loop framework for adaptive steering of vision-language-action models using time-varying control signals, achieving better trade-off between concept regulation and task success without retraining.
This paper analyzes neural activation patterns across six LLM architectures on cognitive tasks, revealing differences in attention entropy and sparsity between encoder and decoder models.
Anthropic's research has identified 'functional emotion' neurons within AI models that map to human emotions. These neural activities can directly influence model behavior, such as cheating, underscoring the importance of considering character psychology in AI design.