activation-control

Tag

Cards List
#activation-control

Measuring Activation Control in Large Language Models

arXiv cs.AI · 16h ago Cached

This paper introduces the Activation Controllability Benchmark to measure how well large language models can modulate their residual stream via natural-language instructions, finding that most models can do so to some extent, which could evade activation-based monitoring methods and pose risks for AI safety.

0 favorites 0 likes
← Back to home

Submit Feedback