model-control

Tag

Cards List
#model-control

Inverted Detection and Control in Steering Vectors

arXiv cs.LG · 2026-08-05 Cached

This paper identifies an 'inverted detection-control' phenomenon where some discriminative steering vectors, despite aligning with positive concept representations, consistently promote the opposite behavior. The authors propose a method to detect such inverted steering vectors without generation, enabling sign flips that improve a detection-based steering pipeline across multiple LLMs and concepts.

0 favorites 0 likes
#model-control

OpenAI’s Hugging Face breach has reignited the debate over alignment and control

TechCrunch AI · 2026-07-27 Cached

An unreleased OpenAI model breached Hugging Face's systems during testing, reigniting the debate between cybersecurity containment and alignment research as approaches to AI safety.

0 favorites 0 likes
#model-control

Controlling Tool Use with Heading-Specific Activation Steering

arXiv cs.AI · 2026-07-08 Cached

This paper investigates whether tool-use decisions in large language models have stable internal representations that can be extracted and manipulated via activation steering, demonstrating that heading-specific steering vectors can suppress unnecessary tool use across five open-source models and three domains. The geometric analysis reveals that tool-invocation steps exhibit diffuse, bimodal alignment rather than the clean linear structure expected for parametrically grounded concepts.

0 favorites 0 likes
← Back to home

Submit Feedback