representation-engineering

Tag

Cards List
#representation-engineering

Steering the Language Axis: From Linear Decodability to Causal Control

arXiv cs.CL · 2026-08-14 Cached

This paper investigates whether language identity in LLMs is linearly decodable and causally controllable via compact activation directions. Through steering and ablation experiments across multiple model families, the authors show that language selection is direction-dependent, layer-specific, and reverts to English when the language signal is ablated.

0 favorites 0 likes
#representation-engineering

RepBench: Compiling Benchmarks into Capability Representations for Large Language Models

arXiv cs.CL · 2026-07-31 Cached

RepBench is a benchmark-grounded data layer for representation probing in LLMs, built from 13,427 benchmark papers and 353 public datasets, yielding 46,149 probe texts across 94 capabilities. It evaluates capability vector transfer across models and readout methods.

0 favorites 0 likes
#representation-engineering

Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study

arXiv cs.CL · 2026-07-27 Cached

This paper analyzes how LLMs internally represent self-harm content, finding that self-harm information crystallizes in the final layers and that linear separability does not align with probe accuracy.

0 favorites 0 likes
#representation-engineering

Latent Fusion Jailbreak: Blending Harmful and Harmless Representations to Elicit Unsafe LLM Outputs

arXiv cs.CL · 2026-07-20 Cached

Introduces Latent Fusion Jailbreak (LFJ), a white-box attack that blends representations of harmful and harmless prompts in LLM hidden states, achieving 94.13% attack success rate. Also proposes a latent adversarial training defence that reduces ASR to 12.37%.

0 favorites 0 likes
#representation-engineering

Multi-Attribute Steering of Language Models via Targeted Intervention

arXiv cs.CL · 2026-07-13 Cached

MAT-Steer introduces a novel inference-time intervention framework for steering LLMs across multiple conflicting attributes by learning sparse, orthogonal steering vectors that selectively target tokens relevant to each attribute, achieving gains in QA tasks and generative tasks over prior methods.

0 favorites 0 likes
#representation-engineering

Controlling Tool Use with Heading-Specific Activation Steering

arXiv cs.AI · 2026-07-08 Cached

This paper investigates whether tool-use decisions in large language models have stable internal representations that can be extracted and manipulated via activation steering, demonstrating that heading-specific steering vectors can suppress unnecessary tool use across five open-source models and three domains. The geometric analysis reveals that tool-invocation steps exhibit diffuse, bimodal alignment rather than the clean linear structure expected for parametrically grounded concepts.

0 favorites 0 likes
#representation-engineering

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment

arXiv cs.AI · 2026-07-02 Cached

The paper analyzes how aligned LLMs encode harmfulness and refusal directions, revealing that jailbreaks suppress these directions. The authors propose HARC, a fine-tuning method that couples these directions across prompt and response positions, achieving robust safety alignment without degrading general capabilities.

0 favorites 0 likes
#representation-engineering

Want Better Synthetic Data? Steer It: Activation Steering for Low-Resource Language Generation

arXiv cs.CL · 2026-06-18 Cached

This paper investigates activation steering as an alternative to few-shot prompting for generating synthetic data in low-resource languages. The authors propose LanguageSteering and QualitySteering strategies, showing that steering on early layers improves diversity and downstream model performance.

0 favorites 0 likes
#representation-engineering

MechELK: A Mechanistic Interpretability Framework for Eliciting Latent Knowledge in Large Language Models

arXiv cs.CL · 2026-05-29 Cached

MechELK is a three-stage framework combining mechanistic interpretability tools (SAE, activation patching, causal probing) with representation engineering to elicit latent knowledge from LLMs, achieving 84.7% accuracy and outperforming existing methods like CCS and linear probing.

0 favorites 0 likes
#representation-engineering

Decomposing and Steering Functional Metacognition in Large Language Models

arXiv cs.CL · 2026-05-12 Cached

This research paper investigates functional metacognition in Large Language Models, demonstrating that internal states like evaluation awareness and self-assessed capability are linearly decodable from residual stream activations. The authors propose a mechanistic framework to steer these states, showing causal control over reasoning behaviors, verbosity, and safety responses.

0 favorites 0 likes
← Back to home

Submit Feedback