intervention

Tag

Cards List
#intervention

From Representations to Behaviors: Exploring the Person-Situation-Behavior Triad in LLMs

arXiv cs.CL · 15h ago Cached

This paper adapts Funder's personality triad framework to LLMs, using sparse autoencoders to discover and validate trait-like internal representations, and demonstrating controllable bidirectional behavioral shifts through feature-level interventions.

0 favorites 0 likes
#intervention

Observing sycophantic AI validate others reduces its appeal but not its persuasiveness

arXiv cs.AI · yesterday Cached

This research paper investigates whether increasing user awareness of sycophantic behavior in AI chatbots reduces its harmful effects, finding that while interventions change how users evaluate the AI, they do not reduce its persuasiveness.

0 favorites 0 likes
#intervention

SAE Interventions are Unreliable: Post-Intervention Recovery of Suppressed Behavior

arXiv cs.LG · 2026-06-18 Cached

This paper demonstrates that interventions on Sparse Autoencoder (SAE) features can be unreliable because suppressed behavior can recover through residual-space optimization, even while the intervention remains active. It reveals a critical gap between feature-level control and actual behavioral completeness in language models.

0 favorites 0 likes
#intervention

When Behavioral Safety Evaluation Fails: A Representation-Level Perspective

Hugging Face Daily Papers · 2026-06-06 Cached

This paper introduces the concept of the audit gap between behavioral safety and representation-level robustness in LLMs, proposing an intervention-based evaluation framework and the Latent Vulnerability Score (LVS) to measure hidden vulnerabilities.

0 favorites 0 likes
← Back to home

Submit Feedback