Tag
This paper proposes a causal model for understanding and counteracting sandbagging in AI models, using interventions like reference grafting and context grafting to restore capabilities in open-weight models.
The paper proposes bipartite graphical causal models (BGCMs) to resolve ambiguities in causal interventions for systems at equilibrium with cyclic dependencies, generalizing existing frameworks like causal Bayesian networks and structural causal models.
This paper introduces Predictive Memory Localization (PML), a method that forecasts selective intervention outcomes from internal model signals, distinguishing calibrated target movement from semantic-neighbor and capability damage. Experiments across 3,000 records and nine datasets show improved selective-path prediction and risk-aware intervention decisions.
This paper adapts Funder's personality triad framework to LLMs, using sparse autoencoders to discover and validate trait-like internal representations, and demonstrating controllable bidirectional behavioral shifts through feature-level interventions.
This research paper investigates whether increasing user awareness of sycophantic behavior in AI chatbots reduces its harmful effects, finding that while interventions change how users evaluate the AI, they do not reduce its persuasiveness.
This paper demonstrates that interventions on Sparse Autoencoder (SAE) features can be unreliable because suppressed behavior can recover through residual-space optimization, even while the intervention remains active. It reveals a critical gap between feature-level control and actual behavioral completeness in language models.
This paper introduces the concept of the audit gap between behavioral safety and representation-level robustness in LLMs, proposing an intervention-based evaluation framework and the Latent Vulnerability Score (LVS) to measure hidden vulnerabilities.