intervention

Tag

Cards List
#intervention

A Causal Model for Locating and Unlocking Sandbagging in Model Organisms

arXiv cs.LG · 2026-09-01 Cached

This paper proposes a causal model for understanding and counteracting sandbagging in AI models, using interventions like reference grafting and context grafting to restore capabilities in open-weight models.

0 favorites 0 likes
#intervention

Causal Reasoning with Bipartite Graphical Causal Models

arXiv cs.AI · 2026-08-21 Cached

The paper proposes bipartite graphical causal models (BGCMs) to resolve ambiguities in causal interventions for systems at equilibrium with cyclic dependencies, generalizing existing frameworks like causal Bayesian networks and structural causal models.

0 favorites 0 likes
#intervention

Predictive Memory Localization: Forecasting Selective Intervention Paths from Internal Signals

arXiv cs.AI · 2026-08-14 Cached

This paper introduces Predictive Memory Localization (PML), a method that forecasts selective intervention outcomes from internal model signals, distinguishing calibrated target movement from semantic-neighbor and capability damage. Experiments across 3,000 records and nine datasets show improved selective-path prediction and risk-aware intervention decisions.

0 favorites 0 likes
#intervention

From Representations to Behaviors: Exploring the Person-Situation-Behavior Triad in LLMs

arXiv cs.CL · 2026-07-30 Cached

This paper adapts Funder's personality triad framework to LLMs, using sparse autoencoders to discover and validate trait-like internal representations, and demonstrating controllable bidirectional behavioral shifts through feature-level interventions.

0 favorites 0 likes
#intervention

Observing sycophantic AI validate others reduces its appeal but not its persuasiveness

arXiv cs.AI · 2026-07-29 Cached

This research paper investigates whether increasing user awareness of sycophantic behavior in AI chatbots reduces its harmful effects, finding that while interventions change how users evaluate the AI, they do not reduce its persuasiveness.

0 favorites 0 likes
#intervention

SAE Interventions are Unreliable: Post-Intervention Recovery of Suppressed Behavior

arXiv cs.LG · 2026-06-18 Cached

This paper demonstrates that interventions on Sparse Autoencoder (SAE) features can be unreliable because suppressed behavior can recover through residual-space optimization, even while the intervention remains active. It reveals a critical gap between feature-level control and actual behavioral completeness in language models.

0 favorites 0 likes
#intervention

When Behavioral Safety Evaluation Fails: A Representation-Level Perspective

Hugging Face Daily Papers · 2026-06-06 Cached

This paper introduces the concept of the audit gap between behavioral safety and representation-level robustness in LLMs, proposing an intervention-based evaluation framework and the Latent Vulnerability Score (LVS) to measure hidden vulnerabilities.

0 favorites 0 likes
← Back to home

Submit Feedback