interventions

Tag

Cards List
#interventions

Graph Surgery and the Do-Operator: A Precise Correspondence for Acyclic Structural Causal Models

arXiv cs.AI · 2026-08-19 Cached

This paper establishes a precise mathematical correspondence between graph surgery and the do-operator in acyclic structural causal models, proving their equivalence in terms of dependency graphs.

0 favorites 0 likes
#interventions

DoTime: A Synthetic Benchmark Generator for Interventional and Counterfactual Time Series

arXiv cs.LG · 2026-07-31 Cached

DoTime is a synthetic benchmark generator for interventional and counterfactual time series, providing scalable TSCM-based data generation with exact ground truth, released as a PyPI package with evaluation suites. It enables training and benchmarking causal foundation models on non-observational time series data.

0 favorites 0 likes
#interventions

When Are Sparse Feature Interventions Actually Localized? Matched Evaluation for SAE-Based Safety Control

arXiv cs.AI · 2026-07-14 Cached

This paper introduces a matched evaluation protocol for sparse feature interventions in language models, showing that the claimed efficiency advantage of SAE-based safety control disappears or reverses when properly comparing against fair dense baselines.

0 favorites 0 likes
#interventions

@swyx: imo this is the most impt part of anthropic's J-space paper today. it's a two-parter: 1) ant proved that they can do "b…

X AI KOLs Timeline · 2026-07-07 Cached

Anthropic's J-space paper demonstrates that they can perform 'brain surgery' interventions into reasoning to change topics midstream, and the model is able to detect what intervention was done, indicating a form of eval awareness.

0 favorites 0 likes
← Back to home

Submit Feedback