标签
DoTime is a synthetic benchmark generator for interventional and counterfactual time series, providing scalable TSCM-based data generation with exact ground truth, released as a PyPI package with evaluation suites. It enables training and benchmarking causal foundation models on non-observational time series data.
本文介绍了一种针对语言模型中稀疏特征干预的匹配评估协议,表明在与公平密集基线进行恰当对比时,基于SAE的安全控制所声称的效率优势会消失或逆转。
Anthropic 的 J-space 论文表明,他们可以对推理进行“脑外科”干预以中途改变话题,并且模型能够检测到所做的干预,这表明了一种评估意识。