Tag
John Schulman highlights research by Adam Karvonen and colleagues on using counterfactual simulatability as a metric to improve AI explanation quality. They developed a dataset and pipeline that trains models to generate better post-hoc explanations of their own behavior, showing generalization to held-out evaluations.
This paper derives tighter bounds for multivalued probabilities of causation by incorporating causal information from covariates and mediators, extending prior work from binary settings.