Tag
This arXiv paper proposes a protocol to measure the predictive credit of scientific explanations for experimental forecasts, finding that matched explanations did not significantly improve prediction accuracy across Tox21 and OpenML benchmarks, though some gains appeared under certain model replays.