Tag
This paper proposes Threat-guided Policy-aware Scene Perturbation (TPSP), a method that augments online reinforcement learning for safe autonomous driving by perturbing scenes in a policy-aware, targeted manner to generate high-value safety-critical experiences. Experiments on NAVSIM v2 show improved safety learning efficiency with about 4 million kilometers of simulated driving data.