标签
Introduces Critic-Free Pretraining (CFP), a method for offline-to-online RL that discards the offline-trained critic and uses a fresh critic with warm-up, matching or improving upon conventional O2O algorithms across tasks.
QPILOTS是一种方法,通过使用从噪声中间状态投影的评论家梯度,在推理时引导流策略,在离线到在线强化学习基准上实现了最先进的性能,并在不修改基础策略的情况下改进了预训练的VLA模型。