Tag
Introduces Critic-Free Pretraining (CFP), a method for offline-to-online RL that discards the offline-trained critic and uses a fresh critic with warm-up, matching or improving upon conventional O2O algorithms across tasks.
QPILOTS is a method that steers flow policies at inference time by using critic gradients projected from noisy intermediate states, achieving state-of-the-art performance on offline-to-online RL benchmarks and improving pretrained VLA models without modifying the base policy.