offline-to-online

Tag

Cards List
#offline-to-online

Critic-Free Pretraining for Efficient Online Reinforcement Learning Fine-Tuning

arXiv cs.LG ↗ · 2026-08-12 Cached

Introduces Critic-Free Pretraining (CFP), a method for offline-to-online RL that discards the offline-trained critic and uses a fresh critic with warm-up, matching or improving upon conventional O2O algorithms across tasks.

0 favorites 0 likes
#offline-to-online

QPILOTS: Efficient Test-Time Q-Steering for Flow Policies

arXiv cs.LG ↗ · 2026-06-16 Cached

QPILOTS is a method that steers flow policies at inference time by using critic gradients projected from noisy intermediate states, achieving state-of-the-art performance on offline-to-online RL benchmarks and improving pretrained VLA models without modifying the base policy.

0 favorites 0 likes
← Back to home

Submit Feedback