Tag
πR^2 introduces a reactive real-time flow policy for robot manipulation that splits conditioning into fast and slow channels and uses a latency-adaptive flow schedule, enabling closed-loop replanning ~4x faster than baseline policies and improving success rates by up to 30%.
Trust Region Q-Adjoint Matching (TRQAM) addresses instability in off-policy reinforcement learning by adaptively controlling path-space KL divergence through projected dual descent, enabling stable fine-tuning of pretrained flow policies. The method consistently outperforms prior arts on 50 OGBench tasks, achieving a 68% success rate in offline RL compared to the strongest baseline's 46%.