offline-rl

Tag

Cards List
#offline-rl

Scaling World-Model Reinforcement Learning Through Diffusion Policy Optimization

arXiv cs.LG · 2026-05-27 Cached

Proposes Model-Based Diffusion Policy Optimization (MBDPO), a framework that unifies search and policy optimization in world models using diffusion policy representations, achieving consistent scaling behavior and superior performance across offline and online reinforcement learning tasks.

0 favorites 0 likes
#offline-rl

Exploiting Local Dynamics Regularity for Reusable Skills in Offline Hierarchical RL

arXiv cs.AI · 2026-05-27 Cached

This paper introduces CARL, a method for offline hierarchical reinforcement learning that exploits local dynamics regularity to learn reusable skills. The approach clusters state-goal pairs requiring similar action sequences, enabling more effective skill reuse and improved performance on complex humanoid tasks.

0 favorites 0 likes
#offline-rl

Generative OOD-regularized Model-based Policy Optimization

arXiv cs.LG · 2026-05-26 Cached

Introduces GORMPO, a density-regularized offline RL algorithm that uses generative density modeling to restrict policy updates to high-density areas, achieving 17% improvement on a real-world medical dataset and outperforming state-of-the-art baselines.

0 favorites 0 likes
#offline-rl

ACSAC: Adaptive Chunk Size Actor-Critic with Causal Transformer Q-Network

arXiv cs.LG · 2026-05-13 Cached

This paper introduces ACSAC, a reinforcement learning method that uses an adaptive chunk size actor-critic algorithm with a causal Transformer Q-network to handle long-horizon, sparse-reward tasks. It demonstrates state-of-the-art performance on manipulation tasks by dynamically adjusting action chunk sizes based on state-dependent needs.

0 favorites 0 likes
#offline-rl

Path-Coupled Bellman Flows for Distributional Reinforcement Learning

arXiv cs.LG · 2026-05-12 Cached

This paper introduces Path-Coupled Bellman Flows (PCBF), a continuous-time distributional reinforcement learning method that uses flow matching to model return distributions without heuristic projections. It addresses boundary mismatch and high-variance issues in previous flow-based approaches by coupling current and successor return flows through shared base noise.

0 favorites 0 likes
#offline-rl

Adaptive Q-Chunking for Offline-to-Online Reinforcement Learning

arXiv cs.LG · 2026-05-08 Cached

This paper introduces Adaptive Q-Chunking (AQC), a reinforcement learning method that dynamically selects action chunk sizes to balance reactive control and long-horizon planning. It achieves state-of-the-art results on OGBench and Robomimic, enhancing the performance of large-scale VLA models in robotics tasks.

0 favorites 0 likes
#offline-rl

Reinforcement Learning via Value Gradient Flow

Hugging Face Daily Papers · 2026-04-15 Cached

Value Gradient Flow (VGF) presents a scalable approach to behavior-regularized reinforcement learning by formulating it as an optimal transport problem solved through discrete gradient flow, achieving state-of-the-art results on offline RL and LLM RL benchmarks. The method eliminates explicit policy parameterization while enabling adaptive test-time scaling by controlling transport budget.

0 favorites 0 likes
#offline-rl

On the Generalization of SFT: A Reinforcement Learning Perspective with Reward Rectification

Papers with Code Trending · 2025-08-07 Cached

This paper analyzes limitations in standard supervised fine-tuning (SFT) from a reinforcement learning perspective and proposes Dynamic Fine-Tuning (DFT), a simple gradient-rescaling method that improves LLM generalization and matches offline RL performance.

0 favorites 0 likes
← Previous
← Back to home

Submit Feedback