Parameter Exploration for RLVR via Variational Learning

Hugging Face Daily Papers Papers

Summary

This paper introduces Perturbed Parameter Policy Optimization (3PO), a family of parameter-space exploration methods for LLM reinforcement learning, showing consistent improvements over GRPO on math and code tasks.

Exploration has been a focus of reinforcement learning research for a long time. Recently, there has been growing evidence that it is also an important ingredient in LLM reinforcement learning recipes that can significantly impact downstream performance. Many existing methods control exploration in the action-space, for example, using temperature scaling. However, these methods cannot reorder tokens but only influence the variance in the output distribution. This limits exploration and can lead to divergence or stalled training. Here, we investigate parameter-space exploration, where rollouts are generated by sampling different policies from a posterior that may each explore different rollouts. Sampling less or more diverse policies is then a complementary control lever over exploration. We introduce a family of methods called Perturbed Parameter Policy Optimization (3PO) which use different sampling strategies and different rollout grouping for reward estimation. Experiments on OLMo-3-1025-7B and Qwen2.5-Math-7B across mathematical reasoning and code generation tasks show that these approaches consistently improve average downstream performance over standard GRPO at a near-identical FLOPs cost. Moreover, using multiple parameter samples consistently produces fewer zero-advantage groups and malformed or incorrect rollouts during training than GRPO and action-space baselines. Overall, our work presents evidence that parameter-space exploration can improve reinforcement learning for LLMs.
Original Article
View Cached Full Text

Cached at: 08/13/26, 03:32 PM

Paper page - Parameter Exploration for RLVR via Variational Learning

Source: https://huggingface.co/papers/2608.09805 https://huggingface.co/login?next=%2Fpapers%2F2608.09805-

Abstract

Parameter-space exploration via perturbed policy sampling improves LLM reinforcement learning by diversifying rollouts and reducing training failures compared to action-space methods.

Exploration has been a focus of reinforcement learning research for a long time. Recently, there has been growing evidence that it is also an important ingredient inLLM reinforcement learningrecipes that can significantly impact downstream performance. Many existing methods control exploration in the action-space, for example, usingtemperature scaling. However, these methods cannot reorder tokens but only influence the variance in the output distribution. This limits exploration and can lead to divergence or stalled training. Here, we investigateparameter-space exploration, where rollouts are generated by sampling different policies from a posterior that may each explore different rollouts. Sampling less or more diverse policies is then a complementary control lever over exploration. We introduce a family of methods calledPerturbed Parameter Policy Optimization (3PO)which use different sampling strategies and differentrollout groupingforreward estimation. Experiments on OLMo-3-1025-7B and Qwen2.5-Math-7B across mathematical reasoning and code generation tasks show that these approaches consistently improve average downstream performance over standardGRPOat a near-identical FLOPs cost. Moreover, using multiple parameter samples consistently produces fewerzero-advantage groupsand malformed or incorrect rollouts during training thanGRPOand action-space baselines. Overall, our work presents evidence thatparameter-space explorationcan improve reinforcement learning for LLMs.

View arXiv pageView PDFGitHub0Add to collection

Community

Paper submitter

about 5 hours ago

Parameter Exploration for RLVR via Variational Learning

Upload images, audio, and videos by dragging in the text input, pasting, orclicking here.

Tap or paste here to upload images

https://huggingface.co/login?next=%2Fpapers%2F2608.09805-

Get this paper in your agent:

hf papers read 2608\.09805

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper14

#### BayesRL/Qwen2.5Math-IVON-SFT-7B Text Generation• 8B• Updatedabout 5 hours ago • 519 #### BayesRL/Olmo3-IVON-SFT-7B Text Generation• 7B• Updatedabout 5 hours ago • 4.4k #### BayesRL/Olmo3-B3PO-7B Text Generation• 7B• Updatedabout 5 hours ago • 9 #### BayesRL/Olmo3-M3PO-7B Text Generation• 7B• Updatedabout 5 hours ago • 6 Browse 14 models citing this paper## Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2608.09805 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2608.09805 in a Space README.md to link it from this page.

Collections including this paper1

Similar Articles

Gradient Extrapolation-Based Policy Optimization

arXiv cs.LG

The article introduces Gradient Extrapolation-Based Policy Optimization (GXPO), a method that approximates multi-step lookahead in RL training for LLMs using only three backward passes. It demonstrates improved reasoning performance on math benchmarks over standard GRPO while maintaining fixed active-phase costs.