parameter-space-exploration

Tag

Cards List
#parameter-space-exploration

Parameter Exploration for RLVR via Variational Learning

Hugging Face Daily Papers · 2026-08-10 Cached

This paper introduces Perturbed Parameter Policy Optimization (3PO), a family of parameter-space exploration methods for LLM reinforcement learning, showing consistent improvements over GRPO on math and code tasks.

0 favorites 0 likes
← Back to home

Submit Feedback