grpo

Tag

Cards List
#grpo

SVR-R1: Bootstrapping Multi-modal Reasoning with Self-verification in Reinforcement Learning

Hugging Face Daily Papers ↗ · 2026-07-13 Cached

Introduces SVR-R1, a multi-turn reinforcement learning framework that uses the model's own verification as a learning signal for multi-modal reasoning, achieving significant accuracy improvements over standard GRPO baselines on vision-language reasoning benchmarks.

0 favorites 0 likes
#grpo

@akshay_pachaar: Andrej Karpathy summarized the entire history of LLM training in three nouns: - text - conversations - and environments…

X AI KOLs Timeline ↗ · 2026-07-12 Cached

Andrej Karpathy frames LLM training as text, conversations, and environments; Prime Intellect's Verifiers is an open-source framework for building and sharing RL environments for LLMs, released under MIT license, with a hub of 2500+ environments.

0 favorites 0 likes
#grpo

@akshay_pachaar: LLM fine-tuning techniques I'd learn if I were to customize them: Bookmark this. 1. LoRA 2. QLoRA 3. Prefix Tuning 4. A…

X AI KOLs Following ↗ · 2026-07-10 Cached

The tweet lists 15 LLM fine-tuning techniques and introduces ART (Agent Reinforcement Trainer), an open-source framework from OpenPipe for training multi-step agents using GRPO, with serverless RL support via W&B Training.

0 favorites 0 likes
#grpo

Reinforcing the Generation Order of Multimodal Masked Diffusion Models

arXiv cs.LG ↗ · 2026-07-10 Cached

This paper introduces a learnable control module trained via Group Relative Policy Optimization (GRPO) to optimize the generation order in multimodal masked diffusion models, achieving improvements in text-to-image alignment and multimodal understanding.

0 favorites 0 likes
#grpo

When Synthetic Speech Is All You Have: Better Call GRPO

arXiv cs.CL ↗ · 2026-07-10 Cached

This paper proposes using Group Relative Policy Optimization (GRPO) for adapting LLM-based ASR models to regulated domains using only synthetic speech, achieving 40-45% relative WER reduction over supervised fine-tuning.

0 favorites 0 likes
#grpo

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning

arXiv cs.CL ↗ · 2026-07-10 Cached

Proposes Tail-Aware Credit Calibration (TACO) to address positive-credit contamination in LLM reinforcement learning by calibrating uniform credit assignment to suppress updates on implausible tail tokens, improving training stability and performance.

0 favorites 0 likes
#grpo

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems

arXiv cs.CL ↗ · 2026-07-09 Cached

This paper introduces AdaPrefix-GRPO, a method that adaptively controls the length of correct solution prefixes provided to a model during GRPO training, maintaining a 50% success rate to maximize gradient signal. It significantly improves accuracy on hard math reasoning problems while reducing computational cost.

0 favorites 0 likes
#grpo

DrugGen 2: A disease-aware language model for enhancing drug discovery

Hugging Face Daily Papers ↗ · 2026-07-09 Cached

DrugGen-2 fine-tunes GPT-2 using supervised learning and reinforcement learning (GRPO) to generate small molecules conditioned on both disease ontology and target protein sequences, achieving superior diversity and binding affinity for drug discovery.

0 favorites 0 likes
#grpo

Self-Review Reinforcement Learning (SRRL) with Cross-Episode Memory and Policy Distillation

arXiv cs.LG ↗ · 2026-07-08 Cached

Introduces Self-Review Reinforcement Learning (SRRL), a training framework that embeds a self-review step into RL episodes to transform sparse environmental feedback into behavioral improvements, using cross-episode memory and selective policy distillation. Evaluated on GSM8K, SRRL outperforms standard RLVR with GRPO across Qwen 3-4B and OLMo-3-7B models.

0 favorites 0 likes
#grpo

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning

Hugging Face Daily Papers ↗ · 2026-07-08 Cached

This paper presents Single-rollout Asynchronous Optimization (SAO) to address stability and off-policy challenges in asynchronous RL for agentic tasks, outperforming GRPO and its variants on coding and reasoning benchmarks. SAO is deployed in the GLM-5.2 model's agentic RL pipeline.

0 favorites 0 likes
#grpo

@cwolferesearch: Agentic RL requires new algorithm modifications. In GRPO, the “group” used starts to change when training agents… In va…

X AI KOLs Timeline ↗ · 2026-07-07 Cached

This thread discusses modifications to GRPO for agentic RL, focusing on different levels of advantage normalization (prompt-level, task-level, environment-level) to handle higher reward variance in multi-task, multi-turn environments.

0 favorites 0 likes
#grpo

NormWorlds-CF: Solver-Verified Counterfactual Normative Reasoning with Metamorphic-Relation GRPO

arXiv cs.CL ↗ · 2026-07-07 Cached

NormWorlds-CF is a solver-verified benchmark for counterfactual normative reasoning. The paper proposes MR-GRPO, a reward mechanism that improves structured reasoning beyond final answers, showing that answer-only accuracy can be misleading in normative tasks.

0 favorites 0 likes
#grpo

@akshay_pachaar: https://x.com/akshay_pachaar/status/2074200571834515574

X AI KOLs Following ↗ · 2026-07-06 Cached

A technical tutorial on building a reinforcement learning environment for LLMs using the open-source Verifiers library, with Othello as a working example.

0 favorites 0 likes
#grpo

TREK: Distill to Explore, Reinforce to Refine

Hugging Face Daily Papers ↗ · 2026-07-06 Cached

TREK is a staged procedure that uses distillation to expand exploration support for policy optimization, improving performance on mathematical reasoning and agentic tasks beyond standard GRPO.

0 favorites 0 likes
#grpo

Mastermind: Strategy-grounded Learning for Repository-Scale Vulnerability Reproduction

arXiv cs.AI ↗ · 2026-07-03 Cached

This paper introduces Mastermind, a dual-loop framework that learns reusable vulnerability-reproduction strategies for repository-scale tasks, achieving an 84.5% pass rate with a frozen executor by separating strategy learning from execution.

0 favorites 0 likes
#grpo

Beyond Next-Token Prediction: An RLVR Proof of Concept for Tool-Use Agents on Atlassian Workflows

arXiv cs.AI ↗ · 2026-07-03 Cached

This paper presents a proof-of-concept using Reinforcement Learning with Verifiable Rewards (RLVR) to train small language models for tool-use in enterprise SaaS workflows like Jira and Confluence. The approach uses synthetic environments and GRPO training to improve tool-call accuracy, achieving significant reward gains over baselines.

0 favorites 0 likes
#grpo

GRPO, Dr. GRPO, and DAPO Are Three Operations on One Number: The Group-Standard-Deviation Identity

arXiv cs.LG ↗ · 2026-07-02 Cached

This paper proves that GRPO, Dr. GRPO, and DAPO are three formulations of the same underlying mechanism: adjusting the standard deviation of rewards within a group of sampled answers. The group-standard-deviation identity shows that unanimous groups teach nothing while split groups drive learning, revealing a unified dial for training language models to reason.

0 favorites 0 likes
#grpo

@neural_avb: https://x.com/neural_avb/status/2072294078805684613

X AI KOLs Timeline ↗ · 2026-07-01 Cached

This paper introduces Autodata, a method that uses an agentic 'data scientist' AI to automate the creation of high-quality synthetic datasets through iterative generation, verification, and refinement, specifically optimized for reinforcement learning (GRPO) to improve reasoning in language models.

0 favorites 0 likes
#grpo

Predictable GRPO: A Closed-Form Model of Training Dynamics

arXiv cs.LG ↗ · 2026-07-01 Cached

Presents a closed-form reduced-order model of GRPO training dynamics, reducing it to a damped oscillator and deriving predictions for stability, group-size invariance, and loss curvature. Validated across multiple models and benchmarks.

0 favorites 0 likes
#grpo

Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners

arXiv cs.AI ↗ · 2026-06-30 Cached

PASS is a middleware that fixes three pathologies in process-supervised RL for LLM reasoners, improving GRPO by independently standardizing streams, chunking by value, and using average value density. It shows consistent gains in math reasoning and multi-hop QA.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback