performance-feedback

Tag

Cards List
#performance-feedback

RLPF: Reinforcement Learning from Performance Feedback for Code Generation

arXiv cs.LG ↗ · 2026-07-31 Cached

RLPF is a reinforcement learning method that trains code models to optimize runtime in addition to correctness, using staged rewards based on execution progress and relative efficiency. Fine-tuning Qwen3-32B with RLPF on PerfCodeBench raises correct-and-runnable solutions from 11.1% to 54.6% and improves relative efficiency from 8.1% to 38.6%.

0 favorites 0 likes
#performance-feedback

SalesLoop: Reinforcement Learning from Performance Feedback for Sales Lead Ranking

arXiv cs.LG ↗ · 2026-07-24 Cached

This paper proposes SalesLoop, a reinforcement learning framework that closes the feedback loop between model predictions and real-world business outcomes for sales lead ranking, using a performance-aware reward and Discriminative GRPO objective. It achieves significant offline improvements and a 160-day production A/B test validates cumulative lift of +4.7% to +8.7%.

0 favorites 0 likes
← Back to home

Submit Feedback