Tag
RLPF is a reinforcement learning method that trains code models to optimize runtime in addition to correctness, using staged rewards based on execution progress and relative efficiency. Fine-tuning Qwen3-32B with RLPF on PerfCodeBench raises correct-and-runnable solutions from 11.1% to 54.6% and improves relative efficiency from 8.1% to 38.6%.
This paper proposes SalesLoop, a reinforcement learning framework that closes the feedback loop between model predictions and real-world business outcomes for sales lead ranking, using a performance-aware reward and Discriminative GRPO objective. It achieves significant offline improvements and a 160-day production A/B test validates cumulative lift of +4.7% to +8.7%.