rl-post-training

Tag

Cards List
#rl-post-training

Beating GPT-5.6 Sol on retrieval with 100x cheaper open models

Hacker News Top · 2026-08-05 Cached

Neon and Castform team up to show how RL post-trained open-weights models combined with Lakebase Search can beat frontier models like GPT-5.6 Sol on retrieval tasks at 100x lower cost and latency.

0 favorites 0 likes
#rl-post-training

PIRL: From Open-Loop Exploration to Closed-Loop Reinforcement Learning [R]

Reddit r/MachineLearning · 2026-07-28

Introduces PIRL (Policy Improvement Reinforcement Learning) and its practical implementation PIPO, a closed-loop framework that verifies policy updates by comparing performance with a historical anchor, enabling correction or reinforcement of previous updates. Experiments show consistent gains in mathematical reasoning, code generation, tool use, and self-distillation when applied on top of existing RL algorithms like PPO and GRPO.

0 favorites 0 likes
#rl-post-training

Flow-Map GRPO: Reinforcement Learning for Few-Step Flow-Map Generators via Anchored Stochastic Composition

arXiv cs.LG · 2026-07-02 Cached

Proposes Flow-Map GRPO, an online RL post-training framework for deterministic few-step flow-map generators, introducing Anchored Stochastic Flow Map Composition (ASFMC) to enable stochastic optimization without altering original model parameterization. Experiments on FLUX-based MeanFlow and sCM show improvement across reward-based, perceptual, and task-level metrics.

0 favorites 0 likes
#rl-post-training

@no_stp_on_snek: verdict up front: it's a "pass" in my book in certain categories, just a narrower one than the 35B. you're buying real …

X AI KOLs Following · 2026-06-27 Cached

The author evaluates Ornith-9B against its base Qwen3.5-9B, finding that RL post-training improves token efficiency and sustained coding coherence but sacrifices single-turn judgment and robustness to misleading inputs, making it a narrower upgrade at 9B compared to the 35B version.

0 favorites 0 likes
#rl-post-training

STAR: SpatioTemporal Adaptive Reward Allocation for Text-to-Image RL Post-Training

arXiv cs.AI · 2026-06-17 Cached

This paper introduces STAR, a method for spatiotemporally adaptive reward allocation in RL post-training for text-to-image diffusion models, improving compositional alignment and text rendering by focusing policy updates on relevant latent regions.

0 favorites 0 likes
← Back to home

Submit Feedback