policy-refinement

Tag

Cards List
#policy-refinement

TEMPO: Scaling Test-time Training for Large Reasoning Models

Hugging Face Daily Papers · 2026-04-21 Cached

TEMPO introduces a test-time training framework that alternates policy refinement with critic recalibration to prevent diversity collapse and sustain performance gains in large reasoning models, boosting AIME 2024 scores for Qwen3-14B from 42.3% to 65.8%.

0 favorites 0 likes
← Back to home

Submit Feedback