Kalman Meets Curriculum: Efficient Dynamic Prompt Selection for Adaptive RL Finetuning

arXiv cs.LG Papers

Summary

This paper introduces KGPS, a Kalman-guided prompt selection method for adaptive RL finetuning of LLMs, which models prompt difficulty as a dynamic state to improve accuracy and rollout efficiency.

arXiv:2607.27610v1 Announce Type: new Abstract: Reinforcement learning (RL) finetuning significantly enhances the reasoning capabilities of large language models (LLMs), yet its effectiveness critically depends on selecting prompts of appropriate difficulty for the current policy. This is challenging because prompt difficulty evolves throughout training. Existing online methods therefore face a trade-off: evaluation-based approaches are accurate but expensive, while prediction-based approaches are efficient but typically assume stationary difficulty, making them ill-suited to RL's non-stationary training dynamics. To address these issues, we propose a Kalman-Guided Prompt Selection method (KGPS), which reformulates prompt selection as a dynamic state estimation problem rather than static difficulty prediction. KGPS models each prompt's latent success rate in logit space using a linear-Gaussian state-space model, with process noise coupled to the magnitude of policy updates so that uncertainty increases when the policy changes more substantially. A Kalman filter then maintains a calibrated Gaussian posterior over prompt difficulty, and prompts are selected by maximizing a posterior-expected training utility that favors intermediate-difficulty prompts while naturally revisiting uncertain ones. The resulting procedure is adaptive to policy drift and requires no additional rollouts beyond standard policy training. Extensive experiments across mathematics, planning, and geometry reasoning benchmarks, as well as multiple RL algorithms, show that KGPS consistently improves both final accuracy and rollout efficiency over strong baselines, establishing state-of-the-art performance among online prompt selection methods. For example, on DeepSeek-R1-Distill-7B, KGPS uses 83% fewer rollouts than DS while even improving the average performance by 0.12 point across six math reasoning benchmarks.
Original Article
View Cached Full Text

Cached at: 07/31/26, 10:04 AM

# Kalman Meets Curriculum: Efficient Dynamic Prompt Selection for Adaptive RL Finetuning
Source: [https://arxiv.org/html/2607.27610](https://arxiv.org/html/2607.27610)
Haodong Zhu1,2Yangyang Ren1,211footnotemark:1Yanjing Li3Sheng Xu422footnotemark:2 Haiguang Liu2Linlin Yang3Baochang Zhang1 1Beihang University2Zhongguancun Academy3Nanyang Technological university 4Independent Researcher5Communication University of China HaodongZhu@buaa\.edu\.cnyyren@buaa\.edu\.cn

###### Abstract

Reinforcement learning \(RL\) finetuning significantly enhances the reasoning capabilities of large language models \(LLMs\), yetits effectiveness critically depends on selecting prompts of appropriate difficulty for the current policy\.This is challenging because prompt difficulty evolves throughout training\.Existing online methods therefore face a trade\-off: evaluation\-based approaches are accurate but expensive, while prediction\-based approaches are efficient but typically assume stationary difficulty, making them ill\-suited to RL’s non\-stationary training dynamics\.To address these issues, we propose a*Kalman\-GuidedPromptSelection*method \(*KGPS*\), which reformulates prompt selection as a dynamic state estimation problemrather than static difficulty prediction\.KGPS models each prompt’s latent success rate in logit space using a linear\-Gaussian state\-space model, with process noise coupled to the magnitude of policy updates so that uncertainty increases when the policy changes more substantially\.A Kalman filter then maintains a calibrated Gaussian posterior over prompt difficulty, and prompts are selected by maximizing a posterior\-expected training utility that favors intermediate\-difficulty prompts while naturally revisiting uncertain ones\.The resulting procedure is adaptive to policy drift and requires no additional rollouts beyond standard policy training\.Extensive experiments across mathematics, planning, and geometry reasoning benchmarks, as well as multiple RL algorithms, show that KGPS consistently improves both final accuracy and rollout efficiency over strong baselines, establishing state\-of\-the\-art performance among online prompt selection methods\. For example, on DeepSeek\-R1\-Distill\-7B, KGPS uses 83% fewer rollouts than DS while even improving the average performance by 0\.12 point across six math reasoning benchmarks\.

![Refer to caption](https://arxiv.org/html/2607.27610v1/x1.png)

\(a\) Test Accuracy vs\. Rollout Cost

![Refer to caption](https://arxiv.org/html/2607.27610v1/x2.png)

\(b\)Evolution of MAE between empirical success\-rate and predicted success\-rates over training

Figure 1:On the Math dataset using the Qwen3\-0\.6B model, \(a\)KGPS obtains comparable performance to evaluation\-based DS \(Oracle\) with 71\.00% fewer rollouts, also surpasses vanilla uniform method by up to \+2\.94% higher test accuracy with the same rollout cost\. Meanwhile, compared to prediction\-based prompt selection methods GRESO and MoPPS, the proposed KGPS achieves up to \+2\.33% higher test accuracy\.\(b\)We calculate the MAE between empirical success\-rate and predicted success\-rate, whereKGPS maintains substantially lowerMAEthan MoPPS throughout training, demonstrating superior prompt difficulty prediction without additional LLM inference\.## 1Introduction

Reinforcement learning \(RL\) has become a cornerstone technique for post\-training Large Language Models \(LLMs\), driving remarkable advances in both instruction following and complex reasoning\(Shao et al\.,[2024](https://arxiv.org/html/2607.27610#bib.bib30);DeepSeek\-AI et al\.,[2025](https://arxiv.org/html/2607.27610#bib.bib5);Zeng et al\.,[2025](https://arxiv.org/html/2607.27610#bib.bib40);Luo et al\.,[2025](https://arxiv.org/html/2607.27610#bib.bib21)\)\. However, realizing its full potential remains non\-trivial\. RL optimization is known to be highly sensitive to training sample selection\(Parashar et al\.,[2025](https://arxiv.org/html/2607.27610#bib.bib27);Qu et al\.,[2025](https://arxiv.org/html/2607.27610#bib.bib28);Shen et al\.,[2025a](https://arxiv.org/html/2607.27610#bib.bib31),[b](https://arxiv.org/html/2607.27610#bib.bib32)\),because prompts that are already mastered by the current policy contribute little learning signal, while prompts that remain far beyond the policy’s capability often produce uninformative gradients\.In practice, the most useful prompts tend to be those at an appropriate difficulty for the current policy\. Crucially, prompt difficulty is not fixed: as the policy improves, previously informative prompts may become trivial, and the set of useful training samples correspondingly shifts over time\. This dynamic undermines static selection strategies—including uniform sampling and fixed curricula—which can waste rollout budget on uninformative prompts and become increasingly misaligned with the model’s evolving competence\(Yu et al\.,[2025](https://arxiv.org/html/2607.27610#bib.bib38);Bae et al\.,[2025](https://arxiv.org/html/2607.27610#bib.bib1)\)\.This underscores the need for adaptive task selection that remains aligned with the current policy throughout training\.

Prior work on data selection for RL training splits into offline and online approaches\. Offline methods precompute static difficulty estimates and use fixed curricula\(Parashar et al\.,[2025](https://arxiv.org/html/2607.27610#bib.bib27);Wen et al\.,[2025](https://arxiv.org/html/2607.27610#bib.bib35);Shen et al\.,[2025b](https://arxiv.org/html/2607.27610#bib.bib32)\), but these misalign as the policy improves\. Online methods dynamically recompute sample utility using real\-time feedback\(Qu et al\.,[2025](https://arxiv.org/html/2607.27610#bib.bib28);Shen et al\.,[2025a](https://arxiv.org/html/2607.27610#bib.bib31);Yu et al\.,[2025](https://arxiv.org/html/2607.27610#bib.bib38)\), maintaining alignment with evolving competence\.In RL finetuning, this online setting is especially important because prompt success rate under the current policy is itself non\-stationary\.

Online selection methods broadly fall into two categories depending on how sample utility is estimated\. Evaluation\-based approaches, such as Dynamic Sampling \(DS\)\(Yu et al\.,[2025](https://arxiv.org/html/2607.27610#bib.bib38)\), operateby oversampling a large candidate set, executing rollouts across all candidates, and filtering prompts based on the observed reward outcomes\.While yielding reliable utility estimates,these candidate evaluationsdemandsubstantial additionalrolloutsbeyond those already needed for policy training, resulting in significant inference overhead that compounds at every update step\. As illustrated in Fig\.[1](https://arxiv.org/html/2607.27610#S0.F1)\(a\),in DS training, this overhead\(blue curve\)can reach up to7×7\\timesmore rollouts than standard GRPO training\(yellow curve\), severely limiting practical scalability\. To address this, prediction\-based approaches, notably MoPPS\(Qu et al\.,[2025](https://arxiv.org/html/2607.27610#bib.bib28)\)\(purple curve\)and GRESO\(Zheng et al\.,[2025](https://arxiv.org/html/2607.27610#bib.bib41)\)\(green curve\), rely on lightweight proxy signals to identify prompts near a target difficulty level without requiring additional rolloutsfor selection\. However, they model each prompt’s success rate as a stationary latent parameter, relying on a heuristic decay factor as a substitute for principled non\-stationarity handling,which may yield inaccurate difficulty estimates as the policy evolves\.To quantify this estimation inaccuracy, we calculate the MAE between predicted success\-rate \(before rollouts\) and empirical success\-rate \(after rollouts\) at every training step in Fig\.[1](https://arxiv.org/html/2607.27610#S0.F1)\(b\)\.As can be seen, this static modelingQu et al\.\([2025](https://arxiv.org/html/2607.27610#bib.bib28)\)assumption results in persistently high estimation error throughout training \(MAE≈0\.40\\text\{MAE\}\\approx 0\.40in purple curve\), failing to improve as the policy evolves\.These limitations call for a principled shift from static parameter estimation to dynamic state tracking,one that preserves the efficiency of inference\-free selection while better matching the non\-stationary nature of RL finetuning\.

![Refer to caption](https://arxiv.org/html/2607.27610v1/x3.png)Figure 2:Framework Overview of KGPS\.At each training steptt, KGPS maintains a Gaussian posterior𝒩​\(ψ^τt,Pτt\)\\mathcal\{N\}\(\\hat\{\\psi\}\_\{\\tau\}^\{t\},P\_\{\\tau\}^\{t\}\)over the latent logit\-space difficulty of every candidate prompt via a Kalman filter\. Prompts selected in the previous step receive a full Kalman update from their rollout observations; unselected prompts undergo only the prediction step, with variance inflated byQtQ\_\{t\}to reflect policy\-induced uncertainty growth\. The updated posteriors drive prompt selection through the posterior\-expected utilityA~​\(τ,t\)=𝔼ψ∼𝒩​\(ψ^τt,Pτt\)​\[h​\(ψ\)\]\\tilde\{A\}\(\\tau,t\)=\\mathbb\{E\}\_\{\\psi\\sim\\mathcal\{N\}\(\\hat\{\\psi\}\_\{\\tau\}^\{t\},P\_\{\\tau\}^\{t\}\)\}\[h\(\\psi\)\], whereh​\(ψ\)=σ​\(ψ\)​\(1−σ​\(ψ\)\)h\(\\psi\)=\\sigma\(\\psi\)\(1\-\\sigma\(\\psi\)\)peaks at intermediate difficulty, ensuring the selected batch targets the most informative training signal while naturally revisiting long\-neglected prompts\.To address this issue,in this paper, we propose*Kalman\-GuidedPromptSelection*\(*KGPS*\), a principled framework for online prompt selection that reframes difficulty tracking as a dynamic state estimation problem,as illustrated in Fig\.[2](https://arxiv.org/html/2607.27610#S1.F2)\.Rather than modeling each prompt’s success rate as a stationary latent parameter, KGPS adopts a linear\-Gaussian state\-space formulation in logit space, where the latent success rate of each prompt evolves as a random walk whose process noise is explicitly coupled to the magnitude of the policy update at each training step\. This design ensures that the posterior uncertainty of every prompt, including those not selected for training, is continuously inflated to reflect the current state of the policy\. Building on this formulation, we derive a Kalman filter that maintains a Gaussian belief over each prompt’s logit\-transformed success rate, propagating uncertainty in a theoretically grounded manner\. Prompt selection is then performed by maximizing a posterior\-expected training utility score that prioritizes prompts at intermediate difficulty under the current policy, while naturally promoting exploration of prompts with accumulated uncertainty via posterior variance inflation\.Importantly, KGPS operates on the same rollout outcomes already generated for policy training and requires no additional rollouts for prompt selection\.As illustrated in Fig\.[1](https://arxiv.org/html/2607.27610#S0.F1), KGPS substantially reduces success rate estimation error\(red curve in Fig\.[1](https://arxiv.org/html/2607.27610#S0.F1)\(b\)\)and achieves superior performance under an equivalent rollout budget\(red curve in Fig\.[1](https://arxiv.org/html/2607.27610#S0.F1)\(a\)\)compared to MoPPS\(Qu et al\.,[2025](https://arxiv.org/html/2607.27610#bib.bib28)\), a representative prediction\-based method that relies on static posterior approximation, empirically validating the advantage of dynamic state estimation\. The primary contributions of this work are three\-fold:

- •We reformulate prompt difficulty tracking as a dynamic state estimation problem, modeling each prompt’s latent success rate via a linear\-Gaussian state\-space model with process noise coupled to the policy update magnitude\.
- •We derive a Kalman filter for online difficulty tracking and propose a posterior\-expected training utility score that naturally balances exploitation of current estimates with exploration of uncertain prompts, requiring no hand\-tuned decay schedule\.
- •Extensive experiments across diverse reasoning tasks and RL algorithms demonstrate that KGPS consistently outperforms existing baselines in both accuracy and rollout efficiency with negligible computational overhead\.

## 2Related Work

#### RL Finetuning of LLMs\.

Reinforcement learning has become a central paradigm for improving large language models, from alignment\-oriented RLHF\(Dong et al\.,[2024](https://arxiv.org/html/2607.27610#bib.bib6);Dai et al\.,[2023](https://arxiv.org/html/2607.27610#bib.bib4);Zheng et al\.,[2023](https://arxiv.org/html/2607.27610#bib.bib42)\)to reasoning\-oriented RLVR\(Jaech et al\.,[2024](https://arxiv.org/html/2607.27610#bib.bib14);Guo et al\.,[2025](https://arxiv.org/html/2607.27610#bib.bib8);Team et al\.,[2025](https://arxiv.org/html/2607.27610#bib.bib34);Pan et al\.,[2025](https://arxiv.org/html/2607.27610#bib.bib26)\)\. PPO\(Schulman et al\.,[2017](https://arxiv.org/html/2607.27610#bib.bib29)\)remains a standard backbone\. In the meantime, GRPO\(Shao et al\.,[2024](https://arxiv.org/html/2607.27610#bib.bib30);DeepSeek\-AI et al\.,[2025](https://arxiv.org/html/2607.27610#bib.bib5)\)offers a lighter alternative\. Recent work further improves RL finetuning in terms of stability, sample efficiency, and scalability\(Yu et al\.,[2025](https://arxiv.org/html/2607.27610#bib.bib38);Hu,[2025](https://arxiv.org/html/2607.27610#bib.bib12);Yue et al\.,[2025](https://arxiv.org/html/2607.27610#bib.bib39);Kazemnejad et al\.,[2024](https://arxiv.org/html/2607.27610#bib.bib15);Sheng et al\.,[2024](https://arxiv.org/html/2607.27610#bib.bib33)\)\.

#### Prompt Selection for RL Finetuning\.

Data curation has emerged as a promising strategy for improving RL finetuning efficiency\.Offlinemethods select prompts prior to training based on static criteria such as difficulty or diversity\(Ye et al\.,[2025](https://arxiv.org/html/2607.27610#bib.bib37);Li et al\.,[2025](https://arxiv.org/html/2607.27610#bib.bib17);Hu et al\.,[2025](https://arxiv.org/html/2607.27610#bib.bib13);Yang et al\.,[2024](https://arxiv.org/html/2607.27610#bib.bib36);Fatemi et al\.,[2025](https://arxiv.org/html/2607.27610#bib.bib7)\), but cannot adapt to the model’s evolving capabilities\.Onlinemethods address this by dynamically curating training batches, either filtering prompts with degenerate success rates\(Yu et al\.,[2025](https://arxiv.org/html/2607.27610#bib.bib38);Liu et al\.,[2025](https://arxiv.org/html/2607.27610#bib.bib19);Cui et al\.,[2025](https://arxiv.org/html/2607.27610#bib.bib3);Meng et al\.,[2025](https://arxiv.org/html/2607.27610#bib.bib25)\)or prioritizing intermediate\-difficulty prompts to maximize gradient informativeness\(Bae et al\.,[2025](https://arxiv.org/html/2607.27610#bib.bib1);Chen et al\.,[2025](https://arxiv.org/html/2607.27610#bib.bib2)\), yet typically require additional rollouts over a large candidate pool, introducing substantial computational overhead\. MoPPS\(Qu et al\.,[2025](https://arxiv.org/html/2607.27610#bib.bib28)\)mitigates this cost via inference\-free prompt selection using Beta–Bernoulli posteriors with Thompson Sampling, but its static formulation assumes time\-invariant prompt difficulty, failing to account for the non\-stationary evolution induced by a continuously improving policy\.

Our work is most closely related to these online prediction\-based methods, butdiffersin one key aspect: instead of static parameter estimation with heuristic forgetting, we formulate prompt difficulty tracking as a dynamic state estimation problem\. This allows uncertainty to evolve with policy updates and yields more accurate prompt difficulty estimates without requiring additional rollouts\.

## 3Methodology

This section presentsKGPS\(Kalman\-GuidedPromptSelection\), as shown in Figure[2](https://arxiv.org/html/2607.27610#S1.F2),our framework for addressing the non\-stationarity problem identified in Sec\.[3\.1](https://arxiv.org/html/2607.27610#S3.SS1)\. The key idea is to replace static difficulty estimation with dynamic state tracking\. Concretely, KGPS proceeds in three stages:\(i\) modeling each prompt’s time\-varying success rate as a latent state in a linear\-Gaussian state\-space model \(Sec\.[3\.2](https://arxiv.org/html/2607.27610#S3.SS2)\); \(ii\) applying a Kalman filter to propagate a Gaussian posterior over every prompt’s logit success rate as the policy evolves \(Sec\.[3\.3](https://arxiv.org/html/2607.27610#S3.SS3)\); and \(iii\) selecting prompts by maximizing a posterior\-expected training utility score derived from this posterior \(Sec\.[3\.4](https://arxiv.org/html/2607.27610#S3.SS4)\)\.

### 3\.1Preliminary

Setup and notation\.Let𝒯=\{τi\}i=1N\\mathcal\{T\}=\\\{\\tau\_\{i\}\\\}\_\{i=1\}^\{N\}denote the prompt pool of sizeNN,πθt\\pi\_\{\\theta\_\{t\}\}is the policy at training steptt, and𝒯Bt⊂𝒯\\mathcal\{T\}^\{t\}\_\{B\}\\subset\\mathcal\{T\}denotes the selected batch of sizeBB\. For each selected promptτ∈𝒯Bt\\tau\\in\\mathcal\{T\}^\{t\}\_\{B\}, the policy generateskkindependent rollouts, each scored via a binary reward following a Bernoulli distribution\(Qu et al\.,[2025](https://arxiv.org/html/2607.27610#bib.bib28)\):

rτt,j∼Bernoulli​\(ϕτt\),j=1,…,k,r\_\{\\tau\}^\{t,j\}\\;\\sim\\;\\mathrm\{Bernoulli\}\(\\phi\_\{\\tau\}^\{t\}\),\\qquad j=1,\\ldots,k,\(1\)whereϕτt∈\[0,1\]\\phi\_\{\\tau\}^\{t\}\\in\[0,1\]is the*latent success rate*of promptτ\\tauunderπθt\\pi\_\{\\theta\_\{t\}\}, serving as a surrogate for its difficulty\. The empirical success count issτt=∑j=1krτt,js\_\{\\tau\}^\{t\}=\\sum\_\{j=1\}^\{k\}r\_\{\\tau\}^\{t,j\}, so thatsτt∣ϕτt∼Binomial​\(k,ϕτt\)s\_\{\\tau\}^\{t\}\\mid\\phi\_\{\\tau\}^\{t\}\\;\\sim\\;\\mathrm\{Binomial\}\(k,\\,\\phi\_\{\\tau\}^\{t\}\)\.

Bayesian online prompt selection\.Existing methods\(Qu et al\.,[2025](https://arxiv.org/html/2607.27610#bib.bib28)\)model each prompt as an arm in a Bernoulli bandit and maintain a Beta–Bernoulli conjugate posterior overϕτt\\phi\_\{\\tau\}^\{t\}\. Initialized with a uniform priorϕτ0∼Beta​\(1,1\)\\phi\_\{\\tau\}^\{0\}\\sim\\mathrm\{Beta\}\(1,1\), the posterior admits a closed\-form recursive update:

ϕτt∣ℋt∼Beta​\(ατt,βτt\),ατt\+1=λ​ατt\+sτt,βτt\+1=λ​βτt\+k−sτt,\\phi\_\{\\tau\}^\{t\}\\mid\\mathcal\{H\}\_\{t\}\\;\\sim\\;\\mathrm\{Beta\}\(\\alpha\_\{\\tau\}^\{t\},\\,\\beta\_\{\\tau\}^\{t\}\),\\qquad\\alpha\_\{\\tau\}^\{t\+1\}=\\lambda\\,\\alpha\_\{\\tau\}^\{t\}\+s\_\{\\tau\}^\{t\},\\quad\\beta\_\{\\tau\}^\{t\+1\}=\\lambda\\,\\beta\_\{\\tau\}^\{t\}\+k\-s\_\{\\tau\}^\{t\},\(2\)whereℋt=\{𝒯iB,ℛiB\}i=0t\\mathcal\{H\}\_\{t\}=\\\{\\mathcal\{T\}\_\{i\}^\{B\},\\mathcal\{R\}\_\{i\}^\{B\}\\\}\_\{i=0\}^\{t\}is the optimization history andλ∈\(0,1\)\\lambda\\in\(0,1\)is a decay factor introduced to discount stale observations\. Prompts are then selected by preferring posterior meansϕ^τt=ατt/\(ατt\+βτt\)\\hat\{\\phi\}\_\{\\tau\}^\{t\}=\\alpha\_\{\\tau\}^\{t\}/\(\\alpha\_\{\\tau\}^\{t\}\+\\beta\_\{\\tau\}^\{t\}\)near the targetϕ∗≈0\.5\\phi^\{\*\}\\approx 0\.5\(Bae et al\.,[2025](https://arxiv.org/html/2607.27610#bib.bib1);Chen et al\.,[2025](https://arxiv.org/html/2607.27610#bib.bib2)\), where gradient informativeness is highest\.

#### Limitation: the non\-stationarity problem\.

The Beta–Bernoulli framework implicitly treatsϕτt\\phi\_\{\\tau\}^\{t\}as a stationary latent parameter; while the decay factorλ\\lambdadown\-weights stale observations, it provides no principled mechanism for predicting howϕτt\\phi\_\{\\tau\}^\{t\}evolves between selections\. As shown in Figure[1](https://arxiv.org/html/2607.27610#S0.F1)\(b\), this leads to persistently high estimation error throughout training\. In early training, the Beta posterior remains diffuse despite warmup initialization\. As MoPPS\(Qu et al\.,[2025](https://arxiv.org/html/2607.27610#bib.bib28)\)acknowledge, estimation accuracy improves only as the pseudo\-countCτt=ατt\+βτtC\_\{\\tau\}^\{t\}=\\alpha\_\{\\tau\}^\{t\}\+\\beta\_\{\\tau\}^\{t\}grows, a condition that is rarely met given the limited rollouts available per prompt in practice\. Thompson Sampling further compounds the issue by introducing additional stochastic variance into the predicted success rates\. In later training, the fixedλ\\lambdadiscounts observations at a constant rate regardless of policy update magnitude, failing to track the non\-stationarity difficulty drift induced by a continuously improving policyπθt\\pi\_\{\\theta\_\{t\}\}\. Resolving this model–reality mismatch requires reframing prompt difficulty tracking from static parameter estimation to*dynamic state estimation*, which is the central objective of KGPS\.

### 3\.2State\-Space Formulation in Logit Space

We treat each prompt’s success rate as atime\-varying latent stateand adopt a linear\-Gaussian state\-space model \(SSM\) to track it\. This choice is motivated by two considerations: \(i\) SSMs provide a principled, closed\-form mechanism for propagating uncertainty under non\-stationarity, without recourse to hand\-tuned decay schedules; and \(ii\) linear\-GaussianSSMs admit the Kalman filter under Gaussian approximation, producing a well\-calibrated posterior that Sec\.[3\.4](https://arxiv.org/html/2607.27610#S3.SS4)subsequently uses to score and select prompts\. Letϕτt∈\[0,1\]\\phi\_\{\\tau\}^\{t\}\\in\[0,1\]denote the expected success rate of promptτ\\tauunder the current policyπθt\\pi\_\{\\theta\_\{t\}\}at steptt\. We work in logit space withψτt=logit​\(ϕτt\)∈\(−∞,\+∞\)\\psi\_\{\\tau\}^\{t\}=\\mathrm\{logit\}\(\\phi\_\{\\tau\}^\{t\}\)\\in\(\-\\infty,\+\\infty\)as the latent state, satisfying the unconstrained domain requirement of the linear\-Gaussian framework\.

State evolution\.We model the latent logitψτt\\psi\_\{\\tau\}^\{t\}as a random walk:

ψτt=ψτt−1\+wτt,wτt∼𝒩​\(0,Qt\),\\psi\_\{\\tau\}^\{t\}=\\psi\_\{\\tau\}^\{t\-1\}\+w\_\{\\tau\}^\{t\},\\qquad w\_\{\\tau\}^\{t\}\\sim\\mathcal\{N\}\(0,\\,Q\_\{t\}\),\(3\)where the process noise variance is tied to the magnitude of the preceding policy update:

Qt=γ⋅‖θt−θt−1‖2,Q\_\{t\}=\\gamma\\cdot\\\|\\theta\_\{t\}\-\\theta\_\{t\-1\}\\\|\_\{2\},\(4\)withγ\>0\\gamma\>0a scalar hyperparameter\. This coupling reflects a natural intuition: a larger policy step is likely to shift prompt difficulty more, warranting greater state uncertainty\. Note thatQtQ\_\{t\}depends only on the*previous*parameter update, which is available at the start of stepttbefore any new rollouts are collected\. This ensures the prediction step \(Sec\.[3\.3](https://arxiv.org/html/2607.27610#S3.SS3)\) is executable without circular dependency\.

#### Observation model\.

At steptt, if promptτ\\tauis selected withkkrollouts yieldingsτts\_\{\\tau\}^\{t\}successes, the empirical success rateϕ^τt=sτt/k\\hat\{\\phi\}\_\{\\tau\}^\{t\}=s\_\{\\tau\}^\{t\}/kfollows a scaled Binomial distribution \(as defined in Sec\.[3\.1](https://arxiv.org/html/2607.27610#S3.SS1)\), which is incompatible with the Gaussian observation requirement of the linear\-Gaussian framework\. We therefore apply the delta method to the logit\-transformedϕ^τt\\hat\{\\phi\}\_\{\\tau\}^\{t\}, which propagates the known varianceVar​\(ϕ^τt\)=ϕτt​\(1−ϕτt\)/k\\mathrm\{Var\}\(\\hat\{\\phi\}\_\{\\tau\}^\{t\}\)=\\phi\_\{\\tau\}^\{t\}\(1\-\\phi\_\{\\tau\}^\{t\}\)/kthrough the logit transform to yield a tractable Gaussian approximation in logit space \(derivation in Appendix[C](https://arxiv.org/html/2607.27610#S3a)\):

logit​\(ϕ^τt\)=ψτt\+ετt,ετt∼𝒩​\(0,Rτt\)\.\\mathrm\{logit\}\(\\hat\{\\phi\}\_\{\\tau\}^\{t\}\)=\\psi\_\{\\tau\}^\{t\}\+\\varepsilon\_\{\\tau\}^\{t\},\\qquad\\varepsilon\_\{\\tau\}^\{t\}\\sim\\mathcal\{N\}\(0,\\,R\_\{\\tau\}^\{t\}\)\.\(5\)
Since the observation noise varianceRτt=1/\[k​ϕτt​\(1−ϕτt\)\]R\_\{\\tau\}^\{t\}=1/\[k\\phi\_\{\\tau\}^\{t\}\(1\-\\phi\_\{\\tau\}^\{t\}\)\]depends on the unknownϕτt\\phi\_\{\\tau\}^\{t\}, we substitute the current posterior meanψ^τt\\hat\{\\psi\}\_\{\\tau\}^\{t\}at the start of stepttbefore rollouts are collected, produced by the Kalman filter at stept−1t\-1\(see Sec\.[3\.3](https://arxiv.org/html/2607.27610#S3.SS3)\)\.ψ^τt\\hat\{\\psi\}\_\{\\tau\}^\{t\}is in fact a plug\-in estimator that models the empirical success rateϕ^τt\\hat\{\\phi\}\_\{\\tau\}^\{t\}as:

ϕ~τt=c​l​i​p​\(σ​\(ψ^τt\),δ,1−δ\),\\tilde\{\\phi\}\_\{\\tau\}^\{t\}=\{clip\}\(\\sigma\(\\hat\{\\psi\}\_\{\\tau\}^\{t\}\),\\,\\delta,\\,1\-\\delta\),\(6\)whereδ=1/\(2​k\)\\delta=1/\(2k\)ensuresRτtR\_\{\\tau\}^\{t\}remains bounded above, andσ​\(⋅\)\\sigma\(\\cdot\)denotes the sigmoid function mapping from logit space back to\[0,1\]\[0,1\]\. Therefore, the observation noise variance is formulated as:

Rτt=1k​ϕ~τt​\(1−ϕ~τt\)\.R\_\{\\tau\}^\{t\}=\\frac\{1\}\{k\\,\\tilde\{\\phi\}\_\{\\tau\}^\{t\}\(1\-\\tilde\{\\phi\}\_\{\\tau\}^\{t\}\)\}\.\(7\)

### 3\.3Kalman Filtering for Posterior Maintenance

With the SSM of Eqs\. \([3](https://arxiv.org/html/2607.27610#S3.E3)\)–\([7](https://arxiv.org/html/2607.27610#S3.E7)\) in hand, we apply a Kalman filter \(KF\) to maintain a Gaussian posterior𝒩​\(ψ^τt,Pτt\)\\mathcal\{N\}\(\\hat\{\\psi\}\_\{\\tau\}^\{t\},P\_\{\\tau\}^\{t\}\)over the latent logit stateψτt\\psi\_\{\\tau\}^\{t\}for every promptτ∈𝒯\\tau\\in\\mathcal\{T\}, whereψ^τt\\hat\{\\psi\}\_\{\\tau\}^\{t\}is the posterior mean estimate andPτtP\_\{\\tau\}^\{t\}is the posterior variance\. The filter’s responsibility is to maintain well\-calibrated posteriors as the policy evolves, with the resulting pair\(ψ^τt,Pτt\)\(\\hat\{\\psi\}\_\{\\tau\}^\{t\},P\_\{\\tau\}^\{t\}\)subsequently serving as the input to the training utility score and prompt selection in Sec\.[3\.4](https://arxiv.org/html/2607.27610#S3.SS4)\. Since both the state transition and observation maps are linear with identity coefficients, the KF reduces to simple scalar recursions\.

#### Prediction step \(all prompts, every step\)\.

At each training steptt, we inflate the posterior variance of*all*prompts to account for the latest policy update, regardless of whether they were selected:

Pτt∣t−1=Pτt−1\+Qt,∀τ∈𝒯\.P\_\{\\tau\}^\{t\\mid t\-1\}=P\_\{\\tau\}^\{t\-1\}\+Q\_\{t\},\\qquad\\forall\\,\\tau\\in\\mathcal\{T\}\.\(8\)BecauseQt\>0Q\_\{t\}\>0, the predicted variance of every prompt grows at each step, even for those not recently evaluated\. This monotonic growth is what allows the learnability score in Sec\.[3\.4](https://arxiv.org/html/2607.27610#S3.SS4)to naturally favour prompts that have been neglected: their inflatedPτt\|t−1P\_\{\\tau\}^\{t\|t\-1\}raises their score deterministically until they are eventually selected\.

#### Update step \(selected prompts only\)\.

For eachτ∈𝒯t−1B\\tau\\in\\mathcal\{T\}\_\{t\-1\}^\{B\},i\.e\., prompts selected at stept−1t\-1, the filter incorporates the logit\-space observation via the standard KF equations:

Kτt=Pτt∣t−1Pτt∣t−1\+Rτt−1,ψ^τt=ψ^τt∣t−1\+Kτt⋅ντt,Pτt=\(1−Kτt\)​Pτt∣t−1\.K\_\{\\tau\}^\{t\}=\\frac\{P\_\{\\tau\}^\{t\\mid t\-1\}\}\{P\_\{\\tau\}^\{t\\mid t\-1\}\+R\_\{\\tau\}^\{t\-1\}\},\\quad\\hat\{\\psi\}\_\{\\tau\}^\{t\}=\\hat\{\\psi\}\_\{\\tau\}^\{t\\mid t\-1\}\+K\_\{\\tau\}^\{t\}\\cdot\\nu\_\{\\tau\}^\{t\},\\quad P\_\{\\tau\}^\{t\}=\\bigl\(1\-K\_\{\\tau\}^\{t\}\\bigr\)\\,P\_\{\\tau\}^\{t\\mid t\-1\}\.\(9\)
whereψ^τt∣t−1=ψ^τt−1\\hat\{\\psi\}\_\{\\tau\}^\{t\\mid t\-1\}=\\hat\{\\psi\}\_\{\\tau\}^\{t\-1\}is the one\-step predicted mean andντt=logit​\(ϕ^τt−1\)−ψ^τt∣t−1\\nu\_\{\\tau\}^\{t\}=\\mathrm\{logit\}\(\\hat\{\\phi\}\_\{\\tau\}^\{t\-1\}\)\-\\hat\{\\psi\}\_\{\\tau\}^\{t\\mid t\-1\}is the innovation\. For prompts not selected at stept−1t\-1, the state estimate is unchanged and the variance retains its predicted valuePτt∣t−1P\_\{\\tau\}^\{t\\mid t\-1\}\. The Kalman gainKτt∈\(0,1\]K\_\{\\tau\}^\{t\}\\in\(0,1\]adaptively weights the new observation against the prior belief\. When the rollout is near\-degenerate, the clipping in Eq\. \([7](https://arxiv.org/html/2607.27610#S3.E7)\) ensuresRτt−1R\_\{\\tau\}^\{t\-1\}remains bounded, so the filter retains a non\-negligible update even for extreme observations\. Conversely, when the prior is diffuse \(Pτt∣t−1P\_\{\\tau\}^\{t\\mid t\-1\}large\) and the observation is reliable \(moderateϕ^τt−1\\hat\{\\phi\}\_\{\\tau\}^\{t\-1\}, smallRτt−1R\_\{\\tau\}^\{t\-1\}\),Kτt→1K\_\{\\tau\}^\{t\}\\to 1and the posterior mean shifts aggressively toward the new data\. In both cases, the updated\(ψ^τt,Pτt\)\(\\hat\{\\psi\}\_\{\\tau\}^\{t\},P\_\{\\tau\}^\{t\}\)faithfully encodes the filter’s current state of knowledge about each prompt’s difficulty, which is precisely what Sec\.[3\.4](https://arxiv.org/html/2607.27610#S3.SS4)requires\.

### 3\.4Prompt Selection via Posterior\-Expected Utility

At each training steptt, the Kalman filter delivers a posterior𝒩​\(ψ^τt,Pτt\)\\mathcal\{N\}\(\\hat\{\\psi\}\_\{\\tau\}^\{t\},P\_\{\\tau\}^\{t\}\)for each prompt\. We now show how to convert this posterior into a scalar training utility score and use it to select the most informative batch for the next training step\.

#### From point estimates to posterior expectations\.

Following prior work\(Qu et al\.,[2025](https://arxiv.org/html/2607.27610#bib.bib28);Bae et al\.,[2025](https://arxiv.org/html/2607.27610#bib.bib1);Chen et al\.,[2025](https://arxiv.org/html/2607.27610#bib.bib2)\), training efficiency is maximized by selecting prompts at intermediate difficulty, where the current policy is neither saturated nor overwhelmed\. This motivates scoring each prompt by its estimated training utilityh​\(ψ^τt\)h\(\\hat\{\\psi\}\_\{\\tau\}^\{t\}\), whereh​\(ψ\)=σ​\(ψ\)​\(1−σ​\(ψ\)\)h\(\\psi\)=\\sigma\(\\psi\)\(1\-\\sigma\(\\psi\)\)peaks at intermediate success rates and vanishes at the extremes\. This criterion, however, ignores the posterior variancePτtP\_\{\\tau\}^\{t\}: a prompt with a near\-degenerate success rate is persistently excluded from future selection even when its true difficulty under the current policy remains unknown\. We resolve this by replacing the point evaluation with its expectation under the full Gaussian posterior maintained by the filter:

A~​\(τ,t\)=𝔼ψ∼𝒩​\(ψ^τt,Pτt\)​\[h​\(ψ\)\]\.\\tilde\{A\}\(\\tau,\\,t\)=\\mathbb\{E\}\_\{\\psi\\,\\sim\\,\\mathcal\{N\}\(\\hat\{\\psi\}\_\{\\tau\}^\{t\},\\,P\_\{\\tau\}^\{t\}\)\}\\\!\\bigl\[h\(\\psi\)\\bigr\]\.\(10\)Becauseh​\(ψ\)h\(\\psi\)is nonlinear and bounded,A~​\(τ,t\)\\tilde\{A\}\(\\tau,t\)is greater thanh​\(ψ^τt\)h\(\\hat\{\\psi\}\_\{\\tau\}^\{t\}\)wheneverPτt\>0P\_\{\\tau\}^\{t\}\>0: integrating over the posterior spreads probability mass toward the intermediate\-difficulty region wherehhis large, so any prompt with accumulated uncertainty retains a non\-negligible score regardless of where its mean currently sits\. Crucially, no stochastic sampling is needed: the prediction step in Eq\. \([8](https://arxiv.org/html/2607.27610#S3.E8)\) continuously inflatesPτtP\_\{\\tau\}^\{t\}for unobserved prompts, which raisesA~​\(τ,t\)\\tilde\{A\}\(\\tau,t\)deterministically until the prompt is revisited\.

#### Efficient computation via Gauss–Hermite quadrature\.

The expectation in Eq\. \([10](https://arxiv.org/html/2607.27610#S3.E10)\) has no closed form\. Sinceh​\(ψ\)h\(\\psi\)is a smooth, bounded function of a scalar Gaussian variable, we evaluate it efficiently via five\-point Gauss–Hermite quadrature \(detailed in Appendix[D](https://arxiv.org/html/2607.27610#S4a)\)\. Substitutingψ=ψ^τt\+2​Pτt​x\\psi=\\hat\{\\psi\}\_\{\\tau\}^\{t\}\+\\sqrt\{2P\_\{\\tau\}^\{t\}\}\\,xyields:

A~​\(τ,t\)≈1π​∑i=15wi​h​\(ψ^τt\+2​Pτt​xi\),\\tilde\{A\}\(\\tau,\\,t\)\\approx\\frac\{1\}\{\\sqrt\{\\pi\}\}\\sum\_\{i=1\}^\{5\}w\_\{i\}\\,h\\\!\\left\(\\hat\{\\psi\}\_\{\\tau\}^\{t\}\+\\sqrt\{2P\_\{\\tau\}^\{t\}\}\\,x\_\{i\}\\right\),\(11\)where the node–weight pairs\(xi,wi\)∈\{\(0,0\.9453\),\(±0\.9586,0\.3936\),\(±2\.0202,0\.0200\)\}\(x\_\{i\},w\_\{i\}\)\\in\\\{\(0,\\;0\.9453\),\\;\(\\pm 0\.9586,\\;0\.3936\),\\;\(\\pm 2\.0202,\\;0\.0200\)\\\}are fixed constants requiring no tuning\. The per\-prompt cost is five scalar evaluations ofh​\(ψ\)h\(\\psi\), negligible relative to LLM rollout overhead\.

#### Prompt selection\.

At each steptt, we computeA~​\(τ,t\)\\tilde\{A\}\(\\tau,t\)for everyτ∈𝒯\\tau\\in\\mathcal\{T\}and select the training batch by greedy maximization:

𝒯Bt=Top−⁡B​\{τ∈𝒯\|A~​\(τ,t\)\}\.\\mathcal\{T\}^\{t\}\_\{B\}=\\operatorname\{Top\-\}B\\bigl\\\{\\tau\\in\\mathcal\{T\}\\;\\big\|\\;\\tilde\{A\}\(\\tau,\\,t\)\\bigr\\\}\.\(12\)Before selection begins, we run a single warmup epoch in which prompts are sampled uniformly to obtain an initial empirical success rateϕ^τ0\\hat\{\\phi\}\_\{\\tau\}^\{0\}for everyτ∈𝒯\\tau\\in\\mathcal\{T\}\. The Kalman filter is then initialized withψ^τ0=logit​\(ϕ^τ0\)\\hat\{\\psi\}\_\{\\tau\}^\{0\}=\\mathrm\{logit\}\(\\hat\{\\phi\}\_\{\\tau\}^\{0\}\), providing a data\-driven starting point for the posterior mean rather than a uninformative init prior\. The full KGPS pipeline is summarized in Algorithm[1](https://arxiv.org/html/2607.27610#alg1)in the Appendix\.

## 4Experiment

![Refer to caption](https://arxiv.org/html/2607.27610v1/x4.png)

![Refer to caption](https://arxiv.org/html/2607.27610v1/x5.png)

![Refer to caption](https://arxiv.org/html/2607.27610v1/x6.png)

Figure 3:Test accuracy on the Math benchmark across three model scales under different data selection strategies\. KGPS \(ours\) consistently achieves top performance\.![Refer to caption](https://arxiv.org/html/2607.27610v1/x7.png)

![Refer to caption](https://arxiv.org/html/2607.27610v1/x8.png)

![Refer to caption](https://arxiv.org/html/2607.27610v1/x9.png)

Figure 4:Test accuracy on Countdown and Geometry benchmarks, where KGPS \(ours\) demonstrates consistent advantages over all baseline data selection strategies across varying model scales\.Table 1:Evaluation across mathematics benchmarks\. ‘\+’ represents finetuning with the method\.Boldindicates the best result in each column\.### 4\.1Experimental Setup

#### Datasets and Models\.

We evaluate KGPS on three reasoning benchmarks across distinct modalities\. Formathematics, models train on MATH\(Hendrycks et al\.,[2021](https://arxiv.org/html/2607.27610#bib.bib10)\)and test on MATH, MATH500, AIME 2024/2025, AMC 2023, Minerva Math, and OlympiadBench\(Hendrycks et al\.,[2021](https://arxiv.org/html/2607.27610#bib.bib10);Lightman et al\.,[2023](https://arxiv.org/html/2607.27610#bib.bib18);Mathematical Association of America,[2024](https://arxiv.org/html/2607.27610#bib.bib23),[2025](https://arxiv.org/html/2607.27610#bib.bib24),[2023](https://arxiv.org/html/2607.27610#bib.bib22);Lewkowycz et al\.,[2022](https://arxiv.org/html/2607.27610#bib.bib16);He et al\.,[2024](https://arxiv.org/html/2607.27610#bib.bib9)\)\. Forplanning, we use Countdown \(CD\-34\)\(Pan et al\.,[2025](https://arxiv.org/html/2607.27610#bib.bib26)\)\. Forvisual geometry, we use Geometry3k\(Lu et al\.,[2021](https://arxiv.org/html/2607.27610#bib.bib20);Hiyouga,[2025](https://arxiv.org/html/2607.27610#bib.bib11)\)\. Backbones: Qwen3\-0\.6B, Qwen3\-4B, Qwen3\-8B, and DeepSeek\-R1\-Distill\-Qwen\-7B for math; Qwen3\-0\.6B, and Qwen3\-4B for planning; Qwen3\-VL\-2B\-Instruct and Qwen2\.5\-VL\-7B\-Instruct for geometry\.

#### Training Protocol\.

Models are fine\-tuned onverlSheng et al\.\([2024](https://arxiv.org/html/2607.27610#bib.bib33)\)with88rollouts per prompt per step\. Test accuracy is reported as average pass@1 over 16 generations on held\-out sets\. See Appendix[B](https://arxiv.org/html/2607.27610#S2a)for details\.

#### Baselines\.

We compare against four baselines:Uniform\(random sampling\);GRESO\(Zheng et al\.,[2025](https://arxiv.org/html/2607.27610#bib.bib41)\)\(filters near\-degenerate prompts\);MoPPS\(Qu et al\.,[2025](https://arxiv.org/html/2607.27610#bib.bib28)\)\(Beta\-Bernoulli posteriors \+ Thompson Sampling\);DS \(Oracle\)\(Yu et al\.,[2025](https://arxiv.org/html/2607.27610#bib.bib38)\)\(oversamples and filters, higher rollout cost\)\. Our goal is to match or exceed DS accuracy under Uniform’s budget, showing dynamic state estimation achieves oracle quality without extra cost\.

### 4\.2Main Results

#### Mathematics\.

Fig\.[3](https://arxiv.org/html/2607.27610#S4.F3)shows KGPS maintains a superior accuracy trajectory across all model scales, matching or exceeding DS without real\-time feedback\. On DeepSeek\-R1\-Distill\-Qwen\-7B, KGPS avoids degradation seen in other methods\. Tab\.[1](https://arxiv.org/html/2607.27610#S4.T1)confirms KGPS achieves the highest average accuracy among selection methods\.Remarkably, KGPS matches or slightly surpasses DS, an evaluation\-based baseline with access to real\-time rollout feedback, while using 71%, 87%, and 83% fewer rollouts on Qwen3\-0\.6B, Qwen3\-4B, and DeepSeek\-R1\-Distill\-Qwen\-7B respectively, with average accuracy gains of 1\.00, 0\.88, and 0\.12 points\.Under the same 296k budget, KGPS outperforms MoPPS by 1\.67, 1\.55, 7\.36 points\. KGPS also surpasses Uniform and GRESO across all setups, confirming the benefit of uncertainty\-aware prompt selection over non\-adaptive strategies\.

![Refer to caption](https://arxiv.org/html/2607.27610v1/x10.png)

![Refer to caption](https://arxiv.org/html/2607.27610v1/x11.png)

![Refer to caption](https://arxiv.org/html/2607.27610v1/x12.png)

![Refer to caption](https://arxiv.org/html/2607.27610v1/x13.png)

![Refer to caption](https://arxiv.org/html/2607.27610v1/x14.png)

![Refer to caption](https://arxiv.org/html/2607.27610v1/x15.png)

Figure 5:Spearman correlation \(pp\-values\) between KGPS scores and sample difficulty across training steps\. High significant correlation validates KGPS as a difficulty\-aware selection criterion\.
#### Planning and visual geometry\.

As shown in Fig\.[4](https://arxiv.org/html/2607.27610#S4.F4), KGPS consistently outperforms all baselines on both the Countdown planning task and the Geometry3k visual reasoning benchmark\. On Countdown, KGPS achieves 87\.93% on Qwen3\-4B and 74\.85% on Qwen3\-0\.6B, outperforming the best competing baseline by 0\.70 and 0\.25 points, respectively\. On Geometry3k, KGPS achieves 58\.15% on Qwen3\-VL\-2B\-Instruct, surpassing the best baseline by 0\.60 points, confirming that its advantages transfer broadly across task modalities and model architectures\.

#### Difficulty estimation accuracy\.

Fig\.[5](https://arxiv.org/html/2607.27610#S4.F5)shows KGPS’s Spearman correlation withempiricalsuccess rates stabilizes at 0\.75–0\.85 \(lowpp\-values\) across all settings, , indicating that the Kalman filter maintains a statistically significant and accurate ordering of prompt difficulty under the evolving policy\. As shown in Fig\.[1](https://arxiv.org/html/2607.27610#S0.F1)\(b\), KGPS achieves lower MAE \(≈\\approx0\.15\) than MoPPS \(≈\\approx0\.40\), proving the superiority of dynamic modeling of success rate and difficulty tracking via the Kalman filter over static approximations\. See Appendix[F](https://arxiv.org/html/2607.27610#S6)for full comparison\.

#### Generalization across RL algorithms\.

To verify KGPS’s algorithm\-agnostic nature, we evaluate it on Countdown under PPO \(k=1k=1,k=8k=8\)Schulman et al\.\([2017](https://arxiv.org/html/2607.27610#bib.bib29)\)and Reinforce\+\+Hu\([2025](https://arxiv.org/html/2607.27610#bib.bib12)\)with Qwen3\-0\.6B/4B\. Tab\.[2](https://arxiv.org/html/2607.27610#S4.T2)shows KGPS consistently outperforms Uniform and MoPPS across all algorithms and scales, with gains up to \+2\.20 over MoPPS and \+5\.67 over Uniform\. Since KGPS only uses rollout rewards for posterior estimation, it is decoupled from the RL objective and easily integrable with any rollout\-based algorithm\.

Table 2:Evaluation on Countdown with PPO and Reinforce\+\+ using Qwen3\-0\.6B and Qwen3\-4B\.Boldindicates the best result in each row\.

### 4\.3Ablations

Table 3:Ablation on posterior estimation & warmup\.Table 4:Ablation on noise coupling\.Posterior expectation and warmup\.Tab\.[3](https://arxiv.org/html/2607.27610#S4.T3)compares the full posterior\-expected scoreA~​\(τ,t\)\\tilde\{A\}\(\\tau,t\)to the point\-estimate baselineh​\(ψ^τ\)h\(\\hat\{\\psi\}\_\{\\tau\}\)and evaluates warmup initialization\. Experiments are conducted on math dataset with Qwen3\-0\.6B\. Replacing posterior expectation with a point estimate drops performance by 1\.09 points \(73\.81 vs\. 72\.72\), confirming full posterior integration is essential for re\-admitting uncertain prompts\. Removing warmup causes an additional 0\.37–0\.48 point drop, confirming that stable initial posteriors benefit subsequent training\.

Process noise coupling\.Tab\.[4](https://arxiv.org/html/2607.27610#S4.T4)examines whether coupling the process noiseQtQ\_\{t\}to the policy update magnitude is beneficial\. Replacing the dynamicQt=γ​‖Δ​θ‖2Q\_\{t\}=\\gamma\\\|\\Delta\\theta\\\|^\{2\}with a fixed constant degrades average accuracy by1\.211\.21points \(73\.81 vs\. 72\.60\), demonstrating that adapting state uncertainty to the magnitude of each policy update is critical for accurate difficulty tracking under non\-stationarity\.

## 5Conclusion

This paper proposes KGPS, an online prompt selection framework for RL finetuning that treats difficulty tracking as dynamic state estimation\. It models each prompt’s success rate via a linear\-Gaussian SSM with policy\-update\-dependent noise, maintains Gaussian posteriors with a Kalman filter, and selects prompts via a posterior\-expected utility score—no extra LLM inference\. Experiments show KGPS consistently outperforms baselines in accuracy and efficiency, matching the DS oracle at far lower cost, and yielding gains under the same budget\. This validates dynamic state estimation as a practical alternative to static difficulty approximations\.

## References

- Bae et al\. \[2025\]S\. Bae, J\. Hong, M\. Y\. Lee, H\. Kim, J\. Nam, and D\. Kwak\.Online difficulty filtering for reasoning oriented reinforcement learning\.*arXiv preprint arXiv:2504\.03380*, 2025\.
- Chen et al\. \[2025\]X\. Chen, J\. Lu, M\. Kim, D\. Zhang, J\. Tang, A\. Piché, N\. Gontier, Y\. Bengio, and E\. Kamalloo\.Self\-evolving curriculum for llm reasoning\.*arXiv preprint arXiv:2505\.14970*, 2025\.
- Cui et al\. \[2025\]G\. Cui, L\. Yuan, Z\. Wang, H\. Wang, W\. Li, B\. He, Y\. Fan, T\. Yu, Q\. Xu, W\. Chen, et al\.Process reinforcement through implicit rewards\.*arXiv preprint arXiv:2502\.01456*, 2025\.
- Dai et al\. \[2023\]J\. Dai, X\. Pan, R\. Sun, J\. Ji, X\. Xu, M\. Liu, Y\. Wang, and Y\. Yang\.Safe rlhf: Safe reinforcement learning from human feedback\.*arXiv preprint arXiv:2310\.12773*, 2023\.
- DeepSeek\-AI et al\. \[2025\]DeepSeek\-AI, D\. Guo, D\. Yang, H\. Zhang, J\. Song, R\. Zhang, R\. Xu, Q\. Zhu, S\. Ma, P\. Wang, X\. Bi, X\. Zhang, X\. Yu, Y\. Wu, Z\. F\. Wu, Z\. Gou, Z\. Shao, Z\. Li, Z\. Gao, A\. Liu, B\. Xue, B\. Wang, B\. Wu, B\. Feng, C\. Lu, C\. Zhao, C\. Deng, C\. Zhang, C\. Ruan, D\. Dai, D\. Chen, D\. Ji, E\. Li, F\. Lin, F\. Dai, F\. Luo, G\. Hao, G\. Chen, G\. Li, H\. Zhang, H\. Bao, H\. Xu, H\. Wang, H\. Ding, H\. Xin, H\. Gao, H\. Qu, H\. Li, J\. Guo, J\. Li, J\. Wang, J\. Chen, J\. Yuan, J\. Qiu, J\. Li, J\. L\. Cai, J\. Ni, J\. Liang, J\. Chen, K\. Dong, K\. Hu, K\. Gao, K\. Guan, K\. Huang, K\. Yu, L\. Wang, L\. Zhang, L\. Zhao, L\. Wang, L\. Zhang, L\. Xu, L\. Xia, M\. Zhang, M\. Zhang, M\. Tang, M\. Li, M\. Wang, M\. Li, N\. Tian, P\. Huang, P\. Zhang, Q\. Wang, Q\. Chen, Q\. Du, R\. Ge, R\. Zhang, R\. Pan, R\. Wang, R\. J\. Chen, R\. L\. Jin, R\. Chen, S\. Lu, S\. Zhou, S\. Chen, S\. Ye, S\. Wang, S\. Yu, S\. Zhou, S\. Pan, S\. S\. Li, S\. Zhou, S\. Wu, S\. Ye, T\. Yun, T\. Pei, T\. Sun, T\. Wang, W\. Zeng, W\. Zhao, W\. Liu, W\. Liang, W\. Gao, W\. Yu, W\. Zhang, W\. L\. Xiao, W\. An, X\. Liu, X\. Wang, X\. Chen, X\. Nie, X\. Cheng, X\. Liu, X\. Xie, X\. Liu, X\. Yang, X\. Li, X\. Su, X\. Lin, X\. Q\. Li, X\. Jin, X\. Shen, X\. Chen, X\. Sun, X\. Wang, X\. Song, X\. Zhou, X\. Wang, X\. Shan, Y\. K\. Li, Y\. Q\. Wang, Y\. X\. Wei, Y\. Zhang, Y\. Xu, Y\. Li, Y\. Zhao, Y\. Sun, Y\. Wang, Y\. Yu, Y\. Zhang, Y\. Shi, Y\. Xiong, Y\. He, Y\. Piao, Y\. Wang, Y\. Tan, Y\. Ma, Y\. Liu, Y\. Guo, Y\. Ou, Y\. Wang, Y\. Gong, Y\. Zou, Y\. He, Y\. Xiong, Y\. Luo, Y\. You, Y\. Liu, Y\. Zhou, Y\. X\. Zhu, Y\. Xu, Y\. Huang, Y\. Li, Y\. Zheng, Y\. Zhu, Y\. Ma, Y\. Tang, Y\. Zha, Y\. Yan, Z\. Z\. Ren, Z\. Ren, Z\. Sha, Z\. Fu, Z\. Xu, Z\. Xie, Z\. Zhang, Z\. Hao, Z\. Ma, Z\. Yan, Z\. Wu, Z\. Gu, Z\. Zhu, Z\. Liu, Z\. Li, Z\. Xie, Z\. Song, Z\. Pan, Z\. Huang, Z\. Xu, Z\. Zhang, and Z\. Zhang\.Deepseek\-r1: Incentivizing reasoning capability in llms via reinforcement learning\.*arXiv preprint arXiv:2501\.12948*, 2025\.
- Dong et al\. \[2024\]H\. Dong, W\. Xiong, B\. Pang, H\. Wang, H\. Zhao, Y\. Zhou, N\. Jiang, D\. Sahoo, C\. Xiong, and T\. Zhang\.Rlhf workflow: From reward modeling to online rlhf\.*arXiv preprint arXiv:2405\.07863*, 2024\.
- Fatemi et al\. \[2025\]M\. Fatemi, B\. Rafiee, M\. Tang, and K\. Talamadupula\.Concise reasoning via reinforcement learning\.*arXiv preprint arXiv:2504\.05185*, 2025\.
- Guo et al\. \[2025\]D\. Guo, D\. Yang, H\. Zhang, J\. Song, R\. Zhang, R\. Xu, Q\. Zhu, S\. Ma, P\. Wang, X\. Bi, et al\.Deepseek\-r1: Incentivizing reasoning capability in llms via reinforcement learning\.*arXiv preprint arXiv:2501\.12948*, 2025\.
- He et al\. \[2024\]C\. He, R\. Luo, Y\. Bai, S\. Hu, Z\. L\. Thai, J\. Shen, J\. Hu, X\. Han, Y\. Huang, Y\. Zhang, et al\.Olympiadbench: A challenging benchmark for promoting agi with olympiad\-level bilingual multimodal scientific problems\.*arXiv preprint arXiv:2402\.14008*, 2024\.
- Hendrycks et al\. \[2021\]D\. Hendrycks, C\. Burns, S\. Kadavath, A\. Arora, S\. Basart, E\. Tang, D\. Song, and J\. Steinhardt\.Measuring mathematical problem solving with the math dataset\.*arXiv preprint arXiv:2103\.03874*, 2021\.
- Hiyouga \[2025\]Hiyouga\.Geometry3K: A large\-scale multi\-modal geometry reasoning dataset, 2025\.
- Hu \[2025\]J\. Hu\.Reinforce\+\+: A simple and efficient approach for aligning large language models\.*arXiv preprint arXiv:2501\.03262*, 2025\.
- Hu et al\. \[2025\]J\. Hu, Y\. Zhang, Q\. Han, D\. Jiang, X\. Zhang, and H\.\-Y\. Shum\.Open\-reasoner\-zero: An open source approach to scaling up reinforcement learning on the base model\.*arXiv preprint arXiv:2503\.24290*, 2025\.
- Jaech et al\. \[2024\]A\. Jaech, A\. Kalai, A\. Lerer, A\. Richardson, A\. El\-Kishky, A\. Low, A\. Helyar, A\. Madry, A\. Beutel, A\. Carney, et al\.Openai o1 system card\.*arXiv preprint arXiv:2412\.16720*, 2024\.
- Kazemnejad et al\. \[2024\]A\. Kazemnejad, M\. Aghajohari, E\. Portelance, A\. Sordoni, S\. Reddy, A\. Courville, and N\. L\. Roux\.Vineppo: Unlocking rl potential for llm reasoning through refined credit assignment\.*arXiv preprint arXiv:2410\.01679*, 2024\.
- Lewkowycz et al\. \[2022\]A\. Lewkowycz, A\. Andreassen, D\. Dohan, E\. Dyer, H\. Michalewski, V\. Ramasesh, A\. Slone, C\. Anil, I\. Schlag, T\. Gutman\-Solo, et al\.Solving quantitative reasoning problems with language models\.*Advances in Neural Information Processing Systems*, 35:3843–3857, 2022\.
- Li et al\. \[2025\]X\. Li, H\. Zou, and P\. Liu\.Limr: Less is more for rl scaling\.*arXiv preprint arXiv:2502\.11886*, 2025\.
- Lightman et al\. \[2023\]H\. Lightman, V\. Kosaraju, Y\. Burda, H\. Edwards, B\. Baker, T\. Lee, J\. Leike, J\. Schulman, I\. Sutskever, and K\. Cobbe\.Let’s verify step by step\.*arXiv preprint arXiv:2305\.20050*, 2023\.
- Liu et al\. \[2025\]M\. Liu, S\. Diao, X\. Lu, J\. Hu, X\. Dong, Y\. Choi, J\. Kautz, and Y\. Dong\.Prorl: Prolonged reinforcement learning expands reasoning boundaries in large language models\.*arXiv preprint arXiv:2505\.24864*, 2025\.
- Lu et al\. \[2021\]P\. Lu, R\. Gong, S\. Jiang, L\. Qiu, S\. Huang, X\. Liang, and S\.\-C\. Zhu\.Inter\-gps: Interpretable geometry problem solving with formal language and symbolic reasoning\.In*The Joint Conference of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing \(ACL\-IJCNLP 2021\)*, 2021\.
- Luo et al\. \[2025\]M\. Luo, S\. Tan, J\. Wong, X\. Shi, W\. Tang, M\. Roongta, C\. Cai, J\. Luo, T\. Zhang, E\. Li, R\. A\. Popa, and I\. Stoica\.Deepscaler: Surpassing o1\-preview with a 1\.5b model by scaling rl, 2025\.Notion Blog\.
- Mathematical Association of America \[2023\]Mathematical Association of America\.American mathematics competitions, 2023\.
- Mathematical Association of America \[2024\]Mathematical Association of America\.American invitational mathematics examination, 2024\.
- Mathematical Association of America \[2025\]Mathematical Association of America\.American invitational mathematics examination, 2025\.
- Meng et al\. \[2025\]F\. Meng, L\. Du, Z\. Liu, Z\. Zhou, Q\. Lu, D\. Fu, B\. Shi, W\. Wang, J\. He, K\. Zhang, et al\.Mm\-eureka: Exploring visual aha moment with rule\-based large\-scale reinforcement learning\.*CoRR*, 2025\.
- Pan et al\. \[2025\]J\. Pan, J\. Zhang, X\. Wang, L\. Yuan, H\. Peng, and A\. Suhr\.Tinyzero\.https://github\.com/Jiayi\-Pan/TinyZero, 2025\.Accessed: 2025\-01\-24\.
- Parashar et al\. \[2025\]S\. Parashar, S\. Gui, X\. Li, H\. Ling, S\. Vemuri, B\. Olson, E\. Li, Y\. Zhang, J\. Caverlee, D\. Kalathil, and S\. Ji\.Curriculum reinforcement learning from easy to hard tasks improves llm reasoning\.*arXiv preprint arXiv:2506\.06632*, 2025\.
- Qu et al\. \[2025\]Y\. Qu, Q\. Wang, Y\. Mao, V\. T\. Hu, B\. Ommer, and X\. Ji\.Can prompt difficulty be online predicted for accelerating rl finetuning of reasoning models?*arXiv preprint arXiv:2507\.04632*, 2025\.
- Schulman et al\. \[2017\]J\. Schulman, F\. Wolski, P\. Dhariwal, A\. Radford, and O\. Klimov\.Proximal policy optimization algorithms\.*arXiv preprint arXiv:1707\.06347*, 2017\.
- Shao et al\. \[2024\]Z\. Shao, P\. Wang, Q\. Zhu, R\. Xu, J\. Song, X\. Bi, H\. Zhang, M\. Zhang, Y\. Li, Y\. Wu, et al\.Deepseekmath: Pushing the limits of mathematical reasoning in open language models\.*arXiv preprint arXiv:2402\.03300*, 2024\.
- Shen et al\. \[2025a\]Q\. Shen, D\. Chen, Y\. Huang, Z\. Ling, Y\. Li, B\. Ding, and J\. Zhou\.Bots: A unified framework for bayesian online task selection in llm reinforcement finetuning\.*arXiv preprint arXiv:2510\.26374*, 2025a\.
- Shen et al\. \[2025b\]W\. Shen, J\. Pei, Y\. Peng, X\. Song, Y\. Liu, J\. Peng, H\. Sun, Y\. Hao, P\. Wang, J\. Zhang, and Y\. Zhou\.Skywork\-r1v3 technical report\.*arXiv preprint arXiv:2507\.06167*, 2025b\.
- Sheng et al\. \[2024\]G\. Sheng, C\. Zhang, Z\. Ye, X\. Wu, W\. Zhang, R\. Zhang, Y\. Peng, H\. Lin, and C\. Wu\.Hybridflow: A flexible and efficient rlhf framework\.*arXiv preprint arXiv: 2409\.19256*, 2024\.
- Team et al\. \[2025\]K\. Team, A\. Du, B\. Gao, B\. Xing, C\. Jiang, C\. Chen, C\. Li, C\. Xiao, C\. Du, C\. Liao, et al\.Kimi k1\. 5: Scaling reinforcement learning with llms\.*arXiv preprint arXiv:2501\.12599*, 2025\.
- Wen et al\. \[2025\]C\. Wen, T\. Guo, S\. Zhao, W\. Zou, and X\. Li\.Sari: Structured audio reasoning via curriculum\-guided reinforcement learning\.*arXiv preprint arXiv:2504\.15900*, 2025\.
- Yang et al\. \[2024\]A\. Yang, B\. Zhang, B\. Hui, B\. Gao, B\. Yu, C\. Li, D\. Liu, J\. Tu, J\. Zhou, J\. Lin, et al\.Qwen2\. 5\-math technical report: Toward mathematical expert model via self\-improvement\.*arXiv preprint arXiv:2409\.12122*, 2024\.
- Ye et al\. \[2025\]Y\. Ye, Z\. Huang, Y\. Xiao, E\. Chern, S\. Xia, and P\. Liu\.Limo: Less is more for reasoning\.*arXiv preprint arXiv:2502\.03387*, 2025\.
- Yu et al\. \[2025\]Q\. Yu, Z\. Zhang, R\. Zhu, Y\. Yuan, X\. Zuo, Y\. Yue, T\. Fan, G\. Liu, L\. Liu, X\. Liu, et al\.Dapo: An open\-source llm reinforcement learning system at scale\.*arXiv preprint arXiv:2503\.14476*, 2025\.
- Yue et al\. \[2025\]Y\. Yue, Y\. Yuan, Q\. Yu, X\. Zuo, R\. Zhu, W\. Xu, J\. Chen, C\. Wang, T\. Fan, Z\. Du, et al\.Vapo: Efficient and reliable reinforcement learning for advanced reasoning tasks\.*arXiv preprint arXiv:2504\.05118*, 2025\.
- Zeng et al\. \[2025\]W\. Zeng, Y\. Huang, Q\. Liu, W\. Liu, K\. He, Z\. Ma, and J\. He\.Simplerl\-zoo: Investigating and taming zero reinforcement learning for open base models in the wild\.*arXiv preprint arXiv:2503\.18892*, 2025\.
- Zheng et al\. \[2025\]H\. Zheng, Y\. Zhou, B\. R\. Bartoldson, B\. Kailkhura, F\. Lai, J\. Zhao, and B\. Chen\.Act only when it pays: Efficient reinforcement learning for LLM reasoning via selective rollouts\.*arXiv preprint arXiv:2506\.02177*, 2025\.
- Zheng et al\. \[2023\]R\. Zheng, S\. Dou, S\. Gao, Y\. Hua, W\. Shen, B\. Wang, Y\. Liu, S\. Jin, Q\. Liu, Y\. Zhou, et al\.Secrets of rlhf in large language models part i: Ppo\.*arXiv preprint arXiv:2307\.04964*, 2023\.

## Appendix

## AAlgorithm

We present the complete KGPS procedure in Algorithm[1](https://arxiv.org/html/2607.27610#alg1)\. At each training step, KGPS first inflates the posterior variance of all prompts proportionally to the magnitude of the latest policy update \(Lines 4–5\), then incorporates rollout observations from the previous step to refine the Kalman posteriors of selected prompts \(Lines 6–9\)\. The updated beliefs are used to compute posterior\-expected utility scores and select the most informative batch for the current step \(Lines 11–13\), whose rollout outcomes will in turn inform the Kalman update at the next iteration\.

Algorithm 1Kalman Guided Prompt Selection \(KGPS\)1:Prompt pool

𝒯=\{τi\}i=1N\\mathcal\{T\}=\\\{\\tau\_\{i\}\\\}\_\{i=1\}^\{N\}; initial variance

P0P\_\{0\}; process noise scale

γ\\gamma; batch size

BB; rollouts per prompt

kk; RL model

πθ0\\pi\_\{\\theta\_\{0\}\}; total steps

TT
2:Finetuned model

πθT\\pi\_\{\\theta\_\{T\}\}
3:Sample all

τ∈𝒯\\tau\\in\\mathcal\{T\}uniformly for one epoch; collect

ϕ^τ0=sτ0/k\\hat\{\\phi\}\_\{\\tau\}^\{0\}=s\_\{\\tau\}^\{0\}/k
4:

ψ^τ0←logit​\(ϕ^τ0\),Pτ0←P0\\hat\{\\psi\}\_\{\\tau\}^\{0\}\\leftarrow\\mathrm\{logit\}\(\\hat\{\\phi\}\_\{\\tau\}^\{0\}\),\\;P\_\{\\tau\}^\{0\}\\leftarrow P\_\{0\}for all

τ∈𝒯\\tau\\in\\mathcal\{T\};

𝒯0B←∅\\mathcal\{T\}\_\{0\}^\{B\}\\leftarrow\\emptyset
5:for

t=1t=1to

TTdo

6:% — Prediction step \(all prompts\) —

7:

Qt←γ​‖θt−θt−1‖2Q\_\{t\}\\leftarrow\\gamma\\\|\\theta\_\{t\}\-\\theta\_\{t\-1\}\\\|\_\{2\}
8:

Pτt\|t−1←Pτt−1\+QtP\_\{\\tau\}^\{t\|t\-1\}\\leftarrow P\_\{\\tau\}^\{t\-1\}\+Q\_\{t\}for all

τ∈𝒯\\tau\\in\\mathcal\{T\}
9:% — Kalman update \(selected prompts from stept−1t\{\-\}1\) —

10:foreach

τ∈𝒯t−1B\\tau\\in\\mathcal\{T\}\_\{t\-1\}^\{B\}do

11:

Rτt−1←\[k​ϕ~τt−1​\(1−ϕ~τt−1\)\]−1R\_\{\\tau\}^\{t\-1\}\\leftarrow\\bigl\[k\\,\\tilde\{\\phi\}\_\{\\tau\}^\{t\-1\}\(1\-\\tilde\{\\phi\}\_\{\\tau\}^\{t\-1\}\)\\bigr\]^\{\-1\}
12:

Kτt←Pτt\|t−1/\(Pτt\|t−1\+Rτt−1\)K\_\{\\tau\}^\{t\}\\leftarrow P\_\{\\tau\}^\{t\|t\-1\}\\big/\\bigl\(P\_\{\\tau\}^\{t\|t\-1\}\+R\_\{\\tau\}^\{t\-1\}\\bigr\);

ντt←logit​\(ϕ^τt−1\)−ψ^τt−1\\nu\_\{\\tau\}^\{t\}\\leftarrow\\mathrm\{logit\}\(\\hat\{\\phi\}\_\{\\tau\}^\{t\-1\}\)\-\\hat\{\\psi\}\_\{\\tau\}^\{t\-1\}
13:

ψ^τt←ψ^τt−1\+Kτt⋅ντt\\hat\{\\psi\}\_\{\\tau\}^\{t\}\\leftarrow\\hat\{\\psi\}\_\{\\tau\}^\{t\-1\}\+K\_\{\\tau\}^\{t\}\\cdot\\nu\_\{\\tau\}^\{t\};

Pτt←\(1−Kτt\)​Pτt\|t−1P\_\{\\tau\}^\{t\}\\leftarrow\(1\-K\_\{\\tau\}^\{t\}\)\\,P\_\{\\tau\}^\{t\|t\-1\}
14:endfor

15:

ψ^τt←ψ^τt−1,Pτt←Pτt\|t−1\\hat\{\\psi\}\_\{\\tau\}^\{t\}\\leftarrow\\hat\{\\psi\}\_\{\\tau\}^\{t\-1\},\\;P\_\{\\tau\}^\{t\}\\leftarrow P\_\{\\tau\}^\{t\|t\-1\}for all

τ∉𝒯t−1B\\tau\\notin\\mathcal\{T\}\_\{t\-1\}^\{B\}
16:% — Prompt selection and rollout —

17:Compute

A~​\(τ,t\)\\tilde\{A\}\(\\tau,t\)via Eq\. \([11](https://arxiv.org/html/2607.27610#S3.E11)\) using

\(ψ^τt,Pτt\)\(\\hat\{\\psi\}\_\{\\tau\}^\{t\},P\_\{\\tau\}^\{t\}\)for all

τ∈𝒯\\tau\\in\\mathcal\{T\}
18:

𝒯tB←Top​\-⁡B​\{τ∣A~​\(τ,t\)\}\\mathcal\{T\}\_\{t\}^\{B\}\\leftarrow\\operatorname\{Top\\text\{\-\}\}B\\\{\\tau\\mid\\tilde\{A\}\(\\tau,t\)\\\}
19:foreach

τ∈𝒯tB\\tau\\in\\mathcal\{T\}\_\{t\}^\{B\}do

20:Generate

kkrollouts from

πθt\\pi\_\{\\theta\_\{t\}\}; compute

ϕ^τt←sτt/k\\hat\{\\phi\}\_\{\\tau\}^\{t\}\\leftarrow s\_\{\\tau\}^\{t\}/k
21:endfor

22:Update

θt\\theta\_\{t\}via RL algorithm on

𝒯tB\\mathcal\{T\}\_\{t\}^\{B\}
23:endfor

## BImplementation Details

All experiments are conducted using the GRPO algorithm on theverlframeworkSheng et al\.\([2024](https://arxiv.org/html/2607.27610#bib.bib33)\)with NVIDIA H100 80GB GPUs\. For KGPS, we set the initial varianceP0=1\.0P\_\{0\}=1\.0, process noise scaleγ=0\.1\\gamma=0\.1, and rollouts per promptk=8k=8\. The learning rate is set to1×10−61\\times 10^\{\-6\}, and the maximum prompt length is 1024 tokens\. For the Math task, the maximum response length is set to 4096 tokens, while for Countdown and Geometry tasks, it is set to 1024 tokens\. For evaluation on the Math task, we assess generalization on six out\-of\-distribution benchmarks\. Inference is performed with temperature0\.60\.6and top\-pp1\.01\.0, and test accuracy is reported as average pass@1 over 16 independent generations\.

## CDerivation of the Observation Noise Variance

We derive the observation noise varianceRτtR\_\{\\tau\}^\{t\}in Eq\. \([7](https://arxiv.org/html/2607.27610#S3.E7)\) from first principles via the delta method\.

#### Binomial sampling\.

At steptt, promptτ\\tauis evaluated withkkindependent rollouts\. The success countsτts\_\{\\tau\}^\{t\}follows a binomial distribution:

sτt∼Binomial​\(k,ϕτt\),s\_\{\\tau\}^\{t\}\\sim\\mathrm\{Binomial\}\(k,\\,\\phi\_\{\\tau\}^\{t\}\),\(13\)whereϕτt∈\(0,1\)\\phi\_\{\\tau\}^\{t\}\\in\(0,1\)is the true success rate\. The empirical success rateϕ^τt=sτt/k\\hat\{\\phi\}\_\{\\tau\}^\{t\}=s\_\{\\tau\}^\{t\}/kis an unbiased estimator ofϕτt\\phi\_\{\\tau\}^\{t\}with variance:

Var​\(ϕ^τt\)=ϕτt​\(1−ϕτt\)k\.\\mathrm\{Var\}\(\\hat\{\\phi\}\_\{\\tau\}^\{t\}\)=\\frac\{\\phi\_\{\\tau\}^\{t\}\(1\-\\phi\_\{\\tau\}^\{t\}\)\}\{k\}\.\(14\)

#### Delta method approximation\.

Letg​\(ϕ\)=logit​\(ϕ\)=log⁡ϕ1−ϕg\(\\phi\)=\\mathrm\{logit\}\(\\phi\)=\\log\\frac\{\\phi\}\{1\-\\phi\}, so that the logit\-space latent state satisfiesψτt=g​\(ϕτt\)\\psi\_\{\\tau\}^\{t\}=g\(\\phi\_\{\\tau\}^\{t\}\)\. The derivative is:

g′​\(ϕ\)=1ϕ​\(1−ϕ\)\.g^\{\\prime\}\(\\phi\)=\\frac\{1\}\{\\phi\(1\-\\phi\)\}\.\(15\)By the central limit theorem,ϕ^τt\\hat\{\\phi\}\_\{\\tau\}^\{t\}is asymptotically Gaussian for largekk\. Applying a first\-order Taylor expansion ofggat the true valueϕτt\\phi\_\{\\tau\}^\{t\}:

g​\(ϕ^τt\)≈g​\(ϕτt\)\+g′​\(ϕτt\)​\(ϕ^τt−ϕτt\),g\(\\hat\{\\phi\}\_\{\\tau\}^\{t\}\)\\approx g\(\\phi\_\{\\tau\}^\{t\}\)\+g^\{\\prime\}\(\\phi\_\{\\tau\}^\{t\}\)\\,\(\\hat\{\\phi\}\_\{\\tau\}^\{t\}\-\\phi\_\{\\tau\}^\{t\}\),\(16\)and taking the variance of both sides:

Var​\(g​\(ϕ^τt\)\)≈\[g′​\(ϕτt\)\]2⋅Var​\(ϕ^τt\)=1\[ϕτt​\(1−ϕτt\)\]2⋅ϕτt​\(1−ϕτt\)k=1k​ϕτt​\(1−ϕτt\)\.\\mathrm\{Var\}\(g\(\\hat\{\\phi\}\_\{\\tau\}^\{t\}\)\)\\approx\[g^\{\\prime\}\(\\phi\_\{\\tau\}^\{t\}\)\]^\{2\}\\cdot\\mathrm\{Var\}\(\\hat\{\\phi\}\_\{\\tau\}^\{t\}\)=\\frac\{1\}\{\[\\phi\_\{\\tau\}^\{t\}\(1\-\\phi\_\{\\tau\}^\{t\}\)\]^\{2\}\}\\cdot\\frac\{\\phi\_\{\\tau\}^\{t\}\(1\-\\phi\_\{\\tau\}^\{t\}\)\}\{k\}=\\frac\{1\}\{k\\phi\_\{\\tau\}^\{t\}\(1\-\\phi\_\{\\tau\}^\{t\}\)\}\.\(17\)Combined with Eq\. \([16](https://arxiv.org/html/2607.27610#S3.E16)\), this establishes the approximate Gaussian observation model in logit space:

logit​\(ϕ^τt\)​∼˙​𝒩​\(ψτt,1k​ϕτt​\(1−ϕτt\)\),\\mathrm\{logit\}\(\\hat\{\\phi\}\_\{\\tau\}^\{t\}\)\\;\\dot\{\\sim\}\\;\\mathcal\{N\}\\\!\\left\(\\psi\_\{\\tau\}^\{t\},\\;\\frac\{1\}\{k\\phi\_\{\\tau\}^\{t\}\(1\-\\phi\_\{\\tau\}^\{t\}\)\}\\right\),\(18\)where∼˙\\dot\{\\sim\}denotes asymptotic distribution\. where∼˙\\dot\{\\sim\}denotes asymptotic distribution\. While this approximation improves with largerkk, the clippingϕ~τt∈\[δ,1−δ\]\\tilde\{\\phi\}\_\{\\tau\}^\{t\}\\in\[\\delta,1\-\\delta\]in the plug\-in estimator ensures the approximation operates in a regime whereϕτt\\phi\_\{\\tau\}^\{t\}is bounded away from the degenerate extremes, mitigating the sensitivity to smallkk\.

#### Plug\-in estimator\.

The true noise variance1/\[k​ϕτt​\(1−ϕτt\)\]1/\[k\\phi\_\{\\tau\}^\{t\}\(1\-\\phi\_\{\\tau\}^\{t\}\)\]depends on the unknownϕτt\\phi\_\{\\tau\}^\{t\}\. We substitute the current posterior mean as a plug\-in estimator\. To ensureRτtR\_\{\\tau\}^\{t\}remains bounded whenψ^τt\\hat\{\\psi\}\_\{\\tau\}^\{t\}approaches the degenerate regime, we clip the implied success rate away from0and11:

Rτt=1k​ϕ~τt​\(1−ϕ~τt\),R\_\{\\tau\}^\{t\}=\\frac\{1\}\{k\\,\\tilde\{\\phi\}\_\{\\tau\}^\{t\}\(1\-\\tilde\{\\phi\}\_\{\\tau\}^\{t\}\)\},\(19\)whereϕ~τt=clip​\(σ​\(ψ^τt\),δ,1−δ\)\\tilde\{\\phi\}\_\{\\tau\}^\{t\}=\\mathrm\{clip\}\(\\sigma\(\\hat\{\\psi\}\_\{\\tau\}^\{t\}\),\\,\\delta,\\,1\-\\delta\)withδ=1/\(2​k\)\\delta=1/\(2k\)\. This clipping ensuresRτtR\_\{\\tau\}^\{t\}remains bounded above, whileRτt→1/\(k​δ​\(1−δ\)\)R\_\{\\tau\}^\{t\}\\to 1/\(k\\delta\(1\-\\delta\)\)asψ^τt\\hat\{\\psi\}\_\{\\tau\}^\{t\}approaches the boundary, which automatically suppresses the Kalman gain for near\-degenerate rollouts, as formalized in Section[3\.3](https://arxiv.org/html/2607.27610#S3.SS3)\.

## DDerivation of the Gauss–Hermite Quadrature Approximation

We derive the quadrature approximation in Eq\. \([11](https://arxiv.org/html/2607.27610#S3.E11)\) from the posterior expectation in Eq\. \([10](https://arxiv.org/html/2607.27610#S3.E10)\)\.

By definition, the expectation ofhhunder the Gaussian posterior𝒩​\(ψ^τt,Pτt\)\\mathcal\{N\}\(\\hat\{\\psi\}\_\{\\tau\}^\{t\},P\_\{\\tau\}^\{t\}\)is:

A~​\(τ,t\)=∫−∞\+∞h​\(ψ\)⋅12​π​Pτt​exp⁡\(−\(ψ−ψ^τt\)22​Pτt\)​𝑑ψ\.\\tilde\{A\}\(\\tau,\\,t\)=\\int\_\{\-\\infty\}^\{\+\\infty\}h\(\\psi\)\\cdot\\frac\{1\}\{\\sqrt\{2\\pi P\_\{\\tau\}^\{t\}\}\}\\exp\\\!\\left\(\-\\frac\{\(\\psi\-\\hat\{\\psi\}\_\{\\tau\}^\{t\}\)^\{2\}\}\{2P\_\{\\tau\}^\{t\}\}\\right\)d\\psi\.\(20\)Substitutingψ=ψ^τt\+2​Pτt​x\\psi=\\hat\{\\psi\}\_\{\\tau\}^\{t\}\+\\sqrt\{2P\_\{\\tau\}^\{t\}\}\\,x, so thatd​ψ=2​Pτt​d​xd\\psi=\\sqrt\{2P\_\{\\tau\}^\{t\}\}\\,dx, the exponent simplifies to−x2\-x^\{2\}and the prefactor reduces as follows:

2​Pτt2​π​Pτt=1π\.\\frac\{\\sqrt\{2P\_\{\\tau\}^\{t\}\}\}\{\\sqrt\{2\\pi P\_\{\\tau\}^\{t\}\}\}=\\frac\{1\}\{\\sqrt\{\\pi\}\}\.\(21\)Eq\. \([20](https://arxiv.org/html/2607.27610#S4.E20)\) therefore becomes

A~​\(τ,t\)=1π​∫−∞\+∞h​\(ψ^τt\+2​Pτt​x\)​e−x2​𝑑x,\\tilde\{A\}\(\\tau,\\,t\)=\\frac\{1\}\{\\sqrt\{\\pi\}\}\\int\_\{\-\\infty\}^\{\+\\infty\}h\\\!\\left\(\\hat\{\\psi\}\_\{\\tau\}^\{t\}\+\\sqrt\{2P\_\{\\tau\}^\{t\}\}\\,x\\right\)e^\{\-x^\{2\}\}\\,dx,\(22\)which is precisely the standard Gauss–Hermite form1π​∫g​\(x\)​e−x2​𝑑x\\frac\{1\}\{\\sqrt\{\\pi\}\}\\int g\(x\)e^\{\-x^\{2\}\}dxwithg​\(x\)=h​\(ψ^τt\+2​Pτt​x\)g\(x\)=h\(\\hat\{\\psi\}\_\{\\tau\}^\{t\}\+\\sqrt\{2P\_\{\\tau\}^\{t\}\}\\,x\)\. Sincehhis smooth and bounded onℝ\\mathbb\{R\}, this integral is well\-approximated by five\-point Gauss–Hermite quadrature:

A~​\(τ,t\)≈1π​∑i=15wi​h​\(ψ^τt\+2​Pτt​xi\),\\tilde\{A\}\(\\tau,\\,t\)\\approx\\frac\{1\}\{\\sqrt\{\\pi\}\}\\sum\_\{i=1\}^\{5\}w\_\{i\}\\,h\\\!\\left\(\\hat\{\\psi\}\_\{\\tau\}^\{t\}\+\\sqrt\{2P\_\{\\tau\}^\{t\}\}\\,x\_\{i\}\\right\),\(23\)where the node–weight pairs\(xi,wi\)∈\{\(0,0\.9453\),\(±0\.9586,0\.3936\),\(±2\.0202,0\.0200\)\}\(x\_\{i\},w\_\{i\}\)\\in\\\{\(0,\\;0\.9453\),\\;\(\\pm 0\.9586,\\;0\.3936\),\\;\(\\pm 2\.0202,\\;0\.0200\)\\\}are the standard five\-point Gauss–Hermite abscissae and weights\. The approximation error decreases rapidly with the number of quadrature points; five points suffice here becausehhis well\-approximated by a low\-degree polynomial over the effective support of the Gaussian integrand\.

## EMore Experiments

![Refer to caption](https://arxiv.org/html/2607.27610v1/x16.png)

\(a\) Qwen3\-8B trained on Math\.

![Refer to caption](https://arxiv.org/html/2607.27610v1/x17.png)

\(b\) Qwen2\.5\-VL\-7B\-Instruct trained on Geometry3k\.

Figure A:Test accuracy across training steps for additional model\-task combinations under different data selection strategies\. KGPS \(ours\) consistently achieves superior performance over all baselines\.Table A:Test accuracy on mathematics benchmarks for Qwen3\-8B\. Bold indicates the best result in each column\.### E\.1Scaling to Larger Models

To further validate the scalability of KGPS, we evaluate on Qwen3\-8B under the same experimental setup as the main paper\. As shown in Figure[A](https://arxiv.org/html/2607.27610#S5.F1)\(a\) and Table[A](https://arxiv.org/html/2607.27610#S5.T1), KGPS achieves the best average accuracy of 62\.16% across six math reasoning benchmarks, outperforming all baselines including DS \(Oracle\) by 0\.87 points while using only 296k rollouts compared to DS’s 1819k, demonstrating that KGPS maintains its advantages at larger model scales\.

### E\.2Additional Results on Geometry3k

We further evaluate KGPS on the Geometry3k visual reasoning benchmark using Qwen2\.5\-VL\-7B\-Instruct\. As shown in Figure[A](https://arxiv.org/html/2607.27610#S5.F1)\(b\), KGPS achieves 55\.38% accuracy at the end of training, outperforming the best baseline DS by 1\.79 points, confirming that its advantages extend to visual reasoning tasks beyond the main paper results\.

## FDifficulty Estimation Error Analysis

To quantitatively assess the quality of prompt difficulty estimation, we compare the Mean Absolute Error \(MAE\) between predicted and empirical success rates for KGPS and MoPPS across all model\-task combinations\. As shown in Figure[B](https://arxiv.org/html/2607.27610#S6.F2), KGPS consistently achieves substantially lower MAE than MoPPS throughout training\. This gap is particularly pronounced in the early training stages, where rapid policy improvement causes prompt difficulty to shift quickly, and MoPPS’s static Beta–Bernoulli formulation fails to track these changes\. In contrast, KGPS’s dynamic state estimation adapts to policy\-induced difficulty drift via the Kalman filter, yielding well\-calibrated posteriors that remain accurate across the full training trajectory\. These results provide direct evidence that accurate difficulty estimation underlies the superior prompt selection quality of KGPS\.

![Refer to caption](https://arxiv.org/html/2607.27610v1/x18.png)

![Refer to caption](https://arxiv.org/html/2607.27610v1/x19.png)

![Refer to caption](https://arxiv.org/html/2607.27610v1/x20.png)

![Refer to caption](https://arxiv.org/html/2607.27610v1/x21.png)

![Refer to caption](https://arxiv.org/html/2607.27610v1/x22.png)

![Refer to caption](https://arxiv.org/html/2607.27610v1/x23.png)

Figure B:Mean Absolute Error \(MAE\) between KGPS\-predicted success rates and empirical success rates across training steps, compared against MoPPS\. KGPS consistently achieves substantially lower estimation error across all model\-task combinations, demonstrating the superiority of dynamic state estimation over static Beta–Bernoulli modeling\.
## GCase Study

To qualitatively illustrate the differences in reasoning quality induced by different prompt selection strategies, we present three representative cases in Table[B](https://arxiv.org/html/2607.27610#S7.T2), Table[C](https://arxiv.org/html/2607.27610#S7.T3), and Table[D](https://arxiv.org/html/2607.27610#S7.T4), where KGPS\-trained models produce correct answers while all baseline methods fail\.

In Case 1 \(Table[B](https://arxiv.org/html/2607.27610#S7.T2)\), a geometry problem requires computing the area of a nonconvex 12\-sided polygon formed by four hexagons surrounding a square\. All baseline methods \(GRPO, DS, GRESO, MoPPS\) make the same error: they naively sum the areas of the square and four complete hexagons, ignoring the overlapping regions at the corners\. In contrast, the KGPS\-trained model correctly identifies the nonconvex structure, extracts the 12 vertices, and applies the Shoelace formula to obtain the correct area16​3−2316\\sqrt\{3\}\-23\.

In Case 2 \(Table[C](https://arxiv.org/html/2607.27610#S7.T3)\), a physics problem asks for the minimum wavelength on the AM band, formatted as an integer\. All baselines use 1605 kHz as the maximum frequency, yielding 186\.9 and rounding to 187\. The KGPS\-trained model instead uses the standard AM upper bound of 1600 kHz, computes the exact value of 187\.5, and correctly rounds to 188, matching the ground truth\.

In Case 3 \(Table[D](https://arxiv.org/html/2607.27610#S7.T4)\), a simple algebra problem asks to determine three ages\. All methods solve the equations correctly, but the baselines output three separate boxed answers \(7\\boxed\{7\},14\\boxed\{14\},18\\boxed\{18\}\), causing the answer extraction to capture only the last value \(18\)\. The KGPS\-trained model recognizes this formatting requirement and outputs a single7,14,18\\boxed\{7,14,18\}, correctly matching the ground truth\.

These cases suggest that KGPS’s difficulty\-aware prompt selection, by consistently providing informative training signals, leads to improved reasoning depth, knowledge accuracy, and output formatting compared to baseline methods\.

Table B:Case study of completions generated by different prompt selection methods on a question from the AMC benchmark\. All responses are produced using the MATH\-trained Qwen3\-4B model\.✓and✗denote correct and incorrect responses, respectively\. KGPS is the only method that correctly identifies the nonconvex polygon structure and applies the Shoelace formula, while all other methods naively sum the square and hexagon areas\.Question:Four regular hexagons surround a square with side length 1, each one sharing an edge with the square\. The area of the resulting 12\-sided outer nonconvex polygon can be written asm​n\+pm\\sqrt\{n\}\+p, wheremm,nn, andppare integers andnnis not divisible by the square of any prime\. What is the absolute value ofm\+n\+pm\+n\+p? Let’s think step by step and output the final answer within box\.Ground Truth: 4GRPO Response✗:<think\>The problem involves a central square of side length 1 with four regular hexagons surrounding the square, each sharing one full edge with the square\. Since the side of the square is 1, the side of each hexagon is also 1\. Thus, the final composite structure is made of the square and four regular hexagons\. Area of the square with side length11is12=11^\{2\}=1\. The area of a regular hexagon of side lengthssis3​32​s2=3​32\\frac\{3\\sqrt\{3\}\}\{2\}s^\{2\}=\\frac\{3\\sqrt\{3\}\}\{2\}\. The area of the four hexagons is4⋅3​32=6​34\\cdot\\frac\{3\\sqrt\{3\}\}\{2\}=6\\sqrt\{3\}\. Therefore, total area of the resulting polygon is1\+6​31\+6\\sqrt\{3\}\. Thus,m=6m=6,n=3n=3,p=1p=1\. Now compute\|m\+n\+p\|=\|6\+3\+1\|=10\|m\+n\+p\|=\|6\+3\+1\|=\\boxed\{10\}</think\>Final Answer10\\boxed\{10\}DS Response✗:<think\>We are given a square of side length 1 surrounded with four regular hexagons, each sharing one edge with the square\. The resulting figure is a nonconvex 12\-sided polygon, and we are to compute its area in the formm​n\+pm\\sqrt\{n\}\+p, wherem,n,pm,n,pare integers andnnis square\-free\. The square has area11\. A regular hexagon with side lengthsshas area3​32​s2\\frac\{3\\sqrt\{3\}\}\{2\}s^\{2\}\. Since each hexagon is attached to a side of the square, the side length of each hexagon is also 1\. Therefore each hexagon has area3​32\\frac\{3\\sqrt\{3\}\}\{2\}\. There are 4 such hexagons, so the total area of the hexagons is4⋅3​32=6​34\\cdot\\frac\{3\\sqrt\{3\}\}\{2\}=6\\sqrt\{3\}\. The total area of the 12\-sided nonconvex polygon is1\+6​31\+6\\sqrt\{3\}\. This is in the formm​n\+pm\\sqrt\{n\}\+p, wherem=6m=6,n=3n=3\(which is square\-free\),p=1p=1\.\|m\+n\+p\|=\|6\+3\+1\|=10\|m\+n\+p\|=\|6\+3\+1\|=\\boxed\{10\}</think\>Final Answer10\\boxed\{10\}GRESO Response✗:<think\>The square is centered and contains four hexagons\. Each hexagon has a side matching the square’s edge of length 1\. Each regular hexagon has the areaAhex=3​32⋅s2A\_\{\\text\{hex\}\}=\\frac\{3\\sqrt\{3\}\}\{2\}\\cdot s^\{2\}\. For side lengths=1s=1:Ahex=3​32A\_\{\\text\{hex\}\}=\\frac\{3\\sqrt\{3\}\}\{2\}\. Thus, total area from the 4 hexagons:4⋅3​32=6​34\\cdot\\frac\{3\\sqrt\{3\}\}\{2\}=6\\sqrt\{3\}\. The square area:Asquare=1A\_\{\\text\{square\}\}=1\. So the total area becomesA=6​3\+1A=6\\sqrt\{3\}\+1\. We identify:m=6m=6,n=3n=3,p=1p=1\. Thus, the sum:\|m\+n\+p\|=\|6\+3\+1\|=10\|m\+n\+p\|=\|6\+3\+1\|=\\boxed\{10\}</think\>Final Answer10\\boxed\{10\}MoPPS Response✗:<think\>If there are four hexagons around square, then outer polygon has 12 sides\. But since the area is simply square \+ 4 hexagons\. So the answer maybe is1\+6​31\+6\\sqrt\{3\}\. Each hexagon has edge length matching the square’s side length \(i\.e\., 1\)\. Area of square:1×1=11\\times 1=1\. Area of a regular hexagon with side lengths=1s=1:3​32⋅s2=3​32\\frac\{3\\sqrt\{3\}\}\{2\}\\cdot s^\{2\}=\\frac\{3\\sqrt\{3\}\}\{2\}\. Combined area from all 4 hexagons:4⋅3​32=6​34\\cdot\\frac\{3\\sqrt\{3\}\}\{2\}=6\\sqrt\{3\}\. Thus, the total area of the outer polygon is1\+6​31\+6\\sqrt\{3\}\. This matchesm​n\+pm\\sqrt\{n\}\+pwithm=6m=6,n=3n=3,p=1p=1\. Therefore:\|6\+3\+1\|=10\|6\+3\+1\|=\\boxed\{10\}</think\>Final Answer10\\boxed\{10\}KGPS Response \(Ours\)✓:<think\>I need to find the area of the 12\-sided nonconvex polygon\. Since the polygon is nonconvex, I cannot simply add the areas of the square and hexagons — the hexagons overlap with each other at the corners\.I identify the 12 vertices of the outer polygon from the Asymptote code:P1=\(0,1−3\)P\_\{1\}=\(0,1\-\\sqrt\{3\}\),P2=\(1,1−3\)P\_\{2\}=\(1,1\-\\sqrt\{3\}\),P3=\(3−3,3−2\)P\_\{3\}=\(3\-\\sqrt\{3\},\\sqrt\{3\}\-2\),P4=\(3,0\)P\_\{4\}=\(\\sqrt\{3\},0\),P5=\(3,1\)P\_\{5\}=\(\\sqrt\{3\},1\),P6=\(3−3,3−3\)P\_\{6\}=\(3\-\\sqrt\{3\},3\-\\sqrt\{3\}\),P7=\(1,3\)P\_\{7\}=\(1,\\sqrt\{3\}\),P8=\(0,3\)P\_\{8\}=\(0,\\sqrt\{3\}\),P9=\(3−2,3−3\)P\_\{9\}=\(\\sqrt\{3\}\-2,3\-\\sqrt\{3\}\),P10=\(1−3,1\)P\_\{10\}=\(1\-\\sqrt\{3\},1\),P11=\(1−3,0\)P\_\{11\}=\(1\-\\sqrt\{3\},0\),P12=\(3−2,3−2\)P\_\{12\}=\(\\sqrt\{3\}\-2,\\sqrt\{3\}\-2\)\.Applying the Shoelace formula, I compute each cross\-product termxi​yi\+1−xi\+1​yix\_\{i\}y\_\{i\+1\}\-x\_\{i\+1\}y\_\{i\}for all 12 edges\. Collecting constants:−1−8−3−6−6−3−8−1−5−5=−46\-1\-8\-3\-6\-6\-3\-8\-1\-5\-5=\-46\. Collecting3\\sqrt\{3\}coefficients:1\+5\+2\+1\+4\+4\+1\+2\+5\+1\+3\+3=321\+5\+2\+1\+4\+4\+1\+2\+5\+1\+3\+3=32\.Therefore total sum is32​3−4632\\sqrt\{3\}\-46\. Area=12​\|32​3−46\|=16​3−23=\\frac\{1\}\{2\}\|32\\sqrt\{3\}\-46\|=16\\sqrt\{3\}\-23\.Som=16m=16,n=3n=3,p=−23p=\-23\. Since 3 is square\-free,\|m\+n\+p\|=\|16\+3\+\(−23\)\|=\|−4\|=4\|m\+n\+p\|=\|16\+3\+\(\-23\)\|=\|\-4\|=\\boxed\{4\}</think\>Final Answer4\\boxed\{4\}Table C:Case study on a Minerva Math problem\. All responses are produced using the MATH\-trained Qwen3\-4B model\.✓and✗denote correct and incorrect responses, respectively\. KGPS is the only method that correctly rounds 187\.5 to 188, while all other methods round down to 187\.Question:Preamble: Determine the following values from a standard radio dial\. What is the minimum wavelength in m for broadcasts on the AM band? Format your answer as an integer\.Ground Truth: 188GRPO Response✗:<think\>The AM band typically spans from 535 kHz to 1605 kHz\. Since wavelength is inversely proportional to frequency, the minimum wavelength occurs at the highest frequency in the AM range\.fmax=1605​kHz=1\.605×106​Hzf\_\{\\text\{max\}\}=1605\\text\{ kHz\}=1\.605\\times 10^\{6\}\\text\{ Hz\}\. Usingλ=c/f\\lambda=c/f:λ=3×1081\.605×106=31\.605×102≈1\.869×100=186\.9​m\\lambda=\\frac\{3\\times 10^\{8\}\}\{1\.605\\times 10^\{6\}\}=\\frac\{3\}\{1\.605\}\\times 10^\{2\}\\approx 1\.869\\times 100=186\.9\\text\{ m\}\. Since the question asks for an integer, the answer is187\\boxed\{187\}\.</think\>Final Answer187\\boxed\{187\}DS Response✗:<think\>The AM band typically spans from 535 kHz to 1605 kHz\. The minimum wavelength occurs at the maximum frequency\.fmax=1605​kHz=1\.605×106​Hzf\_\{\\text\{max\}\}=1605\\text\{ kHz\}=1\.605\\times 10^\{6\}\\text\{ Hz\}\.λ=3×1081\.605×106\\lambda=\\frac\{3\\times 10^\{8\}\}\{1\.605\\times 10^\{6\}\}\.300,000/1,605≈186\.9300\{,\}000/1\{,\}605\\approx 186\.9\. So the minimum wavelength is about 187 meters\. Therefore, the answer is187\\boxed\{187\}\.</think\>Final Answer187\\boxed\{187\}GRESO Response✗:<think\>The AM band is generally defined by frequencies ranging between 535 kHz to 1605 kHz\. The maximum frequency of 1605 kHz leads to the minimum wavelength\.λ=3×1081605×103=31\.605×100≈1\.869×100=186\.9​meters\\lambda=\\frac\{3\\times 10^\{8\}\}\{1605\\times 10^\{3\}\}=\\frac\{3\}\{1\.605\}\\times 100\\approx 1\.869\\times 100=186\.9\\text\{ meters\}\. Rounded to the nearest integer, the answer is187\\boxed\{187\}\.</think\>Final Answer187\\boxed\{187\}MoPPS Response✗:<think\>The standard AM band spans from 520 kHz to 1700 kHz\. The maximum frequency is 1605 kHz, or1\.605×1061\.605\\times 10^\{6\}Hz\.λ=3×1081\.605×106=31\.605×100=1\.869×100=186\.9\\lambda=\\frac\{3\\times 10^\{8\}\}\{1\.605\\times 10^\{6\}\}=\\frac\{3\}\{1\.605\}\\times 100=1\.869\\times 100=186\.9meters\. Rounded to the nearest integer:187\\boxed\{187\}\.</think\>Final Answer187\\boxed\{187\}KGPS Response \(Ours\)✓:<think\>The standard AM band is from 540 kHz to 1600 kHz\. The minimum wavelength corresponds to the maximum frequency:fmax=1600​kHz=1\.6×106​Hzf\_\{\\text\{max\}\}=1600\\text\{ kHz\}=1\.6\\times 10^\{6\}\\text\{ Hz\}\. Usingλ=c/f\\lambda=c/f:λ=3×1081\.6×106=300,000,0001,600,000=187\.5​m\\lambda=\\frac\{3\\times 10^\{8\}\}\{1\.6\\times 10^\{6\}\}=\\frac\{300\{,\}000\{,\}000\}\{1\{,\}600\{,\}000\}=187\.5\\text\{ m\}\. The problem specifies to format the answer as an integer\. 187\.5 rounded to the nearest whole number is 188\.But maybe the AM band extends to 1605 kHz? Let me recalculate:3×1081\.605×106≈186\.9\\frac\{3\\times 10^\{8\}\}\{1\.605\\times 10^\{6\}\}\\approx 186\.9, so approximately 187 meters\. However, if the AM band maximum is 1600 kHz, the exact calculation gives 187\.5, which rounds to 188\. Since the problem says “standard radio dial,” I use 1600 kHz as the maximum frequency\. Therefore, the answer is188\\boxed\{188\}\.</think\>Final Answer188\\boxed\{188\}Table D:Case study on an OlympiadBench problem\. All responses are produced using the MATH\-trained Qwen3\-4B model\.✓and✗denote correct and incorrect responses, respectively\. All methods solve the algebra correctly, but only KGPS formats the answer as a single boxed tuple7,14,18\\boxed\{7,14,18\}matching the ground truth, while other methods output separate boxes, causing the extracted answer to capture only the last value\.Question:Qing is twice as old as Rayna\. Qing is 4 years younger than Paolo\. The average age of Paolo, Qing and Rayna is 13\. Determine their ages\.Ground Truth: 7, 14, 18GRPO Response✗:<think\>LetRRbe Rayna’s age\. Then Qing’s age is2​R2R, and Paolo’s age is2​R\+42R\+4\. The average age is 13, so\(R\+2​R\+2​R\+4\)/3=13\(R\+2R\+2R\+4\)/3=13, which gives5​R\+4=395R\+4=39, soR=7R=7\. Therefore Rayna is 7, Qing is 14, Paolo is 18\.</think\>Final Answer:Rayna is7\\boxed\{7\}, Qing is14\\boxed\{14\}, and Paolo is18\\boxed\{18\}\.DS Response✗:<think\>Let Rayna’s age beRR\. Then Qing’s age is2​R2Rand Paolo’s age is2​R\+42R\+4\. Using the average:\(R\+2​R\+\(2​R\+4\)\)/3=13\(R\+2R\+\(2R\+4\)\)/3=13, so5​R\+4=395R\+4=39,R=7R=7\. Rayna: 7, Qing: 14, Paolo: 18\. Verification:\(7\+14\+18\)/3=39/3=13\(7\+14\+18\)/3=39/3=13\. Correct\.</think\>Final Answer:7\\boxed\{7\},14\\boxed\{14\},18\\boxed\{18\}GRESO Response✗:<think\>LetRRbe Rayna’s age\. FromQ=2​RQ=2RandP=Q\+4=2​R\+4P=Q\+4=2R\+4:P\+Q\+R=\(2​R\+4\)\+2​R\+R=5​R\+4=39P\+Q\+R=\(2R\+4\)\+2R\+R=5R\+4=39, soR=7R=7\. ThenQ=14Q=14,P=18P=18\.</think\>Final Answer:Rayna is7\\boxed\{7\}, Qing is14\\boxed\{14\}, and Paolo is18\\boxed\{18\}\.MoPPS Response✗:<think\>Define Rayna’s age asRR, Qing’s age becomes2​R2R, and Paolo’s age becomes2​R\+42R\+4\. Summing gives:\(R\)\+\(2​R\)\+\(2​R\+4\)=5​R\+4=39\(R\)\+\(2R\)\+\(2R\+4\)=5R\+4=39, soR=7R=7\. Rayna = 7, Qing = 14, Paolo = 18\.</think\>Final Answer:The ages of Rayna, Qing, and Paolo are7\\boxed\{7\},14\\boxed\{14\}, and18\\boxed\{18\}respectively\.KGPS Response \(Ours\)✓:<think\>LetRRbe Rayna’s age,QQbe Qing’s age,PPbe Paolo’s age\. From the relationships:Q=2​RQ=2RandP=Q\+4=2​R\+4P=Q\+4=2R\+4\. Using the average age condition:\(Q\+R\+P\)/3=13\(Q\+R\+P\)/3=13, soQ\+R\+P=39Q\+R\+P=39\. Substituting:2​R\+R\+\(2​R\+4\)=392R\+R\+\(2R\+4\)=39,5​R\+4=395R\+4=39,5​R=355R=35,R=7R=7\. Therefore: Rayna=7=7, Qing=14=14, Paolo=18=18\. Verification:7\+14\+18=397\+14\+18=39, average=13=13\. Correct\.The problem asks to “determine their ages” and the instruction says to return the final answer in\\\\backslashboxed\{\} tags\. Since the answer consists of three numbers, I should put them in a single box separated by commas\.</think\>Final Answer7,14,18\\boxed\{7,14,18\}
## HLimitations

KGPS models each prompt’s difficulty independently via a per\-prompt Kalman filter, without capturing correlations across prompts\. In practice, structurally similar prompts are likely to exhibit correlated difficulty trajectories as the policy evolves\. Exploiting such inter\-prompt structure, for example through shared latent factors or prompt embeddings, could accelerate posterior convergence and improve estimation accuracy\. We leave this extensions to future work\.

Similar Articles

Environment-Grounded Automated Prompt Optimization for LLM Game Agents

arXiv cs.CL

Introduces an automated prompt optimization framework for LLM game agents that decomposes the observation-to-action pipeline into two agents and iteratively refines prompts via an evolutionary loop guided by environment returns. Evaluated on BabyAI tasks, it significantly improves success rates (e.g., from 0% to 72.5% on PutNext) without updating model weights.