token-level-rejection-sampling

Tag

Cards List
#token-level-rejection-sampling

@xennygrimmato_: if you’re wondering how token-level rejection sampling works in this paper, here’s how they do it: M_t = max_v [ pi_the…

X AI KOLs Timeline · 2026-07-11 Cached

Explains token-level rejection sampling for RLHF/PPO, where importance ratio M_t is the maximum over vocabulary and tokens are accepted with Bernoulli sampling based on w_t / M_t.

0 favorites 0 likes
← Back to home

Submit Feedback