Tag
A preprint introduces FLEET, an algorithm that makes Best-of-N sampling reward-aware by attributing external rewards to tokens, storing hidden states with reward metadata in a vector store, and using modified MCTS to reweight logits in subsequent iterations. Results show equivalent or better performance with far fewer sampling iterations on GSM8K and LiveCodeBench with Llama 3.2 3B, plus a reusable metadata store that can serve as a prior for other tasks or enrich SFT/RL.
The paper presents Spectral Feedback, an algorithm that enhances test-time alignment for discrete diffusion models in protein inverse folding by iteratively selecting edit-positions using sparse Fourier representations, resulting in improved performance for reward maximization.