Spectral Feedback for Test-Time Alignment of Protein Diffusion Models

arXiv cs.AI Papers

Summary

The paper presents Spectral Feedback, an algorithm that enhances test-time alignment for discrete diffusion models in protein inverse folding by iteratively selecting edit-positions using sparse Fourier representations, resulting in improved performance for reward maximization.

arXiv:2609.30456v1 Announce Type: new Abstract: Reward maximization alignment methods for discrete diffusion models have primarily focused on steering the reverse process, either by influencing token logits or by selecting favorable sequences at intermediate steps. These approaches largely treat inference as a unidirectional process, lacking mechanisms for revisiting undesirable token selections. We introduce Spectral Feedback, an algorithm that selects edit-positions in a feedback loop, allowing the model to iteratively correct its own generations. This approach leverages the mask structure of discrete diffusion models by re-masking and re-sampling tokens, analogous to image editing methods that reintroduce noisy latents and re-run the reverse process. While prior alignment methods focus on what token labels to assign to maximize a target reward, we instead treat which tokens to revisit as the central alignment problem. Selecting edit-positions is challenging because edit effects are interdependent: the impact of modifying one token depends on which others are edited simultaneously. We define an edit-set as a set of token positions to re-mask and re-sample. Motivated by prior work on sparse interactions in biological systems, we find empirically that edit-set value functions for protein inverse folding admit sparse Fourier representations. This structure enables Spectral Feedback to efficiently learn and optimize the value functions for edit-position selection. Spectral Feedback is model-agnostic and can be applied to pretrained, test-time aligned, and fine-tuned diffusion models. For all of these models, the algorithm improves alignment performance without modifying the underlying generative process. Applied to inverse folding with a protein stability reward oracle, it achieves a 32.3% increase in stable proteins for a pretrained model, 24.8% for Best-of-10, and 5.8% for a state-of-the-art RL fine-tuned diffusion model.
Original Article
View Cached Full Text

Cached at: 09/28/26, 09:36 AM

# Spectral Feedback for Test-Time Alignment of Protein Diffusion Models
Source: [https://arxiv.org/html/2609.30456](https://arxiv.org/html/2609.30456)
Mert CemriLandon ButlerKannan RamchandranAffiliation:Department of Electrical Engineering and Computer SciencesAffiliation:University of California, Berkeley

###### Abstract

Reward maximization alignment methods for discrete diffusion models have primarily focused on steering the diffusion reverse process, either by influencing token logits or by selecting favorable sequences at intermediate steps to eventually yield high\-reward samples\. These approaches largely treat inference as a unidirectional process, lacking effective mechanisms for revisiting undesirable token selections\. We introduceSpectral Feedback, an algorithm that selects edit\-positions in a feedback loop, allowing the model to iteratively correct its own generations\. This approach leverages the masking structure of discrete diffusion models by re\-masking and re\-sampling tokens, analogous to image editing methods that reintroduce noisy latents and re\-run the reverse diffusion process\. While prior alignment methods focus onwhat token labelsto assign to maximize a target reward, we instead treatwhich tokensto revisit as the central alignment problem\. Selecting edit\-positions is challenging because edit effects are interdependent: the impact of modifying one token depends on which others are edited simultaneously\. We define an edit\-set as a set of positions to edit by re\-masking and re\-sampling the corresponding tokens in a sequence\. Motivated by prior work on sparse interactions in biological systems, we find empirically that edit\-set value functions for protein inverse folding admit sparse Fourier representations\. This structure enables Spectral Feedback to efficiently learn and optimize the value functions for edit\-position selection\. Spectral Feedback is model\-agnostic and can be applied to pretrained, test\-time aligned, and fine\-tuned diffusion models\. For all of these models, the algorithm improves alignment performance without modifying the underlying generative process\. Applied to inverse folding with a protein stability reward oracle, it achieves a 32\.3% increase in stable proteins for a pretrained model, 24\.8% for Best\-of\-10, and 5\.8% for a state\-of\-the\-art RL fine\-tuned diffusion model\.

## 1Introduction

Inverse protein folding is the task of designing amino acid sequences that fold into a target protein backbone\. Discrete diffusion models have emerged as a leading architecture for this task\([Yi et al\., 2023](https://arxiv.org/html/2609.30456#bib.bib39);[Gruver et al\., 2023](https://arxiv.org/html/2609.30456#bib.bib5);[Cemri et al\., 2024](https://arxiv.org/html/2609.30456#bib.bib2);[Wang et al\., 2025](https://arxiv.org/html/2609.30456#bib.bib1)\)\. While these models can successfully generate sequences that fold into target backbones, scientists may also target additional attributes such as stability\([Wang et al\., 2025](https://arxiv.org/html/2609.30456#bib.bib1)\)orβ\\beta\-sheets\([Gruver et al\., 2023](https://arxiv.org/html/2609.30456#bib.bib5)\)\. Many sequences can fold into similar structures, yet only a few may satisfy target criteria\([Xiong et al\., 2026](https://arxiv.org/html/2609.30456#bib.bib20)\)\. The alignment problem we study is that given an arbitrary protein model, we want to generate proteins with desirable attributes by efficiently using a reward model that quantifies those attributes\. Test\-time alignment is a category of methods that use additional compute during inference to generate desirable samples\. Existing test\-time alignment of discrete diffusion models can be organized into three groups: token logit alignment, tree\-search, and re\-noising feedback\. Token logit alignment skews probabilities of tokens during the reverse process such as in\([Xiong et al\., 2026](https://arxiv.org/html/2609.30456#bib.bib20)\)and\([Gruver et al\., 2023](https://arxiv.org/html/2609.30456#bib.bib5)\)\. Tree\-search techniques, like those discussed in\([Huang et al\., 2025](https://arxiv.org/html/2609.30456#bib.bib9);[Darmawan et al\., 2025](https://arxiv.org/html/2609.30456#bib.bib8);[Uehara et al\., 2025](https://arxiv.org/html/2609.30456#bib.bib10)\), explore different denoising trajectories in the reverse process and select the highest\-reward sample\. We focus on the third and less studied group of methods: re\-noising feedback\. This method, discussed in[Wang et al\. \(2026\)](https://arxiv.org/html/2609.30456#bib.bib38), naturally exploits the masking structure of discrete diffusion models in the following sequence: generate a candidate sequence, re\-mask a subsetS⊆\[L\]S\\subseteq\[L\]of positions to edit, and re\-sample those positions\. This type of method uniquely treats the model as a black box instead of modifying its denoising trajectory\. The addition and removal of noise have parallels with image editing methods that transform images into partially noisy latents and then re\-sample a segment of a diffusion reverse process\. For example,[Hertz et al\. \(2022\)](https://arxiv.org/html/2609.30456#bib.bib35)and[Mokady et al\. \(2022\)](https://arxiv.org/html/2609.30456#bib.bib36)both use diffusion inversion techniques to recover partially noisy latents and then generate an edited sample using the reverse process conditioned on the edit instructions\. Another work,[Meng et al\. \(2022\)](https://arxiv.org/html/2609.30456#bib.bib37), directly adds Gaussian noise and then executes the reverse process to generate an edited image\. The re\-noising used in these methods is closely related to the Gaussian noise used in the forward processes of their respective continuous diffusion models\. Previous sequence editing methods via re\-masking are also closely related to the mask noise design of their respective discrete diffusion models since they often impose random sampling with independence assumptions\([Lee et al\., 2025](https://arxiv.org/html/2609.30456#bib.bib4);[Reid et al\., 2022](https://arxiv.org/html/2609.30456#bib.bib12);[Wang et al\., 2026](https://arxiv.org/html/2609.30456#bib.bib38);[Gruver et al\., 2023](https://arxiv.org/html/2609.30456#bib.bib5)\)\. In these methods, the actual alignment comes from other techniques such as importance sampling or classifier\-free guidance rather than solely re\-masking\.

We claim that with the removal of independence assumptions, re\-masking can be an effective alignment method entirely on its own\. This approach is useful because its modularity allows any alignment method to be improved within a feedback loop\. Formally, we focus our work on the design of the edit\-setSSto maximize a value functionffthat relates edits with alignment rewards\. However, for an arbitrary value functionf:2\[L\]→ℝf:2^\{\[L\]\}\\to\\mathbb\{R\}, maximization over2L2^\{L\}candidates is intractable, which supports the use of independence assumptions in past works\. Our work identifies structure in edit\-set value functions to enable approximations that avoid combinatorial challenges during optimization\.

Figure 1:Spectral Feedback loop using targeted edits to iteratively improve attributes of a protein sequence with discrete diffusion\.Figure 2:Spectral Feedback in different protein model inference settings, aligning for stability \(Δ​Δ​G\\Delta\\Delta G\- kcal/mol\)\. These are reward evaluation distributions from experiments in section[5](https://arxiv.org/html/2609.30456#S5)\.The expansion offfin the Fourier basis assigns every subsetT⊆\[L\]T\\subseteq\[L\]a coefficientF⁡\(T\)F\(T\)that measures the strength of the interaction among positions inTT\([Li et al\., 2014](https://arxiv.org/html/2609.30456#bib.bib25);[Butler et al\., 2025](https://arxiv.org/html/2609.30456#bib.bib11)\)\. Our central technical claim, which we verify empirically, is that this expansion is sparse for protein sequence edit\-set value functions: a small number of coefficients capture most of the variation inff\. Sparsity of this form is consistent with the epistatic structure of biological sequences\([Poelwijk et al\., 2019](https://arxiv.org/html/2609.30456#bib.bib26);[Aghazadeh et al\., 2021](https://arxiv.org/html/2609.30456#bib.bib23);[Aghazadeh et al\., 2020](https://arxiv.org/html/2609.30456#bib.bib27)\), in which the joint effect of mutating multiple amino acids is dominated by a small number of interacting groups rather than spread uniformly across all2L2^\{L\}subsets\. Because the interactions that determine the physical properties of a protein are sparse, we expect the interactions between edits of these very same amino acids will similarly reflect this structure\.

Importantly, Fourier sparsity makes it tractable to approximatefffrom a limited number of queries\. Instead of requiring an exhaustive2L2^\{L\}reward evaluations, the number of samples needed for sparse Fourier recovery grows only with the sparsity level\([Scheibler et al\., 2015](https://arxiv.org/html/2609.30456#bib.bib33);[Li et al\., 2014](https://arxiv.org/html/2609.30456#bib.bib25)\), which we find to be small in practice\([Kang et al\., 2025](https://arxiv.org/html/2609.30456#bib.bib29);[Butler et al\., 2025](https://arxiv.org/html/2609.30456#bib.bib11)\)\. In proteins where the edit value function spectrum is mostly first\-order, individual amino acid positions largely determine whether re\-masking is useful; in proteins with substantial higher\-order terms, the value of editing an amino acid depends strongly on which other amino acids are edited with it\. While these high\-order terms lead to increased computational complexity, we find that most spectral energy is concentrated in lower\-order terms\.

Motivated by our observations of sparsity in edit\-set value functions, we introduceSpectral Feedback, a feedback\-based test\-time alignment method for discrete diffusion in which sparse Fourier recovery solves the combinatorial edit\-selection problem\. At each iteration, Spectral Feedback queries the value function on a relatively small number of edit\-sets, fits a sparse Fourier approximationf^\\hat\{f\}to the resulting queries, recovers the highest\-valued setS∗=arg⁡maxS​f^​\(S\)S^\{\*\}=\\arg\\max\_\{S\}\\hat\{f\}\(S\), and re\-invokes the reverse process with the positions inS∗S^\{\*\}re\-masked\. Spectral Feedback treats the underlying diffusion model as a black box: it requires only the ability to sample from the model conditioned on a partially masked input, and never modifies its weights or its denoising trajectory\. This makes it compatible with any pretrained discrete diffusion model and composable with other alignment techniques\. We apply Spectral Feedback on top of a pretrained model, an RL fine\-tuned model \(DRAKES\([Wang et al\., 2025](https://arxiv.org/html/2609.30456#bib.bib1)\)\), and on Best\-of\-NNsampling\.

We evaluate Spectral Feedback on protein backbones from the Megascale dataset\([Tsuboyama et al\., 2023](https://arxiv.org/html/2609.30456#bib.bib17)\), aligning sequences to a stability oracle that predictsΔ​Δ​G\\Delta\\Delta G, the difference in Gibbs free energy between a designed sequence and its wild\-type variant\. To detect reward over\-optimization, we separately track self\-consistency RMSD \(s​c​R​M​S​DscRMSD\) between the target backbone and the ESMFold\-predicted structure of each design\([Lin et al\., 2023](https://arxiv.org/html/2609.30456#bib.bib19)\)as well as the sequence naturalness via ProtGPT2 log\-likelihoods\([Ferruz et al\., 2022](https://arxiv.org/html/2609.30456#bib.bib18)\)\. In isolation, Spectral Feedback reaches the reward of Best\-of\-1010in two feedback iterations and Best\-of\-5050in five\. As shown in the evaluationΔ​Δ​G\\Delta\\Delta Greward distribution in Figure[2](https://arxiv.org/html/2609.30456#S1.F2), when used with a pretrained model, Best\-of\-NNsampling on a pretrained model, or an RL fine\-tuned model, Spectral Feedback improves alignment reward for all configurations\.

#### Contributions\.

1. 1\.Formulation\.We formalize a feedback loop that relates the selection of edit\-positions at each iteration to a target reward metric by defining an edit\-set value function\. The test\-time alignment problem is then to maximize this intermediate value function rather than the target metric directly\. This creates a combinatorial challenge that prior refinement methods largely avoid, since the value of editing one position can depend on which other positions are edited simultaneously\.
2. 2\.Sparsity analysis\.We characterize the Fourier spectra and the sparsity structure of edit\-set value functions across protein targets\. Our analysis suggests that some proteins are dominated by first\-order effects, while others involve higher\-order interactions, giving an empirical explanation for performance differences between simpler position\-wise edit methods and high\-order sparse recovery methods\.
3. 3\.Method\.We introduce Spectral Feedback, which learns a sparse Fourier approximation of an edit\-set value function, and then uses the approximation to efficiently identify promising edits at each iteration of a diffusion feedback loop\. Spectral Feedback is model\-agnostic, which allows it to improve the performance of any pretrained, fine\-tuned, or test\-time aligned model\.
4. 4\.Empirical results\.On inverse folding of protein backbones in the Megascale dataset, Spectral Feedback improves alignment rewards for several different underlying protein models\. Compared against several greedy or gradient\-based baseline edit\-set selection methods, Spectral Feedback scales the best in both compute and latency\.

## 2Preliminaries

Discrete Diffusion Models\.We establish the notation here, while a complete treatment appears in Appendix[B](https://arxiv.org/html/2609.30456#A2)\. Consider a vocabularyVV, sequence lengthLL, and sample space𝒳:=VL\\mathcal\{X\}:=V^\{L\}\. A discrete diffusion model defines a forward process that progressively masks tokens via a Continuous Time Markov Chain \(CTMC\) with rate matricesQtQ\_\{t\}, evolving a probability mass trajectory according tod​pt/d​t=Qt​ptdp\_\{t\}/dt=Q\_\{t\}\\,p\_\{t\}withp0∼pdatap\_\{0\}\\sim p\_\{\\text\{data\}\}\([Liang et al\., 2025](https://arxiv.org/html/2609.30456#bib.bib14)\)\. The reverse process, initialized from a fully masked sequencexTx\_\{T\}, progressively unmasks tokens by learning a score functionpt​\(y\)/pt​\(x\)p\_\{t\}\(y\)/p\_\{t\}\(x\)from training data\. A Hamming distance constraint \(Qt​\(x,y\)=0Q\_\{t\}\(x,y\)=0ford⁡\(x,y\)\>1d\(x,y\)\>1\) restricts each transition to a single\-token flip; multiple tokens are updated per step by assuming independent transitions\.

Feedback with Edit\-Positions\.Given a protein sequencexxof lengthLL, feedback selects a subset of positions to re\-mask and then re\-samples them from the diffusion model\. LetS⊆\[L\]S\\subseteq\[L\]denote a set of edit\-positions\. ApplyingSSconstructs a partially masked sequencex~\\tilde\{x\}:

x~i=\{M​A​S​Kif​i∈Sxiif​i∉S\\displaystyle\\tilde\{x\}\_\{i\}=\\left\\\{\\begin\{aligned\} &MASK\\;&\\text\{if \}i\\in S\\\\ &x\_\{i\}\\;&\\text\{if \}i\\notin S\\end\{aligned\}\\right\.\(1\)A new sequencez∼ppre\(⋅\|x~\)z\\sim p\_\{\\text\{pre\}\}\(\\cdot\|\\tilde\{x\}\)is generated by running the reverse process initialized withx~\\tilde\{x\}\. The refinement scheme alternates between selectingSSand re\-sampling, seeking edit\-positions that maximize a value function:S∗=arg⁡maxS⊆\[L\],\|S\|≤k​f​\(S\)S^\{\*\}=\\underset\{S\\subseteq\[L\],\|S\|\\leq k\}\{\\arg\\max\}\\;f\(S\)\.

When samplingzz, we initialize the reverse process at the first time\-step withx~\\tilde\{x\}and run it to completion\. In future work, it could be natural and more efficient to use a later time\-step that relates to the masking pattern ofx~\\tilde\{x\}and noise schedule of the diffusion model\.

We define the total number of stepsNNof a feedback alignment process as the total number of proteins generated during the process\. Our algorithm generates an initial protein sequence before executing the feedback loop, so it hasN−1N\-1feedback iterations\.

The cardinality boundkkbalances reward improvement against the cost of longer reverse processes and the combinatorial complexity of searching over all candidate sets\. Prior methods select edit\-positions randomly or independently\([Lee et al\., 2025](https://arxiv.org/html/2609.30456#bib.bib4);[Reid et al\., 2022](https://arxiv.org/html/2609.30456#bib.bib12);[Gruver et al\., 2023](https://arxiv.org/html/2609.30456#bib.bib5)\)\. Our method instead searches for optimal*sets*of positions by exploiting their joint reward structure, as described in Section[4](https://arxiv.org/html/2609.30456#S4)\.

## 3Related Work

Protein Sequence Design and Inverse Folding\.Recent work has produced strong generative models for inverse folding\. ProteinMPNN\([Dauparas et al\., 2022](https://arxiv.org/html/2609.30456#bib.bib3)\)introduced a graph neural network conditioned on backbone geometry that auto\-regressively generates sequences\. Discrete diffusion variants extend the structured methods to non\-auto\-regressive masked generation\([Yi et al\., 2023](https://arxiv.org/html/2609.30456#bib.bib39);[Gruver et al\., 2023](https://arxiv.org/html/2609.30456#bib.bib5);[Cemri et al\., 2024](https://arxiv.org/html/2609.30456#bib.bib2);[Wang et al\., 2025](https://arxiv.org/html/2609.30456#bib.bib1)\)\. DRAKES\([Wang et al\., 2025](https://arxiv.org/html/2609.30456#bib.bib1)\)fine\-tunes a discrete diffusion inverse folding model with a differentiable reward signal, improving alignment to a stability oracle at the cost of model\-specific training\. Spectral Feedback is complementary to these methods: it operates at test time, leaves model weights untouched, and composes on top of both pretrained and DRAKES\-style fine\-tuned models\.

Inference\-Time Alignment for Discrete Generative Models\.Inference\-time alignment methods spend additional compute at sampling time to improve reward without retraining\. Best\-of\-NNdrawsNNindependent samples and returns the highest\-reward one\([Huang et al\., 2025](https://arxiv.org/html/2609.30456#bib.bib9)\); it is the dominant approach in protein design due to its simplicity\([Yeh et al\., 2023](https://arxiv.org/html/2609.30456#bib.bib21);[Bennett et al\., 2026](https://arxiv.org/html/2609.30456#bib.bib22);[Darmawan et al\., 2025](https://arxiv.org/html/2609.30456#bib.bib8)\), but its reward improvement grows slowly withNN\. Beam Search retains only the highest\-reward partial trajectories at each reverse\-process step\([Uehara et al\., 2025](https://arxiv.org/html/2609.30456#bib.bib10)\)\. Token\-level guidance methods modify the score function during the reverse process: ProteinGuide uses Bayesian conditioning with classifiers trained on partially masked sequences\([Xiong et al\., 2026](https://arxiv.org/html/2609.30456#bib.bib20)\), and NOS performs Langevin updates in the protein embedding space\([Gruver et al\., 2023](https://arxiv.org/html/2609.30456#bib.bib5)\)\. Spectral Feedback differs from all of these in the object it optimizes\. Rather than choosing among full samples or modifying token logits, it chooses which positions of an already\-generated sample should be reopened for resampling\.

Sequence Editing\.Sequence editing techniques often progressively transform an existing sample toward a target objective\.[Lee et al\. \(2025\)](https://arxiv.org/html/2609.30456#bib.bib4)use Multiple\-Try Metropolis with uniform random noising to refine intermediate reverse\-process time\-steps\.[Gruver et al\. \(2023\)](https://arxiv.org/html/2609.30456#bib.bib5)apply embedding\-space Langevin updates at positions sampled with probabilities proportional to reward\-gradient magnitudes, and DiffusER uses random Levenshtein edits as the forward process\([Reid et al\., 2022](https://arxiv.org/html/2609.30456#bib.bib12)\)\. In the language model setting, Self\-Refine\([Madaan et al\., 2023](https://arxiv.org/html/2609.30456#bib.bib6)\)and multi\-agent pipeline optimization\([Xue et al\., 2025](https://arxiv.org/html/2609.30456#bib.bib7)\)use LLM\-generated feedback rather than reward oracles\. A common pattern across these methods is that edit\-positions are selected independently\. Spectral Feedback instead treats edit\-set selection as a combinatorial reward maximization problem and uses sparse Fourier recovery to optimize over interacting position sets directly\.

Sparse Boolean Function Recovery and Epistasis\.The combinatorial structure of edit\-set selection connects to a long line of work on sparse Boolean function recovery\. Epistasis, in which genes or amino acids interact to determine biological traits\([Cordell, 2002](https://arxiv.org/html/2609.30456#bib.bib28)\), is documented to exhibit sparse higher\-order structure\([Poelwijk et al\., 2019](https://arxiv.org/html/2609.30456#bib.bib26);[Aghazadeh et al\., 2021](https://arxiv.org/html/2609.30456#bib.bib23);[Aghazadeh et al\., 2020](https://arxiv.org/html/2609.30456#bib.bib27)\), and similar sparsity patterns have been observed in language and vision data\([Butler et al\., 2025](https://arxiv.org/html/2609.30456#bib.bib11)\)\. LASSO recovers sparse linear models efficiently\([Tibshirani, 1996](https://arxiv.org/html/2609.30456#bib.bib24)\)but requires explicit enumeration of interaction features, which scales exponentially with order\. Recent algorithms\([Scheibler et al\., 2015](https://arxiv.org/html/2609.30456#bib.bib33);[Li et al\., 2014](https://arxiv.org/html/2609.30456#bib.bib25);[Amrollahi et al\., 2019](https://arxiv.org/html/2609.30456#bib.bib32);[Kang et al\., 2025](https://arxiv.org/html/2609.30456#bib.bib29);[Butler et al\., 2025](https://arxiv.org/html/2609.30456#bib.bib11)\)address this by recovering higher\-order Fourier coefficients with sub\-exponential sample complexity\. We use this line of work to recover the edit\-set reward function and to study, for the first time, how the sparsity structure of this function varies across protein targets\.

## 4Spectral Feedback

We introduce Spectral Feedback, a feedback loop that iteratively selects and applies edits to a sequence with a discrete diffusion model\. Our algorithm leverages sparse interactions in an edit\-set value function to efficiently select promising edits\. In this section, we study these sparse interactions and then describe the details of our algorithm\.

### 4\.1Spectral Function Approximation

Figure 3:Spectral analysis of the value functionfa​v​gf\_\{avg\}with the pretrained diffusion model\. The bar graph shows the proportion that each order of Fourier coefficients contributes to the total variance of the coefficients\. The plots compare the 2KVV and R6\-560 \(r6\_560\_TrROS\_Hall\) backbones from the Megascale dataset\. Experimental details and additional plots are in Appendix[D\.2](https://arxiv.org/html/2609.30456#A4.SS2)\.The goal of edit\-set selection is, given a protein sequencexx, target backboneyy, and protein modelp\(⋅\|S,x,y\)p\(\\cdot\|S,x,y\), to find a setS∗S^\{\*\}that maximizes a value functionf:2\[L\]→ℝf:2^\{\[L\]\}\\rightarrow\\mathbb\{R\}\. In terms of edits, this value function should reward more promising edit selections \(i\.e\., those with higher expected rewards after re\-sampling\)\. Given a reward oracler⁡\(⋅\)r\(\\cdot\), we use the following value functions depending on the sampling techniques used for the underlying protein model:

fa​v​g​\(S,xt,y\)=1n​∑i=1nr⁡\(zi\),fm​a​x​\(S,xt,y\)=max⁡\{r⁡\(z1\),…,r⁡\(zn\)\}\\displaystyle f\_\{avg\}\(S,x\_\{t\},y\)=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}r\(z\_\{i\}\),\\quad f\_\{max\}\(S,x\_\{t\},y\)=\\max\\\{r\(z\_\{1\}\),\\ldots,r\(z\_\{n\}\)\\\}\(2\)wherezi∼p⁡\(xt−1\|S,xt,y​andxt−1is unmasked\)z\_\{i\}\\sim p\(x\_\{t\-1\}\|S,x\_\{t\},y\\text\{ and $x\_\{t\-1\}$ is unmasked\}\)for a discrete diffusion modelpp\. Other value functions exist—such as trained predictors for partially masked sequences\([Gruver et al\., 2023](https://arxiv.org/html/2609.30456#bib.bib5);[Xiong et al\., 2026](https://arxiv.org/html/2609.30456#bib.bib20)\)or posterior mean approximations\([Uehara et al\., 2025](https://arxiv.org/html/2609.30456#bib.bib10)\)—but we find the single\-step approximation empirically effective \.

A value functionffadmits a Fourier transformF:2\[L\]→ℝF:2^\{\[L\]\}\\rightarrow\\mathbb\{R\}as follows\([O’Donnell, 2014](https://arxiv.org/html/2609.30456#bib.bib31)\):

Transform:F\(T\)=12L∑S⊆\[L\]\(−1\)\|S∩T\|f\(S\),Inverse:f\(S\)=∑T⊆\[L\]\(−1\)\|T∩S\|F\(T\)\.\\displaystyle\\text\{Transform: \}F\(T\)=\\frac\{1\}\{2^\{L\}\}\\sum\_\{S\\subseteq\[L\]\}\(\-1\)^\{\|S\\cap T\|\}f\(S\),\\quad\\text\{Inverse: \}f\(S\)=\\sum\_\{T\\subseteq\[L\]\}\(\-1\)^\{\|T\\cap S\|\}F\(T\)\.\(3\)
While there are2L2^\{L\}coefficients that fully describeff, if we can use an approximationf^≈f\\hat\{f\}\\approx fwith significantly fewer coefficients, then optimizingf^\\hat\{f\}can become tractable\. Since the Fourier transform is orthonormal, Parseval’s theorem gives us

∑S⊆\[L\]\(f⁡\(S\)−f^​\(S\)\)2=∑T⊆\[L\]\(F⁡\(T\)−F^​\(T\)\)2\.\\displaystyle\\sum\_\{S\\subseteq\[L\]\}\\left\(f\(S\)\-\\hat\{f\}\(S\)\\right\)^\{2\}=\\sum\_\{T\\subseteq\[L\]\}\\left\(F\(T\)\-\\hat\{F\}\(T\)\\right\)^\{2\}\.\(4\)Consequently, ifffadmits a*sparse*Fourier representation—i\.e\., most of its energy is concentrated in a small number of coefficients—then accurately recovering only those dominant coefficients suffices to approximateffwell\. This motivates our use of sparse recovery methods to estimateFFfrom a limited number of evaluations offf\.

R2=1−‖f^−f‖2‖f−f¯‖2,where​‖f‖2=∑S⊆\[L\]f​\(S\)2,f¯=12L​∑S⊂\[L\]f⁡\(S\)\.\\displaystyle R^\{2\}=1\-\\frac\{\|\|\\hat\{f\}\-f\|\|^\{2\}\}\{\|\|f\-\\bar\{f\}\|\|^\{2\}\},\\quad\\text\{where \}\|\|f\|\|^\{2\}=\\sum\_\{S\\subseteq\[L\]\}f\(S\)^\{2\},\\bar\{f\}=\\frac\{1\}\{2^\{L\}\}\\sum\_\{S\\subset\[L\]\}f\(S\)\.\(5\)
We use the sparse recovery algorithm SPEX\([Kang et al\., 2025](https://arxiv.org/html/2609.30456#bib.bib29)\)to accurately compute Fourier coefficients that describe a chosen edit\-set value function during a single feedback iteration\. We study properties of both value functionsfa​v​gf\_\{avg\}andfm​a​xf\_\{max\}for the pretrained protein model by using SPEX\. Across the test set and under a restricted compute budget that limits the total number of Fourier coefficients to be less than one thousand,fa​v​gf\_\{avg\}achieves an averageR2R^\{2\}of 0\.84 andfm​a​xf\_\{max\}achieves an averageR2R^\{2\}of 0\.73\. Notice thatfa​v​gf\_\{avg\}achieves a higherR2R^\{2\}thanfm​a​xf\_\{max\}under the same compute budget\. This is likely because the maximum of random variables tends to have a larger variance than the average\. We also show spectral analysis offa​v​gf\_\{avg\}for two example proteins in Figure[3](https://arxiv.org/html/2609.30456#S4.F3)\. Most coefficients have order at most 3 and both proteins reach anR2R^\{2\}of at least 0\.7\. These results indicate that the Fourier coefficients are sparse and concentrated at lower orders\. Similar results appear for all proteins in the test set and are shown in Appendix[D\.2](https://arxiv.org/html/2609.30456#A4.SS2)along with additional experimental details\. An interesting difference between the two protein examples is that first\-order coefficients dominate for2KVVwhile higher\-order interactions are more significant forR6\-560\. These differences suggest that interactions between edit\-positions depend on the properties of the underlying proteins, which are strongly influenced by interactions among their respective amino acids\. When mainly first\-order coefficients describe the value function, like for2KVV, a first\-order sparse recovery method will likely be sufficient to learn a usefulf^\\hat\{f\}\. We further investigate this in Appendix[G](https://arxiv.org/html/2609.30456#A7)\.

### 4\.2Spectral Feedback Algorithm

Figure 4:Spectral Feedback learns a sparse Fourier representation of the edit\-set value function, then solves for the max\-reward set via integer optimization over the sparse coefficients\.The Spectral Feedback algorithm alternates edit\-position selection and model sampling, using any off\-the\-shelf discrete diffusion model as shown in Figure[1](https://arxiv.org/html/2609.30456#S1.F1)\. The algorithm is described below\.

Step 1 \- Select edit\-positions \(Algorithm[1](https://arxiv.org/html/2609.30456#alg1)\)\.Given an amino acid sequencexxand a max\-orderkk, the first step selects a set of edit\-positionsS⊆\[L\]S\\subseteq\[L\]such that\|S\|≤k\|S\|\\leq k\. This step, depicted in Figure[4](https://arxiv.org/html/2609.30456#S4.F4), is done in the following sub\-steps: sample sets of edit\-positions, evaluate corresponding rewards, fit a sparse Fourier approximation, and solve for the optimal positions\. The sample sets are assigned rewards using Equation[2](https://arxiv.org/html/2609.30456#S4.E2)\. A sparse Fourier approximation is then fit with the sampled set\-reward pairs\. This can be done with any sparse recovery method\. Finally, edit\-positions are selected by solving an optimization problem over the learned sparse Fourier approximation\. The setup of this optimization problem is described in Appendix[D\.1](https://arxiv.org/html/2609.30456#A4.SS1)\.

Step 2 \- Sample the Protein Model \(Algorithm[2](https://arxiv.org/html/2609.30456#alg2)\)\.After generating a setSSof edit\-positions in Step 1, the corresponding amino acids are re\-masked as described in Equation[1](https://arxiv.org/html/2609.30456#S2.E1)\. The partially masked sequence is then passed back through the discrete diffusion reverse process\. Because the sequence is only partially masked, fewer iterations in the reverse process are necessary to generate a new protein than when initializing with a fully masked sequence\.

Require:L,D,k,N∈ℕL,D,k,N\\in\\mathbb\{N\},0<γ<10<\\gamma<1\. VocabularyVVwith mask tokenMM\. Protein backboney∈Yy\\in Y\.

Algorithm 1Spectral Edit\-Set SelectionInput:

x∈VLx\\in V^\{L\}
Sj⊆\[L\]​∀j∈\[D\]S\_\{j\}\\subseteq\[L\]\\;\\forall j\\in\[D\]

i∈Sj​with prob\.​γi\\in S\_\{j\}\\;\\text\{with prob\.\}\\;\\gamma⊳\\trianglerightSample edit sets

r~j←f⁡\(Sj,x,y\)\\tilde\{r\}\_\{j\}\\leftarrow f\(S\_\{j\},x,y\)⊳\\trianglerightEquation[2](https://arxiv.org/html/2609.30456#S4.E2)

f^←ProxySPEX​\(r~,M\)\\hat\{f\}\\leftarrow\\text\{ProxySPEX\}\(\\tilde\{r\},M\)⊳\\trianglerightSparse Recovery

S∗←arg⁡maxS⊆\[L\],\|S\|≤k​f^​\(S\)S^\{\*\}\\leftarrow\\displaystyle\\underset\{S\\subseteq\[L\],\\;\|S\|\\leq k\}\{\\arg\\max\}\\;\\hat\{f\}\(S\)⊳\\trianglerightAppendix[D\.1](https://arxiv.org/html/2609.30456#A4.SS1)

return​S∗\\textbf\{return\}\\;S^\{\*\}

Algorithm 2Protein Model FeedbackInput:

x∈VLx\\in V^\{L\}
xi←M​∀i∈\[L\]x\_\{i\}\\leftarrow M\\;\\forall i\\in\[L\]⊳\\trianglerightFull mask initialization

S←\[L\]S\\leftarrow\[L\]⊳\\trianglerightEdit all tokens initially

for

iiin

1,…,N−11,\\ldots,N\-1do

x∼p\(⋅\|S,x,y\)x\\sim p\(\\cdot\|S,x,y\)⊳\\trianglerightProtein model

S←EditSelection​\(x\)S\\leftarrow\\text\{EditSelection\}\(x\)⊳\\trianglerightAlgorithm[1](https://arxiv.org/html/2609.30456#alg1)

endfor

x∼p\(⋅\|S,x,y\)x\\sim p\(\\cdot\|S,x,y\)⊳\\trianglerightLast feedback iteration

return​x\\textbf\{return\}\\;x

Sparse Recovery Method\. While any sparse recovery method can be used in Spectral Feedback \(e\.g\.,[Kang et al\. \(2025\)](https://arxiv.org/html/2609.30456#bib.bib29);[Li et al\. \(2014\)](https://arxiv.org/html/2609.30456#bib.bib25);[Amrollahi et al\. \(2019\)](https://arxiv.org/html/2609.30456#bib.bib32)\), we use ProxySPEX\([Butler et al\., 2025](https://arxiv.org/html/2609.30456#bib.bib11)\), a method that uses Gradient Boosted Trees to learn a sparse Fourier representation due to its sample efficiency\. To sample for ProxySPEX, we generate binary masks where each element is an independent Bernoulli random variable with parameterγ\\gamma\.

Value Function\. The choice of the value function from equation[2](https://arxiv.org/html/2609.30456#S4.E2)is dependent on the way in which the underlying discrete diffusion model is sampled from\. When Spectral Feedback is applied directly to a diffusion model such as a pretrained or RL fine\-tuned model like DRAKES, we usefa​v​gf\_\{avg\}to approximate the expected reward of the resulting proteins\. However, additional test\-time alignment methods can be applied on top of a diffusion model within the feedback loop\. We demonstrate this by applying Spectral Feedback to Best\-of\-N with a pretrained diffusion model\. To more naturally follow the dynamics of Best\-of\-N, we usefm​a​xf\_\{max\}for the edit\-set value function\.fm​a​xf\_\{max\}should be used instead offa​v​gf\_\{avg\}whenever there is an underlying greedy test\-time alignment method such as Best\-of\-N or Beam Search\. In these methods, only the largest valued sample matters, so we should not penalize an edit\-set for its resulting low\-reward samples by using an average\.

## 5Experiments

![Refer to caption](https://arxiv.org/html/2609.30456v1/ResultsPlot_img.png)Figure 5:\(left\)Spectral Feedback is applied to several underlying protein models: pretrained, Best\-of\-10 with pretrained, and DRAKES\. Spectral Feedback improves the alignment reward of each algorithm in only a few iterations\.\(middle\)Percentage of stable proteins \(evaluationΔ​Δ​G\>0\\Delta\\Delta G\>0\) and the resulting success rate whens​c​R​M​S​D<2scRMSD<2are shown\. Thes​c​R​M​S​DscRMSDconstraint and metric averages are in Table[1](https://arxiv.org/html/2609.30456#S5.T1)\. \(right\) ESMFold protein structures comparing the wild\-type \(red\) and Spectral Feedback \(green\) sequences demonstrate that Spectral Feedback preserves the underlying inverse protein folding model’s capabilities\.### 5\.1Experimental Setup

We use the protein backbones from the Megascale dataset for evaluating our algorithm\. The evaluation metrics arestability,scRMSD, andnaturalness\.

- •Thestabilityoracle represents the quantityΔ​Δ​G=Δ​Gw​i​l​d−Δ​Ga​l​i​g​n\\Delta\\Delta G=\\Delta G\_\{wild\}\-\\Delta G\_\{align\}whereΔ​G\\Delta Gis the Gibbs free energy of the corresponding protein\. A positiveΔ​Δ​G\\Delta\\Delta Gmeans a lowerΔ​G\\Delta Gthan the wild\-type of a protein backbone, and thus a more stable protein\.
- •The self\-consistency RMSD \(scRMSD\) oracle represents the RMSD between the target protein backbone of the inverse folding process and the predicted folded structure of a generated amino acid sequence\. We use ESMFold\([Lin et al\., 2023](https://arxiv.org/html/2609.30456#bib.bib19)\)to predict the structure of a sequence and then calculate the RMSD with the given backbone\.
- •Thenaturalnessoracle represents the log\-likelihood of an amino acid sequence\. We use ProtGPT2\([Ferruz et al\., 2022](https://arxiv.org/html/2609.30456#bib.bib18)\), an auto\-regressive language model that captures the distribution of protein amino acid sequences\. This metric complements scRMSD in identifying overfitting to the stability oracle\.

We use theΔ​Δ​G\\Delta\\Delta Gmetric as the target reward for aligning the underlying protein model\. DRAKES trained twoΔ​Δ​G\\Delta\\Delta Goracles with the Megascale dataset: one for alignment and one for evaluation\([Wang et al\., 2025](https://arxiv.org/html/2609.30456#bib.bib1)\)\. To fairly compare with their work, we also align and evaluateΔ​Δ​G\\Delta\\Delta Gwith each of these respective oracles\. Following previous works\([Campbell et al\., 2024](https://arxiv.org/html/2609.30456#bib.bib15);[Nisonoff et al\., 2025](https://arxiv.org/html/2609.30456#bib.bib16);[Wang et al\., 2025](https://arxiv.org/html/2609.30456#bib.bib1)\), we classify asuccessfulprotein as one that satisfiesΔ​Δ​G\>0\\Delta\\Delta G\>0ands​c​R​M​S​D<2scRMSD<2\. These constraints ensure that proteins are at least as stable as their wild\-type variants and that their folded structure resembles the target structure\.

We use both the pretrained and RL models trained from DRAKES to test our method\. These models are built with the ProteinMPNN architecture which uses a graph neural network to capture local interactions between amino acids and incorporates extracted features from the backbone structure\([Dauparas et al\., 2022](https://arxiv.org/html/2609.30456#bib.bib3);[Wang et al\., 2025](https://arxiv.org/html/2609.30456#bib.bib1)\)\. We also test using Spectral Feedback with Best\-of\-N on the pretrained model\. Following our discussion in section[4\.2](https://arxiv.org/html/2609.30456#S4.SS2), we usefa​v​gf\_\{avg\}for the pretrained and DRAKES models, andfm​a​xf\_\{max\}for the pretrained model with Best\-of\-N\. Spectral Feedback improves the stability for each algorithm and the success rate of the pretrained model with and without Best\-of\-N\. Furthermore, Best\-of\-N has diminishing returns and Spectral Feedback is able to improve Best\-of\-10 beyond these limitations as discussed in Appendix[C](https://arxiv.org/html/2609.30456#A3)\. In Appendix[E](https://arxiv.org/html/2609.30456#A5)we study how the reduction in success rate for DRAKES is related to issues within the original DRAKES sample distribution\. Many other alignment methods have been explored, yet they commonly target one of two distributions: the argmax distribution and a weighted Gibbs distributionpa​\(x\)=pp​r​e​\(x\)​exp⁡\(α⋅r⁡\(x\)\)p\_\{a\}\(x\)=p\_\{pre\}\(x\)\\exp\(\{\\alpha\\cdot r\(x\)\}\)\. Techniques such as Best\-of\-N, Beam Search, and MCTS all approximate an argmax by searching the sequence space for the highest\-reward sample\. Common RL objectives with KL regularization, like the objective used in DRAKES\([Wang et al\., 2025](https://arxiv.org/html/2609.30456#bib.bib1)\), are maximized with weighted Gibbs distributions\. Our experiments show promise that Spectral Feedback will succeed in improving additional alignment methods because of the similarity of the target distributions across alignment techniques\.

### 5\.2Baseline Edit Selection Methods

Spectral Feedback is unique in that it chooses edit\-positions based on interactions between tokens\. To establish that these interactions are important, we compare the algorithm’s performance with methods that do not account for interactions\. These are described below\.

Gradient\-Weighting\. This method selects edit\-positions based on the reward function’s gradient at those positions\. In particular, for a sequencexx, theit​hi^\{th\}token position is assigned a weightw⁡\(i\)=hi​\(x\)T​∇ir​\(hi​\(x\)\)w\(i\)=h\_\{i\}\(x\)^\{T\}\\nabla\_\{i\}r\(h\_\{i\}\(x\)\)wherehi​\(x\)h\_\{i\}\(x\)is the embedding for theit​hi^\{th\}token\. This method is meant to mirror the use of saliency maps in[Gruver et al\. \(2023\)](https://arxiv.org/html/2609.30456#bib.bib5)for edit\-position selection in Langevin dynamics updates, but with signed importance scores\. For edit\-positions, we select the tokens with the most negative scores since these correspond to tokens that lead to lower reward\.

Token Exclusion\. Whereas Spectral Feedback directly selects an optimal set of edit\-positions, Token Exclusion looks at individual tokens to determine edit\-positions\. It works by calculating the expected reward from masking each token in isolation and then selects the highest\-reward positions\.

Argmax Selection\. To show that the sparse recovery methods successfully use observed edit sets to identify a better edit set, we compare against the argmax method which selects the best observed sample edit set\. While this method will suffer if sample edit\-sets are suboptimal, a sparse recovery method still has hope of finding a rare but optimal edit\-set\.

Hill Climbing\. We implement a randomized 1\-flip local search, as described in[Szeider \(2011\)](https://arxiv.org/html/2609.30456#bib.bib34), for an alternative greedy baseline\. This method iteratively improves its edit selection by uniformly randomly adding or removing an edit\-position from the current selection and accepting the change if there was an improvement in reward\. A drawback is that the algorithm is inherently serial so parallelization is limited to single\-set reward calculations\.

### 5\.3Results on Protein Inverse Folding

Table 1:Spectral Feedback improves alignment and evaluation oracle rewards across all three base methods, with the largest gains on the Pretrained and Best\-of\-10 baselines\. DRAKES \+ Spectral achieves the highest evaluationΔ​Δ​G\\Delta\\Delta Goverall \(1\.38\) but trades off structural quality \(scRMSD\)\. Each pair shows baseline followed by its Spectral\-enhanced counterpart\.Modular Integrations\.We first demonstrate how Spectral Feedback can be used in conjunction with different underlying protein models, serving as a modular alignment technique\.

The reward trajectories in Figure[5](https://arxiv.org/html/2609.30456#S5.F5)are of the alignment oracle whereas the bar plot shows the evaluation oracle results\. The evaluation oracle distributions are also visualized in Figure[2](https://arxiv.org/html/2609.30456#S1.F2)\. The results show that Spectral Feedback improves the performance of whichever underlying protein model it samples from\. Additionally, in Table[1](https://arxiv.org/html/2609.30456#S5.T1), the reward increases for both the alignment and evaluation oracles so there was not excessive overfitting\. The relatively stable log\-likelihoods and the high success rates in Figure[5](https://arxiv.org/html/2609.30456#S5.F5)further support Spectral Feedback’s robustness\. However, the averagescRMSDincreases when Spectral Feedback is applied to DRAKES, thus limiting the gains that can be made in protein success rate\. This may be a reflection of the DRAKES alignment distribution rather than the Spectral Feedback algorithm\. Further discussion regarding thescRMSDincrease is in Appendix[E](https://arxiv.org/html/2609.30456#A5)\.

An important aspect of Spectral Feedback is its efficiency\. Figure[5](https://arxiv.org/html/2609.30456#S5.F5)shows that Spectral Feedback achieves the same performance as Best\-of\-10 in only 2 iterations\. Best\-of\-10 requires 10 full reverse processes while Spectral Feedback involves one full reverse process followed by partial processes that fill in only a subset of tokens\. While the value functionffmakes calls to the protein model, if this reward oracle is a distinct smaller neural network as in[Gruver et al\. \(2023\)](https://arxiv.org/html/2609.30456#bib.bib5)and[Xiong et al\. \(2026\)](https://arxiv.org/html/2609.30456#bib.bib20), the performance gains from sample efficiency can be fully realized\.

Baseline Edit\-Position Selection\. Now, we show that Spectral Feedback scales well in alignment performance and latency\. We evaluate Spectral Feedback against baseline edit\-set selection methods, measuring how alignment reward scales with both the number of edit\-set samples and total feedback iteration time\. All results are from a single feedback iteration\. Spectral Feedback and Hill Climbing exhibit similar scaling, but Spectral Feedback achieves a consistently better latency–reward tradeoff due to inherent seriality in Hill Climbing which prevents large batch sizes for reward calculations\. Argmax and Spectral use the exact same samples in these experiments; however, Spectral Feedback incurs additional overhead from ProxySPEX fitting and optimization to find a better edit\-set than in the observed samples\. The latency trade\-off plot shows that this overhead becomes insignificant, allowing for Spectral Feedback to dominate\. Exclusion and Gradient methods have weaker alignment, but are fast alternatives\.

Figure 6:Spectral Feedback shows strong performance in scaling of edit\-set samples and latency\.

## 6Discussion

Conclusion\. In this work, we formulate protein sequence generation with access to stability reward oracles and protein language models as a test\-time scaling problem\. We introduce a novel algorithm, Spectral Feedback, that iteratively identifies the most informative subset of edit\-positions in a generated sequence and re\-samples them through the diffusion reverse process, treating edit\-set selection as the central scaling target\. The crucial insight we develop is that the value of editing an amino acid depends on which other amino acids are edited alongside it, so the reward to optimize is a function over subsets of positions\. We show that this function is sparse under the Fourier transform, consistent with biological epistasis\. We use sparse recovery methods to approximate this function and optimize over it within a black\-box meta\-loop\. For inverse folding with the Megascale dataset, this method improves the alignment reward across different protein models while maintaining structural self\-consistency that also leads to increased success rates\.

Limitations & Future Work\. A limitation of our work is that the value functions we use are non\-deterministic, which can make it harder to reach largerR2R^\{2\}values\. This can also be costly due to the additional protein model calls incurred by simulating a single step of the reverse process\. Instead, it could be beneficial to train a value function for partially masked sequences to allow for approximations to reach greaterR2R^\{2\}and improve runtime\. Another limitation is that it requires a large number of edit\-set samples for sparse recovery to successfully learn a sparse Fourier representation\. However, once in this high\-sample regime, our algorithm scales better than competing methods\. In the future, the scaling of alternative sparse recovery methods could be further explored\. Our algorithm could also be applied to more domains such as DNA design and text generation\. It would be interesting to see how results scale with larger vocabularies and longer sequences\.

## Acknowledgments and Disclosure of Funding

We thank Jennifer Listgarten and her lab group for fruitful discussions and words of advice\. This research used both the DeltaAI advanced computing and data resource, which is supported by the National Science Foundation \(award OAC 2320345\) and the State of Illinois, and the Delta advanced computing and data resource which is supported by the National Science Foundation \(award OAC 2005572\) and the State of Illinois\([Gropp et al\., 2023](https://arxiv.org/html/2609.30456#bib.bib42)\)\. Delta and DeltaAI are joint efforts of the University of Illinois Urbana\-Champaign and its National Center for Supercomputing Applications\. Additionally, this research used the Anvil supercomputer at Purdue University which is also supported by the National Science Foundation \(award OAC 2005632\)\([Song et al\., 2022](https://arxiv.org/html/2609.30456#bib.bib41)\)\.

## References

- Aghazadehet al\.\(2021\)A\. Aghazadeh, H\. Nisonoff, O\. Ocal, D\. H\. Brookes, Y\. Huang, O\. O\. Koyluoglu, J\. Listgarten, and K\. RamchandranEpistatic net allows the sparse spectral regularization of deep neural networks for inferring fitness functions\.Nature Communications12\(1\),pp\. 5225\.External Links:[Document](https://dx.doi.org/10.1038/s41467-021-25371-3),[Link](https://www.nature.com/articles/s41467-021-25371-3)Cited by:[§1](https://arxiv.org/html/2609.30456#S1.p3.1),[§3](https://arxiv.org/html/2609.30456#S3.p4.1)\.
- Aghazadehet al\.\(2020\)A\. Aghazadeh, O\. Ocal, and K\. RamchandranCRISPRLand: interpretable large\-scale inference of dna repair landscape based on a spectral approach\.Bioinformatics36\(Supplement 1\),pp\. i560–i568\.External Links:[Document](https://dx.doi.org/10.1093/bioinformatics/btaa505)Cited by:[§1](https://arxiv.org/html/2609.30456#S1.p3.1),[§3](https://arxiv.org/html/2609.30456#S3.p4.1)\.
- Amrollahiet al\.\(2019\)A\. Amrollahi, A\. Zandieh, M\. Kapralov, and A\. KrauseEfficiently learning fourier sparse set functions\.Advances in Neural Information Processing Systems32\.Cited by:[§3](https://arxiv.org/html/2609.30456#S3.p4.1),[§4\.2](https://arxiv.org/html/2609.30456#S4.SS2.p4.1)\.
- Bennettet al\.\(2026\)N\. R\. Bennett, J\. L\. Watson, R\. J\. Ragotte, A\. J\. Borst, D\. L\. See, C\. Weidle, R\. Biswas, Y\. Yu, E\. L\. Shrock, R\. Ault, P\. J\. Y\. Leung, B\. Huang, I\. Goreshnik, J\. Tam, K\. D\. Carr, B\. Singer, C\. Criswell, B\. I\. M\. Wicky, D\. Vafeados, M\. Garcia Sanchez, H\. M\. Kim, S\. Vázquez Torres, S\. Chan, S\. M\. Sun, T\. T\. Spear, Y\. Sun, K\. O’Reilly, J\. M\. Maris, N\. G\. Sgourakis, R\. A\. Melnyk, C\. C\. Liu, and D\. BakerAtomically accurate de novo design of antibodies with rfdiffusion\.Nature649\(8095\),pp\. 183–193\.External Links:[Document](https://dx.doi.org/10.1038/s41586-025-09721-5),[Link](https://www.nature.com/articles/s41586-025-09721-5)Cited by:[§3](https://arxiv.org/html/2609.30456#S3.p2.1)\.
- Butleret al\.\(2025\)L\. Butler, A\. Agarwal, J\. S\. Kang, Y\. E\. Erginbas, B\. Yu, and K\. RamchandranProxySPEX: inference\-efficient interpretability via sparse feature interactions in llms\.External Links:2505\.17495,[Link](https://arxiv.org/abs/2505.17495)Cited by:[§A\.3](https://arxiv.org/html/2609.30456#A1.SS3.p6.1),[§D\.1](https://arxiv.org/html/2609.30456#A4.SS1.p1.1),[§1](https://arxiv.org/html/2609.30456#S1.p3.1),[§1](https://arxiv.org/html/2609.30456#S1.p4.1),[§3](https://arxiv.org/html/2609.30456#S3.p4.1),[§4\.2](https://arxiv.org/html/2609.30456#S4.SS2.p4.1)\.
- Campbellet al\.\(2024\)A\. Campbell, J\. Yim, R\. Barzilay, T\. Rainforth, and T\. JaakkolaGenerative flows on discrete state\-spaces: enabling multimodal flows with applications to protein co\-design\.External Links:2402\.04997,[Link](https://arxiv.org/abs/2402.04997)Cited by:[§5\.1](https://arxiv.org/html/2609.30456#S5.SS1.p1.2)\.
- Cemriet al\.\(2024\)M\. Cemri, A\. Jalal, and K\. RamchandranDiscrete diffusion posterior sampling for protein design\.InICML 2024 Workshop on Structured Probabilistic Inference & Generative Modeling,Cited by:[§F\.2](https://arxiv.org/html/2609.30456#A6.SS2.p3.1),[§1](https://arxiv.org/html/2609.30456#S1.p1.1),[§3](https://arxiv.org/html/2609.30456#S3.p1.1)\.
- Cordell \(2002\)H\. J\. CordellEpistasis: what it means, what it doesn’t mean, and statistical methods to detect it in humans\.Human Molecular Genetics11\(20\),pp\. 2463–2468\.External Links:[Document](https://dx.doi.org/10.1093/hmg/11.20.2463),[Link](https://doi.org/10.1093/hmg/11.20.2463)Cited by:[§3](https://arxiv.org/html/2609.30456#S3.p4.1)\.
- Darmawanet al\.\(2025\)J\. T\. Darmawan, Y\. Gal, and P\. NotinSampling protein language models for functional protein design\.bioRxiv\.External Links:[Document](https://dx.doi.org/10.1101/2025.09.14.676087),[Link](https://www.biorxiv.org/content/early/2025/09/17/2025.09.14.676087),https://www\.biorxiv\.org/content/early/2025/09/17/2025\.09\.14\.676087\.full\.pdfCited by:[§1](https://arxiv.org/html/2609.30456#S1.p1.1),[§3](https://arxiv.org/html/2609.30456#S3.p2.1)\.
- Dauparaset al\.\(2022\)J\. Dauparas, I\. Anishchenko, N\. Bennett, H\. Bai, R\. J\. Ragotte, L\. F\. Milles, B\. I\. M\. Wicky, A\. Courbet, R\. J\. de Haas, N\. Bethel, P\. J\. Y\. Leung, T\. F\. Huddy, S\. Pellock, D\. Tischer, F\. Chan, B\. Koepnick, H\. Nguyen, A\. Kang, B\. Sankaran, A\. K\. Bera, N\. P\. King, and D\. BakerRobust deep learning–based protein sequence design using proteinmpnn\.Science378\(6615\),pp\. 49–56\.External Links:[Document](https://dx.doi.org/10.1126/science.add2187),[Link](https://www.science.org/doi/abs/10.1126/science.add2187),https://www\.science\.org/doi/pdf/10\.1126/science\.add2187Cited by:[§3](https://arxiv.org/html/2609.30456#S3.p1.1),[§5\.1](https://arxiv.org/html/2609.30456#S5.SS1.p2.1)\.
- Ferruzet al\.\(2022\)N\. Ferruz, S\. Schmidt, and B\. HöckerProtGPT2 is a deep unsupervised language model for protein design\.Nature Communications13\(1\),pp\. 4348\.External Links:[Document](https://dx.doi.org/10.1038/s41467-022-32007-7),[Link](https://doi.org/10.1038/s41467-022-32007-7)Cited by:[§F\.2](https://arxiv.org/html/2609.30456#A6.SS2.p3.1),[§1](https://arxiv.org/html/2609.30456#S1.p6.1),[3rd item](https://arxiv.org/html/2609.30456#S5.I1.i3.p1.1)\.
- Groppet al\.\(2023\)W\. Gropp, T\. Boerner, B\. Bode, and G\. BauerDelta: Balancing GPU Performance with Advanced System Interfaces\.\(en\)\.External Links:[Link](https://hdl.handle.net/2142/117179)Cited by:[Acknowledgments and Disclosure of Funding](https://arxiv.org/html/2609.30456#Sx1.p1.1)\.
- Gruveret al\.\(2023\)N\. Gruver, S\. Stanton, N\. C\. Frey, T\. G\. J\. Rudner, I\. Hotzel, J\. Lafrance\-Vanasse, A\. Rajpal, K\. Cho, and A\. G\. WilsonProtein design with guided discrete diffusion\.External Links:2305\.20009,[Link](https://arxiv.org/abs/2305.20009)Cited by:[§F\.2](https://arxiv.org/html/2609.30456#A6.SS2.p3.1),[§1](https://arxiv.org/html/2609.30456#S1.p1.1),[§2](https://arxiv.org/html/2609.30456#S2.p5.1),[§3](https://arxiv.org/html/2609.30456#S3.p1.1),[§3](https://arxiv.org/html/2609.30456#S3.p2.1),[§3](https://arxiv.org/html/2609.30456#S3.p3.1),[§4\.1](https://arxiv.org/html/2609.30456#S4.SS1.p1.2),[§5\.2](https://arxiv.org/html/2609.30456#S5.SS2.p2.1),[§5\.3](https://arxiv.org/html/2609.30456#S5.SS3.p3.1)\.
- Gupteet al\.\(2021\)T\. M\. Gupte, M\. Ritt, and S\. SivaramakrishnanChapter Seven \- ER/K\-link—Leveraging a native protein linker to probe dynamic cellular interactions\.InLinkers in Biomacromolecules,M\. Merkx \(Ed\.\),Methods in Enzymology, Vol\.647,pp\. 173–208\.External Links:[Link](https://www.sciencedirect.com/science/article/pii/S0076687920303487),[Document](https://dx.doi.org/10.1016/bs.mie.2020.10.002)Cited by:[§G\.1](https://arxiv.org/html/2609.30456#A7.SS1.p2.1)\.
- Gurobi Optimization, LLC \(2026\)Gurobi Optimization, LLCGurobi Optimizer Reference Manual\.External Links:[Link](https://www.gurobi.com/)Cited by:[§D\.1](https://arxiv.org/html/2609.30456#A4.SS1.p5.1)\.
- Hertzet al\.\(2022\)A\. Hertz, R\. Mokady, J\. Tenenbaum, K\. Aberman, Y\. Pritch, and D\. Cohen\-OrPrompt\-to\-prompt image editing with cross attention control\.External Links:2208\.01626,[Link](https://arxiv.org/abs/2208.01626)Cited by:[§1](https://arxiv.org/html/2609.30456#S1.p1.1)\.
- Huanget al\.\(2025\)A\. Huang, A\. Block, Q\. Liu, N\. Jiang, A\. Krishnamurthy, and D\. J\. FosterIs best\-of\-n the best of them? coverage, scaling, and optimality in inference\-time alignment\.External Links:2503\.21878,[Link](https://arxiv.org/abs/2503.21878)Cited by:[§1](https://arxiv.org/html/2609.30456#S1.p1.1),[§3](https://arxiv.org/html/2609.30456#S3.p2.1)\.
- Kanget al\.\(2025\)J\. S\. Kang, L\. Butler, A\. Agarwal, Y\. E\. Erginbas, R\. Pedarsani, K\. Ramchandran, and B\. YuSPEX: scaling feature interaction explanations for llms\.External Links:2502\.13870,[Link](https://arxiv.org/abs/2502.13870)Cited by:[§1](https://arxiv.org/html/2609.30456#S1.p4.1),[§3](https://arxiv.org/html/2609.30456#S3.p4.1),[§4\.1](https://arxiv.org/html/2609.30456#S4.SS1.p6.1),[§4\.2](https://arxiv.org/html/2609.30456#S4.SS2.p4.1)\.
- Leeet al\.\(2025\)S\. Lee, S\. Kim, S\. Kim, J\. Park, and D\. ParkEffective test\-time scaling of discrete diffusion through iterative refinement\.External Links:2511\.05562,[Link](https://arxiv.org/abs/2511.05562)Cited by:[§1](https://arxiv.org/html/2609.30456#S1.p1.1),[§2](https://arxiv.org/html/2609.30456#S2.p5.1),[§3](https://arxiv.org/html/2609.30456#S3.p3.1)\.
- Liet al\.\(2014\)X\. Li, J\. K\. Bradley, S\. Pawar, and K\. RamchandranThe spright algorithm for robust sparse hadamard transforms\.InProceedings of the IEEE International Symposium on Information Theory \(ISIT\),Honolulu, HI, USA,pp\. 1857–1861\.External Links:[Document](https://dx.doi.org/10.1109/ISIT.2014.6875155),[Link](https://ieeexplore.ieee.org/document/6875155)Cited by:[§1](https://arxiv.org/html/2609.30456#S1.p3.1),[§1](https://arxiv.org/html/2609.30456#S1.p4.1),[§3](https://arxiv.org/html/2609.30456#S3.p4.1),[§4\.2](https://arxiv.org/html/2609.30456#S4.SS2.p4.1)\.
- Lianget al\.\(2025\)Y\. Liang, Y\. Liang, L\. Lai, and N\. ShroffDiscrete diffusion models: novel analysis and new sampler guarantees\.External Links:2509\.16756,[Link](https://arxiv.org/abs/2509.16756)Cited by:[Appendix B](https://arxiv.org/html/2609.30456#A2.p1.3),[Appendix B](https://arxiv.org/html/2609.30456#A2.p3.2),[§2](https://arxiv.org/html/2609.30456#S2.p1.1)\.
- Linet al\.\(2023\)Z\. Lin, H\. Akin, R\. Rao, B\. Hie, Z\. Zhu, W\. Lu, N\. Smetanin, R\. Verkuil, O\. Kabeli, Y\. Shmueli, A\. dos Santos Costa, M\. Fazel\-Zarandi, T\. Sercu, S\. Candido, and A\. RivesEvolutionary\-scale prediction of atomic\-level protein structure with a language model\.Science379\(6637\),pp\. 1123–1130\.External Links:[Document](https://dx.doi.org/10.1126/science.ade2574),[Link](https://www.science.org/doi/abs/10.1126/science.ade2574),https://www\.science\.org/doi/pdf/10\.1126/science\.ade2574Cited by:[§1](https://arxiv.org/html/2609.30456#S1.p6.1),[2nd item](https://arxiv.org/html/2609.30456#S5.I1.i2.p1.1)\.
- Madaanet al\.\(2023\)A\. Madaan, N\. Tandon, P\. Gupta, S\. Hallinan, L\. Gao, S\. Wiegreffe, U\. Alon, N\. Dziri, S\. Prabhumoye, Y\. Yang, S\. Gupta, B\. P\. Majumder, K\. Hermann, S\. Welleck, A\. Yazdanbakhsh, and P\. ClarkSelf\-refine: iterative refinement with self\-feedback\.External Links:2303\.17651,[Link](https://arxiv.org/abs/2303.17651)Cited by:[§3](https://arxiv.org/html/2609.30456#S3.p3.1)\.
- Menget al\.\(2022\)C\. Meng, Y\. He, Y\. Song, J\. Song, J\. Wu, J\. Zhu, and S\. ErmonSDEdit: guided image synthesis and editing with stochastic differential equations\.External Links:2108\.01073,[Link](https://arxiv.org/abs/2108.01073)Cited by:[§1](https://arxiv.org/html/2609.30456#S1.p1.1)\.
- Mokadyet al\.\(2022\)R\. Mokady, A\. Hertz, K\. Aberman, Y\. Pritch, and D\. Cohen\-OrNull\-text inversion for editing real images using guided diffusion models\.External Links:2211\.09794,[Link](https://arxiv.org/abs/2211.09794)Cited by:[§1](https://arxiv.org/html/2609.30456#S1.p1.1)\.
- Nisonoffet al\.\(2025\)H\. Nisonoff, J\. Xiong, S\. Allenspach, and J\. ListgartenUnlocking guidance for discrete state\-space diffusion and flow models\.External Links:2406\.01572,[Link](https://arxiv.org/abs/2406.01572)Cited by:[§5\.1](https://arxiv.org/html/2609.30456#S5.SS1.p1.2)\.
- O’Donnell \(2014\)R\. O’DonnellAnalysis of boolean functions\.Cambridge University Press\.Cited by:[§4\.1](https://arxiv.org/html/2609.30456#S4.SS1.p2.1)\.
- Poelwijket al\.\(2019\)F\. J\. Poelwijk, M\. Socolich, and R\. RanganathanLearning the pattern of epistasis linking genotype and phenotype in a protein\.Nature Communications10\(1\),pp\. 4213\.External Links:[Document](https://dx.doi.org/10.1038/s41467-019-12130-8),[Link](https://www.nature.com/articles/s41467-019-12130-8)Cited by:[§1](https://arxiv.org/html/2609.30456#S1.p3.1),[§3](https://arxiv.org/html/2609.30456#S3.p4.1)\.
- Reidet al\.\(2022\)M\. Reid, V\. J\. Hellendoorn, and G\. NeubigDiffusER: discrete diffusion via edit\-based reconstruction\.External Links:2210\.16886,[Link](https://arxiv.org/abs/2210.16886)Cited by:[§1](https://arxiv.org/html/2609.30456#S1.p1.1),[§2](https://arxiv.org/html/2609.30456#S2.p5.1),[§3](https://arxiv.org/html/2609.30456#S3.p3.1)\.
- Scheibleret al\.\(2015\)R\. Scheibler, S\. Haghighatshoar, and M\. VetterliA fast hadamard transform for signals with sublinear sparsity in the transform domain\.IEEE Transactions on Information Theory61\(4\),pp\. 2115–2132\.External Links:[Document](https://dx.doi.org/10.1109/TIT.2015.2404441)Cited by:[§1](https://arxiv.org/html/2609.30456#S1.p4.1),[§3](https://arxiv.org/html/2609.30456#S3.p4.1)\.
- Songet al\.\(2022\)X\. C\. Song, P\. Smith, R\. Kalyanam, X\. Zhu, E\. Adams, K\. Colby, P\. Finnegan, E\. Gough, E\. Hillery, R\. Irvine, A\. Maji, and J\. St\. JohnAnvil \- system architecture and experiences from deployment and early user operations\.InPractice and Experience in Advanced Research Computing 2022: Revolutionary: Computing, Connections, You,PEARC ’22,New York, NY, USA\.External Links:ISBN 9781450391610,[Link](https://doi.org/10.1145/3491418.3530766),[Document](https://dx.doi.org/10.1145/3491418.3530766)Cited by:[Acknowledgments and Disclosure of Funding](https://arxiv.org/html/2609.30456#Sx1.p1.1)\.
- Sunet al\.\(2023\)H\. Sun, L\. Yu, B\. Dai, D\. Schuurmans, and H\. DaiScore\-based continuous\-time discrete diffusion models\.External Links:2211\.16750,[Link](https://arxiv.org/abs/2211.16750)Cited by:[Appendix B](https://arxiv.org/html/2609.30456#A2.p1.4)\.
- Szeider \(2011\)S\. SzeiderThe parameterized complexity of k\-flip local search for sat and max sat\.Discrete Optimization8\(1\),pp\. 139–145\.Note:Parameterized Complexity of Discrete OptimizationExternal Links:ISSN 1572\-5286,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.disopt.2010.07.003),[Link](https://www.sciencedirect.com/science/article/pii/S1572528610000526)Cited by:[§5\.2](https://arxiv.org/html/2609.30456#S5.SS2.p5.1)\.
- Tibshirani \(1996\)R\. TibshiraniRegression shrinkage and selection via the lasso\.Journal of the Royal Statistical Society: Series B \(Methodological\)58\(1\),pp\. 267–288\.External Links:[Link](https://www.jstor.org/stable/2346178)Cited by:[§3](https://arxiv.org/html/2609.30456#S3.p4.1)\.
- Tsuboyamaet al\.\(2023\)K\. Tsuboyama, J\. Dauparas, J\. Chen, E\. Laine, Y\. M\. Behbahani, J\. J\. Weinstein, N\. M\. Mangan, S\. Ovchinnikov, and G\. J\. RocklinMega\-scale experimental analysis of protein folding stability in biology and design\.Nature620\(7973\),pp\. 434–444\.External Links:ISBN 1476\-4687,[Document](https://dx.doi.org/10.1038/s41586-023-06328-6),[Link](https://www.nature.com/articles/s41586-023-06328-6)Cited by:[§1](https://arxiv.org/html/2609.30456#S1.p6.1)\.
- Ueharaet al\.\(2025\)M\. Uehara, Y\. Zhao, C\. Wang, X\. Li, A\. Regev, S\. Levine, and T\. BiancalaniInference\-time alignment in diffusion models with reward\-guided generation: tutorial and review\.External Links:2501\.09685,[Link](https://arxiv.org/abs/2501.09685)Cited by:[§1](https://arxiv.org/html/2609.30456#S1.p1.1),[§3](https://arxiv.org/html/2609.30456#S3.p2.1),[§4\.1](https://arxiv.org/html/2609.30456#S4.SS1.p1.2)\.
- Wanget al\.\(2025\)C\. Wang, M\. Uehara, Y\. He, A\. Wang, T\. Biancalani, A\. Lal, T\. Jaakkola, S\. Levine, H\. Wang, and A\. RegevFine\-tuning discrete diffusion models via reward optimization with applications to dna and protein design\.arXiv preprint arXiv:2410\.13643\.External Links:2410\.13643,[Link](https://arxiv.org/abs/2410.13643)Cited by:[§1](https://arxiv.org/html/2609.30456#S1.p1.1),[§1](https://arxiv.org/html/2609.30456#S1.p5.1),[§3](https://arxiv.org/html/2609.30456#S3.p1.1),[§5\.1](https://arxiv.org/html/2609.30456#S5.SS1.p1.2),[§5\.1](https://arxiv.org/html/2609.30456#S5.SS1.p2.1)\.
- Wanget al\.\(2026\)G\. Wang, Y\. Schiff, S\. S\. Sahoo, and V\. KuleshovRemasking discrete diffusion models with inference\-time scaling\.External Links:2503\.00307,[Link](https://arxiv.org/abs/2503.00307)Cited by:[§1](https://arxiv.org/html/2609.30456#S1.p1.1)\.
- Xionget al\.\(2026\)J\. Xiong, I\. Gaur, M\. Lukarska, H\. Nisonoff, L\. M\. Oltrogge, D\. F\. Savage, and J\. ListgartenProteinGuide: on\-the\-fly property guidance for protein sequence generative models\.External Links:2505\.04823,[Link](https://arxiv.org/abs/2505.04823)Cited by:[Appendix C](https://arxiv.org/html/2609.30456#A3.p1.1),[§1](https://arxiv.org/html/2609.30456#S1.p1.1),[§3](https://arxiv.org/html/2609.30456#S3.p2.1),[§4\.1](https://arxiv.org/html/2609.30456#S4.SS1.p1.2),[§5\.3](https://arxiv.org/html/2609.30456#S5.SS3.p3.1)\.
- Xueet al\.\(2025\)E\. Xue, K\. Chen, Z\. Huang, Y\. Ji, and H\. WangIMPROVE: iterative model pipeline refinement and optimization leveraging llm experts\.External Links:2502\.18530,[Link](https://arxiv.org/abs/2502.18530)Cited by:[§3](https://arxiv.org/html/2609.30456#S3.p3.1)\.
- Yehet al\.\(2023\)A\. H\. Yeh, C\. Norn, Y\. Kipnis, D\. Tischer, S\. J\. Pellock, D\. Evans, P\. Ma, G\. R\. Lee, J\. Z\. Zhang, I\. Anishchenko, B\. Coventry, L\. Cao, J\. Dauparas, S\. Halabiya, M\. DeWitt, L\. Carter, K\. N\. Houk, and D\. BakerDe novo design of luciferases using deep learning\.Nature614\(7949\),pp\. 774–780\.External Links:[Document](https://dx.doi.org/10.1038/s41586-023-05696-3),[Link](https://www.nature.com/articles/s41586-023-05696-3)Cited by:[§3](https://arxiv.org/html/2609.30456#S3.p2.1)\.
- Yiet al\.\(2023\)K\. Yi, B\. Zhou, Y\. Shen, P\. Liò, and Y\. G\. WangGraph denoising diffusion for inverse protein folding\.External Links:2306\.16819,[Link](https://arxiv.org/abs/2306.16819)Cited by:[§1](https://arxiv.org/html/2609.30456#S1.p1.1),[§3](https://arxiv.org/html/2609.30456#S3.p1.1)\.

## Appendix AHyperparameter Settings

### A\.1Computational Resources

The experiments used NVIDIA GH200, H100, and L40S GPUs\. All timing experiments were run on GH200s\. The timing experiments also used 32 GB of RAM and 16 CPU cores\. The CPU parallelism is useful for speeding up the Gurobi solver used during sparse Fourier optimization\.

### A\.2Protein Design Methods

PretrainedWe use the ProteinMPNN discrete diffusion model that was developed in the original DRAKES paper\.

DRAKESThe KL\-penalty isβ=0\.001\\beta=0\.001\. Like for the pretrained model, we use the RL fine\-tuned ProteinMPNN model developed in the original DRAKES paper\.

Best\-of\-NThe performance of Best\-of\-N will increase monotonically asNNincreases, but with diminishing returns\. We use Spectral Feedback in conjunction with Best\-of\-N to improve alignment beyond these limitations\.

### A\.3Feedback Mechanisms

Unless otherwise written, we usen=64n=64samples for the value functions in equation[2](https://arxiv.org/html/2609.30456#S4.E2)\.

Feedback LoopWe selectk=20k=20edit\-positions andD=8192D=8192mask samples as our experimental parameters by following the results in the validation curves below\. The alignment reward continues to improve with more edit\-positions and more mask samples, though with diminishing returns\. IncreasingkkandDDalso increases the computational cost so we found our selected parameters to be a reasonable balance\. The mask sampling experiment does not use cross\-validation during edit\-selection, resulting in worse relative performance than the edit\-positions experiment which does use cross\-validation\. We used pre\-selected settings \(Max Depth=None, Leaves=50, Rate=0\.01,λ\\lambda=0\.0001\) because we faced compute resource limitations when increasingDDto 32768\. We used 3 feedback iterations in these experiments for similar reasons\.

Figure 7:Validation curves for feedback loop parameters\. Log\-scale shows diminishing returns\.Spectral Feedback \(ProxySPEX\)\. The ProxySPEX subroutine used in Spectral Feedback fits a GBT ensemble\. We used the validation protein set, optimizing for each hyperparameter in isolation\.

First\-order LASSO\. We use LASSO to implement a first\-order sparse recovery method\. In Spectral Feedback, we are given training setsSj⊆\[L\]​∀j∈\[D\]S\_\{j\}\\subseteq\[L\]\\forall j\\in\[D\]\. We define corresponding bit masksbjb\_\{j\}wherebj​\[i\]=1⟺i∈Sjb\_\{j\}\[i\]=1\\Longleftrightarrow i\\in S\_\{j\}\. A first\-order sparse recovery method models the value function as a linear functionf\(S\)=∑ici⋅𝟙\{i∈S\}f\(S\)=\\sum\_\{i\}c\_\{i\}\\cdot\\mathbbm\{1\}\\\{i\\in S\\\}wherec∈ℝLc\\in\\mathbb\{R\}^\{L\}\. We can then define the value function for the bit masks asf⁡\(b\)=cT​bf\(b\)=c^\{T\}b\. EachSjS\_\{j\}andbjb\_\{j\}have a corresponding training rewardrjr\_\{j\}\. LASSO is natural for this setting because we assume a sparse Fourier transform and theℓ1\\ell\_\{1\}norm induces sparsity\. Applying LASSO with the penalty coefficientλ∈ℝ\\lambda\\in\\mathbb\{R\}yields the following optimization problem:

c~=arg⁡minc∈ℝL​1D​∑j=1D\(cT​bj−rj\)2\+λ​‖c‖1\\tilde\{c\}=\\underset\{c\\in\\mathbb\{R\}^\{L\}\}\{\\arg\\min\}\\;\\frac\{1\}\{D\}\\sum\_\{j=1\}^\{D\}\(c^\{T\}b\_\{j\}\-r\_\{j\}\)^\{2\}\+\\lambda\|\|c\|\|\_\{1\}
We can then write the maximization problem off⁡\(b\)f\(b\)like so:

b∗=arg⁡maxb∈\{0,1\}L,∑ib⁡\[i\]≤k​c~T​bb^\{\*\}=\\underset\{b\\in\\\{0,1\\\}^\{L\},\\;\\sum\_\{i\}b\[i\]\\leq k\}\{\\arg\\max\}\\tilde\{c\}^\{T\}bThe solution is then the top\-kkcoefficients ofc~\\tilde\{c\}that are also positive:

b∗\[i\]=\{1​if​c~i\>0​and​c~i∈topkcoefficients0​elseb^\{\*\}\[i\]=\\left\\\{\\begin\{aligned\} &1\\text\{ if \}\\tilde\{c\}\_\{i\}\>0\\text\{ and \}\\tilde\{c\}\_\{i\}\\in\\text\{top $k$ coefficients\}\\\\ &0\\text\{ else \}\\end\{aligned\}\\right\.
Like for Spectral Feedback, when using LASSO for edit\-position selection, we run cross\-validation to be consistent with the original ProxySPEX paper\[[Butler et al\., 2025](https://arxiv.org/html/2609.30456#bib.bib11)\]\.

ParameterValidation Valuesλ\\lambda55values withλm​i​n=0\.00001,λm​a​x=0\.1\\lambda\_\{min\}=0\.00001,\\lambda\_\{max\}=0\.1andλ=0\.0\\lambda=0\.0Exclusion & Gradient\. The only hyperparameter for these edit\-position methods is the number of edit\-positions, which we already select for Spectral Feedback on the validation set\. We use the same number of positions \(20 in our experiments\) to make them fair baselines\.

## Appendix BDiscrete Diffusion Models

The discrete diffusion model is a type of generative language model that has been popularized for biology\-related tasks such as inverse folding\. We describe these models formally and explain how to sample from them\. First, consider a finite vocabularyVV, a sequence lengthLL, and a corresponding sample space𝒳:=VL\\mathcal\{X\}:=V^\{L\}\. We then model a corresponding probability mass trajectorypt∈ℝNp\_\{t\}\\in\\mathbb\{R\}^\{N\}whereN:=\|𝒳\|=\|V\|LN:=\|\\mathcal\{X\}\|=\|V\|^\{L\}\. Like in continuous diffusion, the discrete diffusion process involves evolvingptp\_\{t\}by a differential equation\. The evolution ofptp\_\{t\}forward in time is called the forward process\. In the continuous case, it is common to use the Fokker\-Planck equation, a partial differential equation related to the Itô diffusion process\. In the discrete setting, the following linear ordinary differential equation is used:

d​ptd​t=Qt​pt,p0∼pd​a​t​a\\displaystyle\\frac\{dp\_\{t\}\}\{dt\}=Q\_\{t\}p\_\{t\},\\quad p\_\{0\}\\sim p\_\{data\}Here,Qt∈ℝN×NQ\_\{t\}\\in\\mathbb\{R\}^\{N\\times N\}are the rate matrices of a Continuous Time Markov Chain \(CTMC\) and correspondingly must satisfy

Qt​\(i,j\)\\displaystyle Q\_\{t\}\(i,j\)≥0∀i≠j\\displaystyle\\geq 0\\quad\\forall i\\neq j∑iQt​\(i,j\)\\displaystyle\\sum\_\{i\}Q\_\{t\}\(i,j\)=0∀j\.\\displaystyle=0\\quad\\forall j\.The first constraint is due to the non\-negativity of probability masses, whereas the second constraint ensures that the total probability mass of the system does not change\. Like for continuous diffusion models, to make such a process useful for inference, we need a way to reverse the process\. In particular, if we structure the process to converge to a probability distributionpTp\_\{T\}from which we know how to sample, and then sample from the reverse process initialized bypTp\_\{T\}, we will converge to the target distributionp0p\_\{0\}\. The following is a commonly used reverse process\[[Liang et al\., 2025](https://arxiv.org/html/2609.30456#bib.bib14)\]:

d​pT−td​t=Q¯T−tpT−t,Q¯t\(x,y\)=\{pt​\(y\)pt​\(x\)​Qt​\(x,y\)x≠y−∑x′≠xQ¯t\(x′,x\)x=y\\displaystyle\\frac\{dp\_\{T\-t\}\}\{dt\}=\\bar\{Q\}\_\{T\-t\}p\_\{T\-t\},\\quad\\bar\{Q\}\_\{t\}\(x,y\)=\\left\\\{\\begin\{aligned\} \\frac\{p\_\{t\}\(y\)\}\{p\_\{t\}\(x\)\}Q\_\{t\}\(x,y\)&\\quad x\\neq y\\\\ \-\\sum\_\{x^\{\\prime\}\\neq x\}\\bar\{Q\}\_\{t\}\(x^\{\\prime\},x\)&\\quad x=y\\end\{aligned\}\\right\.Like before, the reverse process is a CTMC andQ¯t\\bar\{Q\}\_\{t\}are the corresponding rate matrices\. We can design the process to converge to a distribution of our choosing by selectingQtQ\_\{t\}accordingly\. Thus, to learnQ¯t\\bar\{Q\}\_\{t\}so that we can sample from the reverse process, we just need to learnpt​\(y\)/pt​\(x\)p\_\{t\}\(y\)/p\_\{t\}\(x\)\. This ratio is commonly called the score function\[[Sun et al\., 2023](https://arxiv.org/html/2609.30456#bib.bib13)\]\.

In its most abstract form, designing a discrete diffusion model requires defining a forward process with some generatorQtQ\_\{t\}and then learning a corresponding score functionpt​\(y\)/pt​\(x\)p\_\{t\}\(y\)/p\_\{t\}\(x\)from the training data\. A popular choice for the forward process is random masking\. The CTMC has some probability of transitioning tokens into masks, and once a token is masked, it stays masked\. The reverse process is initialized with a fully masked sequencexTx\_\{T\}, and then progressively flips tokens from masks to unmasked elements in the vocabulary\. Analogous to the forward process, a token will stay unmasked once it leaves its masked state\.

The Euler\-Maruyama scheme, a first\-order accurate iterative method, is commonly used to approximate the solution of continuous diffusion models\. It is then natural to apply Euler’s method to approximate the reverse discrete diffusion evolution:

pn\+1←pn\+h​Q¯T−tn​pn\\displaystyle p\_\{n\+1\}\\leftarrow p\_\{n\}\+h\\bar\{Q\}\_\{T\-t\_\{n\}\}p\_\{n\}wherehhis the step size\. This method is unfortunately impractical since it would require keeping track of the scores for an exponential number of state transitionsx→yx\\rightarrow y\. Instead, the following Hamming distance constraint is often used:Qt​\(x,y\)=0Q\_\{t\}\(x,y\)=0ifd⁡\(x,y\)\>1d\(x,y\)\>1\[[Liang et al\., 2025](https://arxiv.org/html/2609.30456#bib.bib14)\]\. This allows only state transitions that flip one token\. Still, to support multiple token updates during a single step for faster inference, the transitions of each token are assumed to be independent\.

## Appendix CCompute Complexity

The Spectral Feedback pipeline involves calls to several models with varying computational costs\. The protein diffusion model is called at each iteration of the reverse process\. In the first iteration, a protein sequence is generated from a fully masked sequence\. Subsequent iterations initialize the reverse process with partially noisy sequences via the feedback process\. Recall thatNNis the total number of generated proteins so the number of feedback iterations isN−1N\-1\. In each feedback iteration Spectral Feedback makesDDcalls to the mask value function\. LetMMbe the number of reverse process steps\. Then Spectral Feedback makesM×NM\\times Ncalls to the diffusion model andD×\(N−1\)D\\times\(N\-1\)calls to the mask value function\. The value functions in equation[2](https://arxiv.org/html/2609.30456#S4.E2)call the diffusion model once to get state change probabilities of a single reverse process step, then samplennstate changes and pass each generated sequence to the alignment reward oracle\. With these value functions, Spectral Feedback makesM×N\+D×\(N−1\)M\\times N\+D\\times\(N\-1\)diffusion model calls andn×D×\(N−1\)n\\times D\\times\(N\-1\)alignment reward oracle calls\. Best\-of\-NNis cheaper in compute since it makesM×NM\\times Ncalls to the diffusion model andNNcalls to the alignment reward oracle\. A key limitation of Best\-of\-N, however, is that it has diminishing returns asNNincreases and levels off in reward\. With protein design, desirable proteins may be rare as discussed in[Xiong et al\. \[2026\]](https://arxiv.org/html/2609.30456#bib.bib20)\. Figure[8](https://arxiv.org/html/2609.30456#A3.F8)demonstrates that Spectral Feedback can improve Best\-of\-N beyond its diminishing returns and identify these rare yet more stable proteins\. The modularity of the algorithm allows it to also improve other alignment methods such as Beam Search and RL as shown in Figure[12](https://arxiv.org/html/2609.30456#A6.F12)\.

Figure 8:Spectral Feedback with Best\-of\-10 reaches better performance than scaling Best\-of\-N\. Hyperparameters:D=8192D=8192,k=20k=20, and 5 feedback iterations\.
## Appendix DSparse Fourier Approximation

### D\.1Sparse Fourier Optimization

The derivation in this section originally appeared in Appendix A\.3 of[Butler et al\. \[2025\]](https://arxiv.org/html/2609.30456#bib.bib11)\. We restate it here\.

Assumef^:2\[n\]→ℝ\\hat\{f\}:2^\{\[n\]\}\\to\\mathbb\{R\}has a sparse, low\-degree Fourier expansion with support

𝒜=\{T⊆\[n\]:F^\(T\)≠0,\|T\|≤d\},\|𝒜\|≪2n,\\mathcal\{A\}\\;=\\;\\bigl\\\{\\,T\\subseteq\[n\]\\,:\\,\\hat\{F\}\(T\)\\neq 0,\\;\|T\|\\leq d\\,\\bigr\\\},\\qquad\|\\mathcal\{A\}\|\\ll 2^\{n\},such that

f^​\(S\)=∑T∈𝒜\(−1\)\|S∩T\|​F^​\(T\)\.\\hat\{f\}\(S\)\\;=\\;\\sum\_\{T\\in\\mathcal\{A\}\}\(\-1\)^\{\|S\\cap T\|\}\\,\\hat\{F\}\(T\)\.To formulate the optimization as an integer program, we re\-expressf^\\hat\{f\}in the Möbius basis, which replaces the parity functions\(−1\)\|S∩T\|∈\{−1,\+1\}\(\-1\)^\{\|S\\cap T\|\}\\in\\\{\-1,\+1\\\}with subset\-indicator functions𝟙\[T⊆S\]∈\{0,1\}\\mathbbm\{1\}\[T\\subseteq S\]\\in\\\{0,1\\\}\. Through a change of variable, the Möbius coefficients are

M^​\(T\)=\(−2\)\|T\|​∑S∈𝒜,S⊇TF^​\(S\)\.\\hat\{M\}\(T\)\\;=\\;\(\-2\)^\{\|T\|\}\\sum\_\{\\begin\{subarray\}\{c\}S\\in\\mathcal\{A\},\\\\ S\\supseteq T\\end\{subarray\}\}\\hat\{F\}\(S\)\.Letting𝒜\+=\{R⊆T\|T∈𝒜\}\\mathcal\{A\}^\{\+\}=\\bigl\\\{\\,R\\subseteq T\\,\\bigm\|\\,T\\in\\mathcal\{A\}\\,\\bigr\\\}denote the downward closure of𝒜\\mathcal\{A\}, the inverse Möbius transform gives

f^​\(S\)=∑R∈𝒜\+,R⊆SM^​\(R\)\.\\hat\{f\}\(S\)\\;=\\;\\sum\_\{\\begin\{subarray\}\{c\}R\\in\\mathcal\{A\}^\{\+\},\\\\ R\\subseteq S\\end\{subarray\}\}\\hat\{M\}\(R\)\.The optimization problem can then be expressed as a polynomial over\{0,1\}n\\\{0,1\\\}^\{n\}\. Let𝐱∈\{0,1\}n\\mathbf\{x\}\\in\\\{0,1\\\}^\{n\}be the characteristic vector ofSS, so thatxi=1x\_\{i\}=1if and only ifi∈Si\\in S\. We focus on the maximization problem \(minimization follows analogously\):

maxS⊆\[n\],\|S\|≤k⁡f^​\(S\)=max⁡∑R∈𝒜\+𝐱∈\{0,1\}n,∑ixi≤k⁡M^​\(R\)​∏i∈Rxi\.\\max\_\{\\begin\{subarray\}\{c\}S\\subseteq\[n\],\\\\ \|S\|\\leq k\\end\{subarray\}\}\\hat\{f\}\(S\)\\;=\\;\\max\_\{\\begin\{subarray\}\{c\}\\mathbf\{x\}\\in\\\{0,1\\\}^\{n\},\\\\ \\sum\_\{i\}x\_\{i\}\\leq k\\end\{subarray\}\}\\sum\_\{R\\in\\mathcal\{A\}^\{\+\}\}\\hat\{M\}\(R\)\\prod\_\{i\\in R\}x\_\{i\}\.
To reduce the problem to a linear integer program, each monomial∏i∈Rxi\\prod\_\{i\\in R\}x\_\{i\}is replaced with a binary decision variableyR∈\{0,1\}y\_\{R\}\\in\\\{0,1\\\}\. We augment𝒜\+\\mathcal\{A\}^\{\+\}to include all singletons\{i\}\\\{i\\\}fori∈\[n\]i\\in\[n\], settingM^​\(\{i\}\)=0\\hat\{M\}\(\\\{i\\\}\)=0for any newly added singletons, so thaty\{i\}y\_\{\\\{i\\\}\}exists for everyii\. Linking constraints enforceyR=∏i∈Rxiy\_\{R\}=\\prod\_\{i\\in R\}x\_\{i\}:

max𝐲∈\{0,1\}\|𝒜\+\|\\displaystyle\\max\_\{\\mathbf\{y\}\\in\\\{0,1\\\}^\{\|\\mathcal\{A\}^\{\+\}\|\}\}\\quad∑R∈𝒜\+M^​\(R\)​yR\\displaystyle\\sum\_\{R\\in\\mathcal\{A\}^\{\+\}\}\\hat\{M\}\(R\)\\,y\_\{R\}s\.t\.yR≤yQ\\displaystyle y\_\{R\}\\;\\leq\\;y\_\{Q\}∀Q⊂R,R∈𝒜\+\\displaystyle\\quad\\forall\\,Q\\subset R,\\;R\\in\\mathcal\{A\}^\{\+\}∑i∈Ry\{i\}≤\|R\|−1\+yR\\displaystyle\\sum\_\{i\\in R\}y\_\{\\\{i\\\}\}\\;\\leq\\;\|R\|\-1\+y\_\{R\}∀R∈𝒜\+\\displaystyle\\quad\\forall\\,R\\in\\mathcal\{A\}^\{\+\}∑i∈\[n\]y\{i\}≤k\.\\displaystyle\\sum\_\{i\\in\[n\]\}y\_\{\\\{i\\\}\}\\;\\leq\\;k\.The first constraint guarantees that whenever a monomial is activated \(i\.e\.,xi=1x\_\{i\}=1for alli∈Ri\\in R\), all of its subsets are also activated\. The second ensures that if a monomial is deactivated \(i\.e\.,xi=0x\_\{i\}=0for somei∈Ri\\in R\), then at least one of its constituent singletonsy\{i\}y\_\{\\\{i\\\}\}is likewise deactivated\. The third imposes the cardinality constraint, and after solving the program, the solution𝐱\\mathbf\{x\}is read off fromy\{i\}y\_\{\\\{i\\\}\}\.

Lets=\|𝒜\|s=\|\\mathcal\{A\}\|be the sparsity off^\\hat\{f\}\. The resulting program has at mosts⋅2d\+ns\\cdot 2^\{d\}\+ndecision variables: up to2d2^\{d\}subsets per element of𝒜\\mathcal\{A\}, plus thennaugmented singletons\. The constraints decompose as at most\(s⋅2d\+n\)​\(2d−1\)\(s\\cdot 2^\{d\}\+n\)\(2^\{d\}\-1\)subset\-activation constraints \(one per pairQ⊂RQ\\subset R\),s⋅2d\+ns\\cdot 2^\{d\}\+ndeactivation constraints, and a single cardinality constraint, giving at mosts⋅4d\+n⋅2d\+1s\\cdot 4^\{d\}\+n\\cdot 2^\{d\}\+1constraints in total\. The program is therefore tractable whenf^\\hat\{f\}is both sparse \(sssmall\) and low\-degree \(ddsmall\), even for largenn\.

We solve the program using Gurobi’s default branch\-and\-cut algorithm via thegurobipyinterface\[[Gurobi Optimization, LLC, 2026](https://arxiv.org/html/2609.30456#bib.bib30)\]\.

### D\.2Sparsity Experiments

Data in Figure[3](https://arxiv.org/html/2609.30456#S4.F3)is computed using SPEX with a compute budget of 100k and a max interaction order of 5\. The studied value functions usen=64n=64empirical samples for each evaluation\. The underlying protein model is the pretrained model\. We show additional spectral analysis results below for each protein in the test set, using both thefm​a​xf\_\{max\}and thefa​v​gf\_\{avg\}value functions\. We evaluate the sparsity of the functions as well as the necessity of higher\-order terms\. We first study theΔ​Δ​G\\Delta\\Delta Galignment oracle that was used in the paper’s main results\. We repeat this study for the ProtGPT2 Log\-Likelihood oracle\. In the main paper, we used ProtGPT2 to evaluate the naturalness of generated proteins; however, it can also be a target for alignment\. The results show that both value functions exhibit sparse Fourier transforms as they can be approximated with high faithfulness \(R2R^\{2\}\) by using relatively few Fourier coefficients\. Additionally, most coefficients are concentrated at lower interaction orders\.

Δ​Δ​G\\Delta\\Delta GAlignment Oracle

ProtGPT2 Log\-Likelihood Oracle

### D\.3High\-Sensitivity Value Function Ablation

High globalR2R^\{2\}is not by itself a sufficient certificate for argmax recovery\. One can easily construct counterexamples where an approximation has arbitrarily high globalR2R^\{2\}, yet selects a poor maximizer\. Given our strong experimental results using the global\-R2R^\{2\}\-driven approximation, we do not believe our value functions fall into this pathological regime\. Nevertheless, this is an important distinction\. In follow\-up analyses, we have been exploring objectives that are more directly aligned with maximization\. In particular, one can replace any value functionf⁡\(S\)f\(S\)with a monotone shaping

g⁡\(S\)=ϕ⁡\(f⁡\(S\)\),g\(S\)=\\phi\(f\(S\)\),which preserves the optimizer but makesL2L\_\{2\}approximation increasingly sensitive to high\-value subsets\. One concrete choice is an exponentially tilted objective,

gβ​\(S\)=exp⁡\(β​f⁡\(S\)−fmaxfmax−fmin\),g\_\{\\beta\}\(S\)=\\exp\\left\(\\beta\\frac\{f\(S\)\-f\_\{\\max\}\}\{f\_\{\\max\}\-f\_\{\\min\}\}\\right\),wherefmaxf\_\{\\max\}andfminf\_\{\\min\}denote the maximum and minimum values offfover the sampled set\. Under this normalization, the maximizer has value11, while lower\-value subsets are exponentially downweighted towardexp⁡\(−β\)\\exp\(\-\\beta\)\. Largerβ\\betatherefore places more emphasis on accurately approximating the top of the value function\.

This transformation has interesting implications for spectral sparsity\. Asβ→∞\\beta\\to\\infty,gβg\_\{\\beta\}approaches a Dirac delta at the maximizerS⋆S^\{\\star\}\. A Dirac delta on the Boolean cube is spectrally dense\. We have repeated our experiments above by analyzing spectral sparsity under different severities of exponential tilt\. We find that for modest choices ofβ\\beta\(0\.25\-8\), the sparsity levels of the function remain similar to the unnormalizedf⁡\(S\)f\(S\), and the proportion of spectral variance in higher\-degree terms remains high\. For larger values ofβ\\beta, the spectrum places more energy on higher\-order terms and becomes increasingly dense, making it difficult for spectral recovery techniques to achieve highR2R^\{2\}\. Thus, increasingβ\\betatrades better argmax alignment for weaker Fourier sparsity\. An improved version of Spectral Feedback could try learninggβg\_\{\\beta\}for a range ofβ\\betausing the same samples, and then cross\-validate to select the highestβ\\betafor a pre\-specified approximation quality\.

Figure 9:Exponential tilt function Fourier analysis for 2KRU with the pretrained model\.Figure 10:Exponential tilt function Fourier analysis for r6\_560\_TrROS\_Hall with pretrained model\.

## Appendix ESpectral FeedbackscRMSDDiscussion

Table[1](https://arxiv.org/html/2609.30456#S5.T1)shows an increase inscRMSDwhen Spectral Feedback is applied to DRAKES\. This is an undesired result since this means the inverse folding is less accurate\. Recall that one of the two conditions for a protein to be classified as successful is thats​c​R​M​S​D<2scRMSD<2\. An increase in the average then limits the gains that can be made in protein success rate, as seen in Figure[5](https://arxiv.org/html/2609.30456#S5.F5)\.

To better understand this behavior, we analyzed scatter and contour plots relatings​c​R​M​S​DscRMSDwith alignmentΔ​Δ​G\\Delta\\Delta G\. The results in the corresponding figure below indicate that Spectral Feedback inherits the properties of both the underlying model and the reward objective\. In the DRAKES distribution, the highest\-reward region contains structurally inconsistent samples, including sequences with scRMSD values between 8 and 10, well above the desired threshold of 2\. In contrast, the pretrained model has fewer of these high\-scRMSD samples, and its highest\-reward samples generally retain lower scRMSD\. As a result, the average scRMSD is greater for DRAKES than for the pretrained model: 1\.17 vs 1\.10\. This difference increases with the addition of Spectral Feedback as reported in the paper\. DRAKES also has a lower average log\-likelihood than the pretrained model, as seen in Table[1](https://arxiv.org/html/2609.30456#S5.T1), which further indicates inherent overfitting to the alignment oracle that leads to more out\-of\-distribution proteins\.

![Refer to caption](https://arxiv.org/html/2609.30456v1/scatter_contour_exps2_pretrained_spectral_20.png)

![Refer to caption](https://arxiv.org/html/2609.30456v1/scatter_contour_exps2_drakes_spectral_20.png)

Figure 11:scRMSDvsΔ​Δ​G\\Delta\\Delta Gdistributions show a worse tradeoff withscRMSDfor DRAKES when applying Spectral Feedback\.Spectral Feedback further optimizes the supplied stability reward and can, therefore, move samples toward regions where stability reward improves but structural consistency degrades\. However, this is a reflection of the DRAKES distribution in high\-reward regions\. We conclude that Spectral Feedback improves a specified alignment objective, while preservation of auxiliary properties depends on the quality of the reward and the distribution induced by the underlying model\. In settings such as DRAKES, a regularized or multi\-objective reward that explicitly includes structural consistency or correlated statistics would be needed to prevent this trade\-off\. We discuss using a multi\-objective reward in Appendix[F\.2](https://arxiv.org/html/2609.30456#A6.SS2)\.

## Appendix FAdditional Scaling and Alignment Studies

### F\.1Spectral Feedback Applications

Our experiments show that Spectral Feedback improves a broad set of alignment techniques\. In applications like protein design where computational resources or latency may not be an issue, algorithms that can improve an arbitrary method like Best\-of\-N or DRAKES in the face of diminishing returns would be very useful\. Feedback is an important technique to improve alignment methods beyond their limits and Spectral Feedback achieves this in various settings shown in Figure[12](https://arxiv.org/html/2609.30456#A6.F12)\.

Figure 12:Spectral Feedback withk=20k=20,D=8192D=8192, and 5 feedback iterations\.In general, feedback can use substantial compute resources when many samples are used to make edit\-selections; however, it is useful when a target alignment model reaches insufficient reward values\. Still, it is important to make each feedback step as efficient as possible and Figure[6](https://arxiv.org/html/2609.30456#S5.F6)shows that Spectral Feedback with ProxySPEX scales well in alignment performance and latency when compared to competing edit\-position selection methods\.

### F\.2Additional Alignment Oracles

The bulk of this paper is spent focusing on alignment of inverse protein folding diffusion models with aΔ​Δ​G\\Delta\\Delta Greward oracle, though it is important to understand the flexibility of Spectral Feedback across not just different protein models but also different reward oracles\. We study its performance with two additional alignment oracles: ProtGPT2 Log\-Likelihood and a multi\-objective reward that balances log\-likelihood withΔ​Δ​G\\Delta\\Delta G\.

ProtGPT2 Log\-Likelihood Oracle

The ProtGPT2 Log\-Likelihood oracle computes log\-likelihoods of diffusion model sample sequences by using the ProtGPT2 auto\-regressive language model\. This has been done in previous works such as[Gruver et al\. \[2023\]](https://arxiv.org/html/2609.30456#bib.bib5)and[Cemri et al\. \[2024\]](https://arxiv.org/html/2609.30456#bib.bib2)\. These log\-likelihoods can be used as a measure of naturalness of protein sequences because ProtGPT2 was trained to capture the distribution of protein amino acid sequences\[[Ferruz et al\., 2022](https://arxiv.org/html/2609.30456#bib.bib18)\]\. The figure below shows Spectral Feedback applied to the pretrained model with the ProtGPT2 oracle, using first\-order LASSO as the edit\-selection strategy\.

Figure 13:ProtGPT2 Log\-Likelihood alignment trajectory with Spectral Feedback\. Hyperparameters arek=10k=10,D=2048D=2048,n=16n=16\(reward oracle calls per value function call\)\.Balanced Oracle

We define the multi\-objective reward oracle asr⁡\(x\)=α⋅Δ​Δ​G​\(x\)\+0\.005⋅\(1−α\)⋅P​r​o​t​G​P​T​2​\(x\)r\(x\)=\\alpha\\cdot\\Delta\\Delta G\(x\)\+0\.005\\cdot\(1\-\\alpha\)\\cdot ProtGPT2\(x\)\. The weightα\\alphabalances between aligning forΔ​Δ​G\\Delta\\Delta Gand aligning for log\-likelihood\. The constant0\.0050\.005rescales the ProtGPT2 oracle to make the choice ofα\\alphamore linear for weighting the importance of the two oracles\. We show results for two proteins in the test set which demonstrate differing tradeoffs\. R6\-560 has a direct tradeoff between the oracles whereas 2KRU has higher log\-likelihoods when primarily aligning forΔ​Δ​G\\Delta\\Delta G\.

Figure 14:An example multi\-objective oracle balances betweenΔ​Δ​G\\Delta\\Delta Gand ProtGPT2 Log\-Likelihood\.These results show that a multi\-objective oracle can be useful for targeting and balancing between multiple objectives with Spectral Feedback\. This can be useful as a form of regularization to avoid overfitting, which is typical for alignment methods with a maximizing objective\. For example, one could make an oracle that balances betweenΔ​Δ​G\\Delta\\Delta Gands​c​R​M​S​DscRMSDto avoid the highs​c​R​M​S​DscRMSDproteins identified in Figure[11](https://arxiv.org/html/2609.30456#A5.F11)\.

## Appendix GHigh\-Order Interactions Case Study \(Δ​Δ​G\\Delta\\Delta GAlignment\)

### G\.1Comparing ProxySPEX with LASSO

In our work, we observed that the necessity of multivariate interactions for determining edit\-positions is highly dependent on the target protein\. While Spectral Feedback performs well when paired with ProxySPEX, a high\-order sparse recovery method such as ProxySPEX may not always be necessary\. We study how LASSO performs as a first\-order sparse recovery method to better understand the impact of high\-order sparse recovery\. We first compare using ProxySPEX and LASSO for Spectral Feedback with two protein backbones: R6\-560 and 2KVV\. We run a single Spectral Feedback process for each backbone and initialize each process with the same corresponding protein sequences that were studied in Figure[3](https://arxiv.org/html/2609.30456#S4.F3)\. We then compare ProxySPEX with LASSO across the entire test set\.

In the first experiment, ProxySPEX had an 8\.3% larger final reward than LASSO for2KVVand a 24% larger reward forR6\-560\. A possible explanation for these results lies in the structural differences between the protein backbones\. Figure[5](https://arxiv.org/html/2609.30456#S5.F5)shows that the2KVVbackbone contains notableα\\alpha\-helices while this structure is less prominent inR6\-560\. Anα\\alpha\-helix is a secondary protein structure that is formed by structured local interactions between theit​hi^\{th\}and\(i\+4\)t​h\(i\+4\)^\{th\}amino acids in a protein sequence\[[Gupte et al\., 2021](https://arxiv.org/html/2609.30456#bib.bib40)\]\. The primarily local interactions that create this structure may then lead to less complex interactions between edit\-positions\. In Figure[3](https://arxiv.org/html/2609.30456#S4.F3), forfa​v​gf\_\{avg\},2KVVhas strong first\-order interactions whileR6\-560has more significant high\-order interactions that ProxySPEX can leverage\. Differences in backbone structure may influence differences in the value function Fourier spectra, which then relate to the performance of ProxySPEX and LASSO for edit\-selection\. We show detailed results of the exact edit\-selections and amino\-acid updates for this experiment in Appendices[G\.2](https://arxiv.org/html/2609.30456#A7.SS2)and[G\.3](https://arxiv.org/html/2609.30456#A7.SS3)\.

In the second experiment, ProxySPEX slightly outperforms LASSO for the pretrained model and Best\-of\-10\. Additionally, LASSO slightly outperforms ProxySPEX for DRAKES\. The close performance may be because most proteins in the test set have dominant first\-order interactions for both value functions \(see Appendix[D\.2](https://arxiv.org/html/2609.30456#A4.SS2)\)\. It would be interesting to see if the corresponding spectral profiles of the value functions for DRAKES are also more dominantly first\-order, given the better performance of LASSO\. It would also be interesting to find cases where high\-order interactions are more dominant and where ProxySPEX would stand out more\. Perhaps better performance could be reached in cases where the value function is deterministic since then largerR2R^\{2\}values would be achievable and the sparse recovery from ProxySPEX would be more accurate\.

Figure 15:Spectral Feedback performance comparison with ProxySPEX vs first\-order LASSO as the edit selection method\. We use the pretrained model and constrain edit\-selections to at mostk=20k=20positions\.\(left\)Visualization of one held\-out feedback trajectory for the R6\-560 protein backbone\.\(right\)Average Spectral Feedback alignment across the test set, comparing ProxySPEX and LASSO as the edit\-selection strategies\.
### G\.2Target Protein: R6\-560 \(r6\_560\_TrROS\_Hall\)

Initial: SKPPKVVTVEVAVTKPDGKTELVKVTFTNLPRELKPGDTVTIPETGQKATVVKIIP\\boxed\{\\begin\{aligned\} &\\text\{Initial: SKPPKVVTVEVAVTKPDGKTELVKVTFTNLPRELKPGDTVTIPETGQKATVVKIIP\}\\\\ \\end\{aligned\}\}
ProxySPEX\-fa​v​gf\_\{avg\}

Iteration 1

Targets:4, 6,9, 12,15, 18, 19, 20, 24, 25,26, 28, 34, 40,41, 43, 48, 49,50,52

Changes: 4 \(K→\\rightarrowR\), 9 \(E→\\rightarrowV\), 15 \(P→\\rightarrowA\), 26 \(F→\\rightarrowL\), 41 \(I→\\rightarrowL\), 50 \(V→\\rightarrowI\), 52 \(K→\\rightarrowE\)

Result:SKPPRVVTVVVAVTKADGKTELVKVTLTNLPRELKPGDTVTLPETGQKATIVEIIP

Iteration 2

Targets:0,1,9, 13,14, 17, 18, 19, 25, 27,28, 32,34, 36,38, 43, 46,47,54

Changes: 0 \(S→\\rightarrowA\), 1 \(K→\\rightarrowP\), 9 \(V→\\rightarrowL\), 14 \(K→\\rightarrowR\), 28 \(N→\\rightarrowG\), 34 \(K→\\rightarrowR\), 38 \(T→\\rightarrowV\), 47 \(K→\\rightarrowE\), 54 \(I→\\rightarrowL\)

Result:APPPRVVTVLVAVTRADGKTELVKVTLTGLPRELRPGDVVTLPETGQEATIVEILP

Iteration 3

Targets: 0, 8, 11, 12, 17,18, 20,22,23, 28, 30, 36, 42, 43, 44, 45,47, 49, 55

Changes: 18 \(K→\\rightarrowR\), 22 \(V→\\rightarrowR\), 23 \(K→\\rightarrowR\), 47 \(E→\\rightarrowR\)

Result:APPPRVVTVLVAVTRADGRTELRRVTLTGLPRELRPGDVVTLPETGQRATIVEILP

Iteration 4

Targets: 0,7, 8, 13,14, 16, 19,20, 28, 29, 30, 34, 35, 38, 40, 45,46,47, 48, 54

Changes: 7 \(T→\\rightarrowE\), 14 \(R→\\rightarrowE\), 20 \(E→\\rightarrowV\), 46 \(Q→\\rightarrowE\), 47 \(R→\\rightarrowE\)

Result:APPPRVVEVLVAVTEADGRTVLRRVTLTGLPRELRPGDVVTLPETGEEATIVEILP

Iteration 5

Targets: 0, 1, 6, 10, 16, 17, 24, 26, 29, 30, 31, 35, 36, 39, 42,47, 51, 52, 54, 55

Changes: 47 \(E→\\rightarrowR\)

Result:APPPRVVEVLVAVTEADGRTVLRRVTLTGLPRELRPGDVVTLPETGERATIVEILP

Reward Trajectory: \[0\.0769, 0\.4864, 0\.5541, 0\.5915, 0\.6773, 0\.6843\]

LASSO\-fa​v​gf\_\{avg\}

Iteration 1

Targets:0, 1,4, 6, 11, 13, 15, 18, 24, 25,26, 27, 28, 29, 34,38,41, 48,50,52

Changes: 0 \(S→\\rightarrowA\), 4 \(K→\\rightarrowR\), 26 \(F→\\rightarrowL\), 38 \(T→\\rightarrowV\), 41 \(I→\\rightarrowL\), 50 \(V→\\rightarrowI\), 52 \(K→\\rightarrowR\)

Result:AKPPRVVTVEVAVTKPDGKTELVKVTLTNLPRELKPGDVVTLPETGQKATIVRIIP

Iteration 2

Targets:1, 2, 6, 10, 12, 13,15, 18,23, 24,27,28, 29, 34, 35, 36, 43, 44, 48

Changes: 1 \(K→\\rightarrowP\), 15 \(P→\\rightarrowA\), 23 \(K→\\rightarrowT\), 27 \(T→\\rightarrowE\), 28 \(N→\\rightarrowG\)

Result:APPPRVVTVEVAVTKADGKTELVTVTLEGLPRELKPGDVVTLPETGQKATIVRIIP

Iteration 3

Targets: 6, 12,14, 17, 18, 24,27,28,29,34, 35, 36, 43, 44, 48, 52

Changes: 14 \(K→\\rightarrowR\), 27 \(E→\\rightarrowR\), 28 \(G→\\rightarrowD\), 29 \(L→\\rightarrowT\), 34 \(K→\\rightarrowR\)

Result:APPPRVVTVEVAVTRADGKTELVTVTLRDTPRELRPGDVVTLPETGQKATIVRIIP

Iteration 4

Targets:1, 2, 5, 6, 8, 11, 12, 13, 18, 22, 24,27,29,34, 36, 43, 44, 48, 52, 55

Changes: 1 \(P→\\rightarrowA\), 27 \(R→\\rightarrowT\), 29 \(T→\\rightarrowL\), 34 \(R→\\rightarrowK\)

Result:AAPPRVVTVEVAVTRADGKTELVTVTLTDLPRELKPGDVVTLPETGQKATIVRIIP

Iteration 5

Targets: 3, 6, 10, 12, 13, 17, 24, 29, 30,34, 36, 43, 44, 45, 48, 49

Changes: 34 \(K→\\rightarrowR\)

Result:AAPPRVVTVEVAVTRADGKTELVTVTLTDLPRELRPGDVVTLPETGQKATIVRIIP

Reward Trajectory: \[0\.0679, 0\.4345, 0\.4868, 0\.4770, 0\.5528, 0\.5518\]

### G\.3Target Protein: 2KVV \(v2K43S\_2KVV\)

Initial: EKWIEQNELMKETGLKRSTITKLRKTKLKEGEHYKRVSKDGKPSKDATILYNLEKIKKLLK\\boxed\{\\begin\{aligned\} &\\text\{Initial: EKWIEQNELMKETGLKRSTITKLRKTKLKEGEHYKRVSKDGKPSKDATILYNLEKIKKLLK\}\\\\ \\end\{aligned\}\}
ProxySPEX\-fa​v​gf\_\{avg\}

Iteration 1

Targets:1,5, 6, 10, 19, 26, 28, 31,34, 38, 41, 43,44, 50, 53, 54, 56,57, 59, 60

Changes: 1 \(K→\\rightarrowR\), 5 \(Q→\\rightarrowE\), 34 \(K→\\rightarrowR\), 44 \(K→\\rightarrowP\), 57 \(K→\\rightarrowE\)

Result:ERWIEENELMKETGLKRSTITKLRKTKLKEGEHYRRVSKDGKPSPDATILYNLEKIKELLK

Iteration 2

Targets: 3, 6, 9, 10, 11, 15, 16,21, 24, 26, 28, 31, 34, 41,42, 48, 54, 55,56, 60

Changes: 21 \(K→\\rightarrowR\), 42 \(P→\\rightarrowD\), 56 \(K→\\rightarrowL\)

Result:ERWIEENELMKETGLKRSTITRLRKTKLKEGEHYRRVSKDGKDSPDATILYNLEKILELLK

Iteration 3

Targets: 3,6, 9,10,11,15, 16, 17,24,26,28,31, 40,41, 48, 51, 53,54, 55,60

Changes: 6 \(N→\\rightarrowR\), 10 \(K→\\rightarrowA\), 11 \(E→\\rightarrowA\), 15 \(K→\\rightarrowA\), 24 \(K→\\rightarrowR\), 26 \(K→\\rightarrowR\), 28 \(K→\\rightarrowE\), 31 \(E→\\rightarrowR\), 41 \(K→\\rightarrowR\), 54 \(K→\\rightarrowA\), 60 \(K→\\rightarrowA\)

Result:ERWIEERELMAATGLARSTITRLRRTRLEEGRHYRRVSKDGRDSPDATILYNLEAILELLA

Iteration 4

Targets: 1, 2, 8, 9, 14,15, 19,20, 21, 25,26, 35, 36,38,47, 51, 52, 55,57, 58

Changes: 15 \(A→\\rightarrowR\), 20 \(T→\\rightarrowA\), 26 \(R→\\rightarrowA\), 38 \(K→\\rightarrowA\), 47 \(T→\\rightarrowR\), 57 \(E→\\rightarrowA\)

Result:ERWIEERELMAATGLRRSTIARLRRTALEEGRHYRRVSADGRDSPDARILYNLEAILALLA

Iteration 5

Targets: 0, 2, 8, 13, 14,15, 16,19,21, 24, 25, 29, 36, 39, 43, 49, 52, 53, 58, 60

Changes: 15 \(R→\\rightarrowA\), 19 \(I→\\rightarrowL\), 21 \(R→\\rightarrowA\)

Result:ERWIEERELMAATGLARSTLAALRRTALEEGRHYRRVSADGRDSPDARILYNLEAILALLA

Reward Trajectory: \[\-0\.1766, 0\.3299, 0\.3596, 0\.5910, 0\.6261, 0\.6717\]

LASSO\-fa​v​gf\_\{avg\}

Iteration 1

Targets:5, 6, 10, 12, 19, 25, 28, 29, 31,34, 38, 40, 41, 44, 50, 54, 55, 56,57, 60

Changes: 5 \(Q→\\rightarrowE\), 34 \(K→\\rightarrowR\), 57 \(K→\\rightarrowE\)

Result:EKWIEENELMKETGLKRSTITKLRKTKLKEGEHYRRVSKDGKPSKDATILYNLEKIKELLK

Iteration 2

Targets:1, 6, 9, 10, 15, 17, 19, 24, 26, 28, 30,31, 34,38,44, 54, 55, 56, 57, 60

Changes: 1 \(K→\\rightarrowR\), 31 \(E→\\rightarrowI\), 38 \(K→\\rightarrowA\), 44 \(K→\\rightarrowP\)

Result:ERWIEENELMKETGLKRSTITKLRKTKLKEGIHYRRVSADGKPSPDATILYNLEKIKELLK

Iteration 3

Targets: 3,6, 9, 10,15, 19,21,24, 26, 28, 29, 39, 43, 44,54, 55, 56,60

Changes: 6 \(N→\\rightarrowR\), 15 \(K→\\rightarrowA\), 21 \(K→\\rightarrowR\), 24 \(K→\\rightarrowR\), 54 \(K→\\rightarrowA\), 60 \(K→\\rightarrowA\)

Result:ERWIEERELMKETGLARSTITRLRRTKLKEGIHYRRVSADGKPSPDATILYNLEAIKELLA

Iteration 4

Targets: 0, 3, 9, 10, 17, 19, 22,26, 28, 29,42, 44, 55,57

Changes: 26 \(K→\\rightarrowR\), 42 \(P→\\rightarrowD\), 57 \(E→\\rightarrowA\)

Result:ERWIEERELMKETGLARSTITRLRRTRLKEGIHYRRVSADGKDSPDATILYNLEAIKALLA

Iteration 5

Targets: 0, 3, 9,10, 17, 19, 28, 39, 42, 45, 55, 60

Changes: 10 \(K→\\rightarrowR\)

Result:ERWIEERELMRETGLARSTITRLRRTRLKEGIHYRRVSADGKDSPDATILYNLEAIKALLA

Reward Trajectory: \[\-0\.1766, 0\.1106, 0\.3755, 0\.5710, 0\.6129, 0\.6203\]

Similar Articles

Spectral Prior for Reducing Exposure Bias in Diffusion Models

Hugging Face Daily Papers

This paper proposes Spectral Alignment (SPA), a lightweight guidance-based method that reduces exposure bias in diffusion models by calibrating the power spectrum of intermediate predictions, showing consistent improvements across pixel-space, latent, and flow-matching models with minimal computational overhead.

Learning to Discretize: Diffusion-Based Adaptive Mesh with Spectral Guidance

arXiv cs.LG

This paper proposes a diffusion-based framework for learning adaptive mesh discretization conditioned on observed PDE dynamics, using spectral guidance and physics constraints to allocate resolution where needed. The method achieves competitive or superior performance across five PDE regimes.