Variable-Length Generative Protein Design via Generalized Poisson Flow

arXiv cs.LG Papers

Summary

Introduces Generalized Poisson Flow (GPFlow), a variable-length generative framework for protein design that learns an inhomogeneous generalized Poisson process, enabling flexible length exploration and improving designability across structure, sequence, and peptide co-design tasks.

arXiv:2607.09039v1 Announce Type: new Abstract: The ability to generate variable-length proteins is crucial in protein design, where the optimal length is often unknown and tightly coupled to designability. Current diffusion- and flow-based generative models typically require the protein length to be specified before sampling, limiting their flexibility in exploring the feasible design space. To address this limitation, we introduce Generalized Poisson Flow (GPFlow), a variable-length generative framework that learns the rate function of an inhomogeneous generalized Poisson process by minimizing its negative log-likelihood. We establish population-level guarantees for recovering the joint multimodal distribution and derive an upper bound on the KL divergence between the data and generated distributions. We comprehensively evaluate GPFlow across structure and sequence design, motif scaffolding, and peptide co-design, spanning Euclidean, categorical, and Riemannian modalities to fully validate its variable-length generation quality. In unconditional design, GPFlow improves structural designability and achieves the best distributional fitness for sequence design compared to their corresponding fixed-length baselines, while perfectly recovering the length distribution. In conditional motif scaffolding, GPFlow ranks first on 10 of 16 structure-based design tasks with significantly more unique successes and also achieves more passed tasks in sequence-based design. In peptide co-design, GPFlow remains competitive even without access to a native-length oracle.
Original Article
View Cached Full Text

Cached at: 07/13/26, 07:57 AM

# Variable-Length Generative Protein Design via Generalized Poisson Flow
Source: [https://arxiv.org/html/2607.09039](https://arxiv.org/html/2607.09039)
\\useunder

\\ul\\newkeytheoremtheorem\[name=Theorem, numberwithin=section\]\\newkeytheoremproposition\[name=Proposition, numberlike=theorem\]

Chaoran Cheng∗, Zhanghan Ni∗, Yanru Qu∗, Yuxin Chen, Ruihan Guo, Jiajun Fan, Ge Liu Siebel School of Computing and Data Science University of Illinois Urbana\-Champaign

###### Abstract

The ability to generate variable\-length proteins is crucial in protein design, where the optimal length is often unknown and tightly coupled to designability\. Current diffusion\- and flow\-based generative models typically require the protein length to be specified before sampling, limiting their flexibility in exploring the feasible design space\. To address this limitation, we introduceGeneralized Poisson Flow \(GPFlow\), a variable\-length generative framework that learns the rate function of an inhomogeneous generalized Poisson process by minimizing its negative log\-likelihood\. We establish population\-level guarantees for recovering the joint multimodal distribution and derive an upper bound on the KL divergence between the data and generated distributions\. We comprehensively evaluate GPFlow across structure and sequence design, motif scaffolding, and peptide co\-design, spanning Euclidean, categorical, and Riemannian modalities to fully validate its variable\-length generation quality\. In unconditional design, GPFlow improves structural designability and achieves the best distributional fitness for sequence design compared to their corresponding fixed\-length baselines, while perfectly recovering the length distribution\. In conditional motif scaffolding, GPFlow ranks first on 10 of 16 structure\-based design tasks with significantly more unique successes and also achieves more passed tasks in sequence\-based design\. In peptide co\-design, GPFlow remains competitive even without access to a native\-length oracle\.

11footnotetext:Equal contribution\.††footnotetext:Correspondence tochaoran7@illinois\.edu\.![[Uncaptioned image]](https://arxiv.org/html/2607.09039v1/x1.png)

## 1Introduction

Diffusion models\[[22](https://arxiv.org/html/2607.09039#bib.bib22),[42](https://arxiv.org/html/2607.09039#bib.bib42)\]and flow matching\[[34](https://arxiv.org/html/2607.09039#bib.bib34)\]have become widely used frameworks for protein generative modeling and have produced strong results in protein design\[[40](https://arxiv.org/html/2607.09039#bib.bib40),[18](https://arxiv.org/html/2607.09039#bib.bib18),[31](https://arxiv.org/html/2607.09039#bib.bib31)\]\. However, their standard formulations operate on fixed\-dimensional state spaces\. Such a constraint is restrictive for tasks like motif scaffolding, where the optimal length is often unknown*a priori*and is tightly coupled to designability\. Sampling sweep over a pre\-specified length range increases inference cost and may miss feasible designs when the selected lengths are incompatible with the conditional constraints\.

To bridge this gap, we introduceGeneralized Poisson Flow \(GPFlow\), a variable\-length generative framework in which length evolves under an inhomogeneous generalized Poisson process\. GPFlow couples this length process to a within\-length generator, thereby accommodating continuous, discrete, Riemannian, and mixed length\-dependent modalities within the same construction\. Building on a marginalization theorem for the conditional process, we derive a tractable objective from the negative log\-likelihood \(NLL\) of the stochastic process and further extend the construction to multimodal generation with guarantees on the joint distribution recovery\. Furthermore, we derive a generator\-based KL upper bound whose terms align with the additive losses, thereby complementing the population\-level recovery guarantee at the optimal marginal generator\. While the final rate objective aligns with existing work, our formulation provides a common probability\-flow construction for continuous, discrete, Riemannian, and mixed modalities, with deeper theoretical connections\.

We consider five diverse protein design scenarios: \(a\)unconditional structure design\(Section[5\.1](https://arxiv.org/html/2607.09039#S5.SS1)\), \(b\)unconditional sequence design\(Section[5\.2](https://arxiv.org/html/2607.09039#S5.SS2)\), \(c\)structure\-based motif scaffolding\(Section[5\.3](https://arxiv.org/html/2607.09039#S5.SS3)\), \(d\)sequence\-based motif scaffolding\(Appendix[E\.4](https://arxiv.org/html/2607.09039#A5.SS4)\), and \(e\)peptide structure\-sequence co\-design\(Section[5\.4](https://arxiv.org/html/2607.09039#S5.SS4)\)\. This task landscape evaluates how the same variable\-length construction adapts across modalities and generative goals\. By retaining the corresponding fixed\-length base architectures, we obtain direct comparisons with their variable\-length extensions\. Across these tasks, our experiments provide evidence that dynamic\-length modeling can improve corresponding fixed\-length systems in the evaluated settings while removing the need to specify a target length\. For unconditional generation, our small structure model significantly improves designability over the corresponding Proteina base model, while the 642M sequence model more closely matches the UniRef50 statistics and maintains higher diversity than the DPLM base model\. On the motif\-scaffolding benchmark, the structure model ranks first on 10 of 16 tasks and yields more than a 10\-fold increase in unique successes on several challenging targets, while the sequence model also achieves more passed tasks\. In peptide co\-design, where translation, rotation, residue types, and torsion angles are learned jointly with insertion, GPFlow remains competitive and outperforms the base PepFlow model even without access to a native\-length oracle\. Our contributions are summarized as follows:

- •A unified variable\-dimensional probability flow\.We formulate the variable\-length generative modeling in GPFlow with inhomogeneous generalized Poisson probability paths and couple them with continuous, discrete, Riemannian, or mixed within\-length dynamics\. A single marginalization principle covers these modality types for multimodal generation\.
- •Simulation\-free likelihood learning with theoretical guarantees\.We derive the rate objective from the generalized Poisson trajectory likelihood while analytically marginalizing its auxiliary insertion times, so training is fully simulation\-free\. We further prove population\-level recovery of the joint multimodal distribution and establish a generator\-based KL upper bound\.
- •Variable\-length generation across five protein\-design settings\.We instantiate the framework for protein structure design, sequence design, motif scaffolding, and peptide structure\-sequence co\-design\. Across all tasks, direct comparisons with the fixed\-length base model demonstrate GPFlow’s superior performance and flexibility\. We achieve superior structural designability, better sequence fitness, more unique successes and passed tasks in motif scaffolding, and stable peptide designs without a native\-length oracle\.

## 2Related Work

##### Protein Design Models

Building on diffusion and flow matching, recent generative models for protein design have achieved strong performance\. FrameDiff\[[51](https://arxiv.org/html/2607.09039#bib.bib51)\]adopts Riemannian diffusion to learn the positions and orientations of each amino acid, and FrameFlow\[[50](https://arxiv.org/html/2607.09039#bib.bib50)\]further extends it with Riemannian flow matching\. Proteina\[[18](https://arxiv.org/html/2607.09039#bib.bib18)\]is a CA\-only flow\-based structure generative model that achieves remarkable scaling results\. RFDiffusion\[[49](https://arxiv.org/html/2607.09039#bib.bib49),[8](https://arxiv.org/html/2607.09039#bib.bib8)\]is a family of diffusion\-based models that incorporates the structure\-prediction capabilities of RoseTTAFold\[[3](https://arxiv.org/html/2607.09039#bib.bib3)\]into generative modeling\. Generally, these structure generative models also support motif scaffolding, in which fragments of functionally important protein structures, known as*motifs*, are provided to the model as the conditions\. There are also dedicated models for motif scaffolding, e\.g\., Genie2\[[31](https://arxiv.org/html/2607.09039#bib.bib31)\]is a diffusion\-based generative model with a specialized multi\-motif framework that designs co\-occurring motifs\. These models either adopt an inpainting\-style sampling that enforces hard constraints by fixing the motif residues and structures, or inject motif information as an explicit but soft conditioning signal\. Additional design models for sequences and peptides are available in Appendix[C](https://arxiv.org/html/2607.09039#A3)\.

##### Variable\-Length Generative Models

Standard diffusion or flow matching models require a predetermined sampling length, limiting their applicability to variable\-length data\. Non\-diffusion approaches primarily rely on autoregressive modeling\[[47](https://arxiv.org/html/2607.09039#bib.bib47)\]\. Two existing works incorporate data length through related stochastic\-process constructions\. TDDM\[[9](https://arxiv.org/html/2607.09039#bib.bib9)\]models dimension changes with a learnable jump kernel coupled to continuous diffusion dynamics, but does not treat a categorical length\-dependent modality\. EditFlow\[[20](https://arxiv.org/html/2607.09039#bib.bib20)\]instead derives insertion rates from discrete flow matching\[[17](https://arxiv.org/html/2607.09039#bib.bib17)\]; it applies to categorical modalities but not to continuous coordinates\. Importantly, neither of these models has been applied to the protein design domain\. More recently, SCISOR\[[4](https://arxiv.org/html/2607.09039#bib.bib4)\]approaches variable\-length protein sequence generation via a deletion\-based process, where random residues are inserted during forward noising and removed by the learned reverse process\.

GPFlow differs primarily in scope and probabilistic interpretation\. It uses one unified variable\-dimensional probability\-flow formulation for continuous, discrete, Riemannian, and mixed modalities\. Our rate objective follows from the generalized Poisson trajectory likelihood, in which auxiliary insertion times are analytically marginalized to yield a simulation\-free objective\. In their overlapping regimes, GPFlow recovers the insertion\-only, modality\-specific instantiations of TDDM and EditFlow and yields the same rate objective\. Our contribution is therefore the unified construction and likelihood interpretation with the new KL bound on the training dynamics\. While the original TDDM and EditFlow papers do not evaluate protein design tasks, our experiments instantiate the framework across five comprehensive settings for unconditional and conditional protein design\.

## 3Generalized Poisson Flow

To construct a probability path\{pt\}t∈\[0,1\]\\\{p\_\{t\}\\\}\_\{t\\in\[0,1\]\}for a stochastic stateYtY\_\{t\}using flow matching \(or diffusion\), one needs \(a\)p1≈qp\_\{1\}\\approx q, whereqqis the data distribution, \(b\) simulation\-free sampling fromptp\_\{t\}for feasible training, and \(c\) a tractable loss for learning the marginal flow\. We first derive theoretical results for the length\-only stateYt=Xt∈ℕY\_\{t\}=X\_\{t\}\\in\\mathbb\{N\}and then extend them to the multimodal stateYt=\(Xt,St\)∈𝖸Y\_\{t\}=\(X\_\{t\},S\_\{t\}\)\\in\\mathsf\{Y\}, whereStS\_\{t\}may comprise arbitrary discrete or continuous length\-dependent modalities or their mixes\. We demonstrate how to construct feasible probability paths using the generalized Poisson process\.

### 3\.1Length\-Only Generalized Poisson Flow

A*generalized Poisson process*\{Xt\}0≤t≤T\\\{X\_\{t\}\\\}\_\{0\\leq t\\leq T\}is a continuous\-time Markov process defined by a*rate function*λt:\[0,T\]×ℕ→ℝ≥0\\lambda\_\{t\}:\[0,T\]\\times\\mathbb\{N\}\\to\\mathbb\{R\}^\{\\geq 0\}, such thatPr⁡\(Xt\+h=Xt\+1\)=λt​\(Xt\)​h\+o​\(h\)\\Pr\(X\_\{t\+h\}=X\_\{t\}\+1\)=\\lambda\_\{t\}\(X\_\{t\}\)h\+o\(h\)andPr⁡\(Xt\+h=Xt\)=1−λt​\(Xt\)​h\+o​\(h\)\\Pr\(X\_\{t\+h\}=X\_\{t\}\)=1\-\\lambda\_\{t\}\(X\_\{t\}\)h\+o\(h\)\. Consistent with the flow matching convention, we will assume the generalized Poisson process to be defined on the time interval\[0,1\]\[0,1\]withX0=0X\_\{0\}=0\. We now show one feasible construction of the probability pathpt​\(X\)p\_\{t\}\(X\)using the generalized Poisson process with the following theorem:

\{theorem\}

\[store=forwardrate\] Letκ:\[0,1\]→\[0,1\]\\kappa:\[0,1\]\\to\[0,1\]be a nondecreasing, absolutely continuous scheduler withκ0=0\\kappa\_\{0\}=0andκ1=1\\kappa\_\{1\}=1, and define the*completion time*τκ:=inf\{t:κt=1\}\\tau\_\{\\kappa\}:=\\inf\\\{t:\\kappa\_\{t\}=1\\\}\. For any fixedL≥1L\\geq 1, assumeκt<1\\kappa\_\{t\}<1fort<τκt<\\tau\_\{\\kappa\}and define the insertion*hazard*and*conditional rate*by

ht:=\{κ˙t1−κt,t<τκ,0,t≥τκ,λt​\(Xt\|L\):=\(L−Xt\)​ht\.h\_\{t\}:=\\begin\{cases\}\\dfrac\{\\dot\{\\kappa\}\_\{t\}\}\{1\-\\kappa\_\{t\}\},&t<\\tau\_\{\\kappa\},\\\\ 0,&t\\geq\\tau\_\{\\kappa\},\\end\{cases\}\\qquad\\lambda\_\{t\}\(X\_\{t\}\|L\):=\(L\-X\_\{t\}\)h\_\{t\}\.\(1\)Then, the resulting process satisfiesXt∼B​\(L,κt\)X\_\{t\}\\sim B\(L,\\kappa\_\{t\}\)fort<τκt<\\tau\_\{\\kappa\}and is absorbed atXt=LX\_\{t\}=Lfort≥τκt\\geq\\tau\_\{\\kappa\}\. See the proof in Appendix[A\.1](https://arxiv.org/html/2607.09039#A1.SS1)\. Theorem[3\.1](https://arxiv.org/html/2607.09039#S3.SS1)offers a simulation\-free approach for samplingXtX\_\{t\}and guaranteesX1=LX\_\{1\}=Lalmost surely\. Similar to standard flow matching for continuous modalities, we then establish the following marginalization theorem:\{theorem\}\[store=marginal\] For any schedulerκt\\kappa\_\{t\}defined above, letqqbe a target length distribution, definept​\(k\|y\):=Pr⁡\(Xt=k\|X1=y\)p\_\{t\}\(k\|y\):=\\Pr\(X\_\{t\}=k\|X\_\{1\}=y\), and letpt​\(k\):=∑ypt​\(k\|y\)​q​\(y\)p\_\{t\}\(k\):=\\sum\_\{y\}p\_\{t\}\(k\|y\)q\(y\)\. For everykkwithpt​\(k\)\>0p\_\{t\}\(k\)\>0, the posterior probability isp​\(X1=y\|Xt=k\):=pt​\(k\|y\)​q​\(y\)/pt​\(k\)p\(X\_\{1\}=y\|X\_\{t\}=k\):=\{p\_\{t\}\(k\|y\)q\(y\)\}/\{p\_\{t\}\(k\)\}\. Then the*marginal rate*

λt∗​\(k\):=𝔼X1∼p\(⋅\|Xt=k\)​\[λt​\(k\|X1\)\]\\lambda\_\{t\}^\{\*\}\(k\):=\\mathbb\{E\}\_\{X\_\{1\}\\sim p\(\\cdot\|X\_\{t\}=k\)\}\[\\lambda\_\{t\}\(k\|X\_\{1\}\)\]\(2\)generates the marginal pathptp\_\{t\}and hencep1=qp\_\{1\}=q\.

See Appendix[A\.2](https://arxiv.org/html/2607.09039#A1.SS2)for the proof, which relies primarily on the Kolmogorov forward equation as the continuity equation\. Theorem[1](https://arxiv.org/html/2607.09039#S3.E1)guarantees that any length distributionqqcan be modeled using a generalized Poisson process and also ensures the data constraintp1=qp\_\{1\}=qfor constructing a probability flowptp\_\{t\}over the discrete length space\. We therefore refer to our generative framework as*Generalized Poisson Flow*\. To effectively learn the flow, in Proposition[A\.3](https://arxiv.org/html/2607.09039#A1.SS3), we prove that the negative log\-likelihood of a realization of a generalized Poisson process can be calculated asℒNLL=∫01λt​d​t−∑k=1Llog⁡λtk\\mathcal\{L\}\_\{\\text\{NLL\}\}=\\int\_\{0\}^\{1\}\\lambda\_\{t\}\\mathop\{\}\\\!\\mathrm\{d\}t\-\\sum\_\{k=1\}^\{L\}\\log\\lambda\_\{t\_\{k\}\}, where the two terms can be viewed as the NLL for the survival and event terms, respectively\. The insertion times are auxiliary variables used to construct the probability path, and by the integration formula for a stochastic point process\[[7](https://arxiv.org/html/2607.09039#bib.bib7)\], the expected sum over insertion times can be marginalized into an integral against the ground\-truth marginal rateλt∗\\lambda^\{\*\}\_\{t\}\. Combining the marginalization result in Theorem[1](https://arxiv.org/html/2607.09039#S3.E1)and the conditional rate in Eq\.[1](https://arxiv.org/html/2607.09039#S3.E1), we have:

𝔼\{tk\}k=1L,L​\[−∑k=1Llog⁡λtk\]=𝔼t​\[−λt∗​log⁡λt\]=𝔼t,X1∼q​\(X\)​\[−ht​\(X1−Xt\)​log⁡λt\]\.\\displaystyle\\mathbb\{E\}\_\{\\\{t\_\{k\}\\\}\_\{k=1\}^\{L\},L\}\\left\[\-\\sum\_\{k=1\}^\{L\}\\log\\lambda\_\{t\_\{k\}\}\\right\]=\\mathbb\{E\}\_\{t\}\[\-\\lambda^\{\*\}\_\{t\}\\log\\lambda\_\{t\}\]=\\mathbb\{E\}\_\{t,X\_\{1\}\\sim q\(X\)\}\\left\[\-h\_\{t\}\(X\_\{1\}\-X\_\{t\}\)\\log\\lambda\_\{t\}\\right\]\.\(3\)Therefore, we arrive at a tractable, simulation\-free NLL loss that requires only an endpointX1X\_\{1\}, a sampled timett, and the corresponding intermediate stateXtX\_\{t\}obtained from Theorem[3\.1](https://arxiv.org/html/2607.09039#S3.SS1):

ℒGP=𝔼t,X1∼q​\(X\)​\[λt−ht​\(X1−Xt\)​log⁡λt\]\.\\mathcal\{L\}\_\{\\text\{GP\}\}=\\mathbb\{E\}\_\{t,X\_\{1\}\\sim q\(X\)\}\\left\[\\lambda\_\{t\}\-h\_\{t\}\(X\_\{1\}\-X\_\{t\}\)\\log\\lambda\_\{t\}\\right\]\.\(4\)We note that, unlike continuous flow matching, in which the ground\-truth marginal vector field is intractable, Theorem[3\.1](https://arxiv.org/html/2607.09039#S3.SS1)provides a closed\-form conditional rate target\.

### 3\.2Multimodal Generalized Poisson Flow

To truly unlock variable\-length protein structure and sequence design, we now extend the length\-only formulation to incorporate other length\-dependent modalities such as continuous coordinates or discrete types\. Let𝖲\\mathsf\{S\}be a component space and define the variable\-length state space𝖸:=⨆k≥0\{k\}×𝖲k\\mathsf\{Y\}:=\\bigsqcup\_\{k\\geq 0\}\\\{k\\\}\\times\\mathsf\{S\}^\{k\}\. The component space may be continuous \(with a Lebesgue or manifold\-volume reference measure\), discrete \(with counting measure\), or a product of such spaces\. We writeYt=\(Xt,St\)∈𝖸Y\_\{t\}=\(X\_\{t\},S\_\{t\}\)\\in\\mathsf\{Y\}, whereSt=\(St1,…,StXt\)S\_\{t\}=\(S\_\{t\}^\{1\},\\dots,S\_\{t\}^\{X\_\{t\}\}\)\. An*insertion*of a component can be described by a rateλt\\lambda\_\{t\}and a sampling distributionρt\\rho\_\{t\}that selects the new component value \(and, in the ordered case, its insertion position\)\. Between insertions, a within\-length generatorℬt\\mathcal\{B\}\_\{t\}refines the existing components\. For continuous modalities,ℬt​f=vt⋅∇f\\mathcal\{B\}\_\{t\}f=v\_\{t\}\\cdot\\nabla f\(with an optional diffusion term\); for discrete modalities,ℬt\\mathcal\{B\}\_\{t\}is the corresponding continuous\-time Markov\-chain generator\.

The key observation is that the conditional generators associated with individual clean targets can be marginalized under the posterior of the target given the current state\. The following theorem states this construction; a measure\-theoretic formulation and proof are provided in Appendix[A\.4](https://arxiv.org/html/2607.09039#A1.SS4)\.\{theorem\}\[store=multimodal\] LetY1=\(X1,S1\)∼qY\_\{1\}=\(X\_\{1\},S\_\{1\}\)\\sim qindex conditional probability pathspt\(⋅\|Y1\)p\_\{t\}\(\\cdot\|Y\_\{1\}\)on𝖸\\mathsf\{Y\}, withp1\(⋅\|Y1\)=δY1p\_\{1\}\(\\cdot\|Y\_\{1\}\)=\\delta\_\{Y\_\{1\}\}\. At statey∈𝖸y\\in\\mathsf\{Y\}, suppose each conditional process has within\-length generatorℬtY1\\mathcal\{B\}\_\{t\}^\{Y\_\{1\}\}, insertion rateλt​\(y\|Y1\)\\lambda\_\{t\}\(y\|Y\_\{1\}\), and sampling distributionρt​\(d​m\|y,Y1\)\\rho\_\{t\}\(\\mathop\{\}\\\!\\mathrm\{d\}m\|y,Y\_\{1\}\)for the insertion variablemm\. For anyyywithpt​\(y\)\>0p\_\{t\}\(y\)\>0, the posterior of the clean target isp​\(d​y1\|Yt=y\):=pt​\(y\|y1\)​q​\(d​y1\)/pt​\(y\)p\(\\mathop\{\}\\\!\\mathrm\{d\}y\_\{1\}\|Y\_\{t\}=y\):=\{p\_\{t\}\(y\|y\_\{1\}\)q\(\\mathop\{\}\\\!\\mathrm\{d\}y\_\{1\}\)\}/\{p\_\{t\}\(y\)\}, wherept=∫pt\(⋅\|y1\)q\(dy1\)p\_\{t\}=\\int p\_\{t\}\(\\cdot\|y\_\{1\}\)q\(\\mathop\{\}\\\!\\mathrm\{d\}y\_\{1\}\)\. Define the marginals by expectations over this posterior:

ℬt∗​f​\(y\):=𝔼Y1∼p\(⋅\|Yt=y\)​\[ℬtY1​f​\(y\)\],λt∗​\(y\):=𝔼Y1∼p\(⋅\|Yt=y\)​\[λt​\(y\|Y1\)\],\\displaystyle\\mathcal\{B\}\_\{t\}^\{\*\}f\(y\):=\\mathbb\{E\}\_\{Y\_\{1\}\\sim p\(\\cdot\|Y\_\{t\}=y\)\}\[\\mathcal\{B\}\_\{t\}^\{Y\_\{1\}\}f\(y\)\],\\quad\\lambda\_\{t\}^\{\*\}\(y\):=\\mathbb\{E\}\_\{Y\_\{1\}\\sim p\(\\cdot\|Y\_\{t\}=y\)\}\[\\lambda\_\{t\}\(y\|Y\_\{1\}\)\],\(5\)ρt∗​\(d​m\|y\):=𝔼Y1∼p\(⋅\|Yt=y\)​\[λt​\(y\|Y1\)​ρt​\(d​m\|y,Y1\)\]𝔼Y1∼p\(⋅\|Yt=y\)​\[λt​\(y\|Y1\)\],\\displaystyle\\rho\_\{t\}^\{\*\}\(\\mathop\{\}\\\!\\mathrm\{d\}m\|y\):=\\frac\{\\mathbb\{E\}\_\{Y\_\{1\}\\sim p\(\\cdot\|Y\_\{t\}=y\)\}\[\\lambda\_\{t\}\(y\|Y\_\{1\}\)\\rho\_\{t\}\(\\mathop\{\}\\\!\\mathrm\{d\}m\|y,Y\_\{1\}\)\]\}\{\\mathbb\{E\}\_\{Y\_\{1\}\\sim p\(\\cdot\|Y\_\{t\}=y\)\}\[\\lambda\_\{t\}\(y\|Y\_\{1\}\)\]\},\(6\)whereρt∗\\rho\_\{t\}^\{\*\}can be arbitrary when the marginal rate is zero\. Then the process withℬt∗\\mathcal\{B\}\_\{t\}^\{\*\}and insertion kernelQt∗=λt∗​ρt∗Q\_\{t\}^\{\*\}=\\lambda\_\{t\}^\{\*\}\\rho\_\{t\}^\{\*\}generates the marginalptp\_\{t\}and hencep1=qp\_\{1\}=q\. The proof follows the same posterior\-marginalization principle as Theorem[1](https://arxiv.org/html/2607.09039#S3.E1)\. In the continuous case,ℬt∗\\mathcal\{B\}\_\{t\}^\{\*\}reduces to the usual marginal vector field; in the discrete case, the same posterior expectation marginalizes the within\-length transition rates\. We learnλt∗\\lambda\_\{t\}^\{\*\}using the Poisson NLL in Eq\.[4](https://arxiv.org/html/2607.09039#S3.E4)\. The within\-length generator is learned with the appropriate continuous or discrete flow\-matching lossℒFM\\mathcal\{L\}\_\{\\text\{FM\}\}\. As the sampling distributionρt\(⋅\|y,Y1\)\\rho\_\{t\}\(\\cdot\|y,Y\_\{1\}\)is a mixture over theX1−XtX\_\{1\}\-X\_\{t\}upcoming insertions, the general reconstruction NLL is:

ℒrec=𝔼Y1,Yt​\[−ht​∑k=1X1−Xtlog⁡ρt​\(Yt,sk\)\],\\mathcal\{L\}\_\{\\text\{rec\}\}=\\mathbb\{E\}\_\{Y\_\{1\},Y\_\{t\}\}\\left\[\-h\_\{t\}\\sum\_\{k=1\}^\{X\_\{1\}\-X\_\{t\}\}\\log\\rho\_\{t\}\(Y\_\{t\},s\_\{k\}\)\\right\],\(7\)where the marginal rate term\(X1−Xt\)​ht\(X\_\{1\}\-X\_\{t\}\)h\_\{t\}is combined with the marginal length denominator of1/\(X1−Xt\)1/\(X\_\{1\}\-X\_\{t\}\)to give the final coefficient\. Generally, the final loss for multimodal GPFlow is

ℒ=ℒGP\+wrec​ℒrec\+wFM​ℒFM,\\mathcal\{L\}=\\mathcal\{L\}\_\{\\text\{GP\}\}\+w\_\{\\text\{rec\}\}\\mathcal\{L\}\_\{\\text\{rec\}\}\+w\_\{\\text\{FM\}\}\\mathcal\{L\}\_\{\\text\{FM\}\},\(8\)wherewrec,wFMw\_\{\\text\{rec\}\},w\_\{\\text\{FM\}\}are hyperparameters\. We describe the parameterization choices for learning the flow and sampling distributions in Appendix[B\.3](https://arxiv.org/html/2607.09039#A2.SS3)and[B\.4](https://arxiv.org/html/2607.09039#A2.SS4)\. Sampling from a multimodal GPFlow generally follows the procedure discussed above, in which flow sampling for the existing length\-dependent modalities is performed using the standard approach with the learnedℬt\\mathcal\{B\}\_\{t\}\. An event is then sampled according to the learned rateλt\\lambda\_\{t\}, and if successful, it will draw the new insertion value from the learned sampling distributionsnew∼ρts\_\{\\text\{new\}\}\\sim\\rho\_\{t\}\. We outline the sampling procedure in Algorithm[1](https://arxiv.org/html/2607.09039#alg1)\.

In practice, because proteins have a natural ordering according to the residue index, we need to extend our formulation to an order\-preserving multimodal GPFlow\. Intuitively, this can be easily obtained by modelingXt\+1X\_\{t\}\+1separate rates and sampling distributions to represent the different insertion positions while preserving the order\. A more rigorous formulation together with a modified characterization of the sampling distribution is available in Appendix[B\.1](https://arxiv.org/html/2607.09039#A2.SS1)\.

Algorithm 1Sampling from Multimodal GPFlow \(Euler\)1:

Y0=\(X0=0,S0=∅\)Y\_\{0\}=\(X\_\{0\}=0,S\_\{0\}=\\emptyset\)\.

2:for

i←1,2,…,Ni\\leftarrow 1,2,\\dots,Ndo

3:

t←ti,Δ​t←ti\+1−tit\\leftarrow t\_\{i\},\\Delta t\\leftarrow t\_\{i\+1\}\-t\_\{i\}, where

tN\+1:=1t\_\{N\+1\}:=1\.

4:Predict the rate

λθ​\(Yt,t\)\\lambda\_\{\\theta\}\(Y\_\{t\},t\), the vector field

vθ​\(Yt,t\)v\_\{\\theta\}\(Y\_\{t\},t\), and the sampling distribution

ρθ​\(Yt,t\)\\rho\_\{\\theta\}\(Y\_\{t\},t\)\.

5:Advance the generative flow for each

StS\_\{t\}with

vθv\_\{\\theta\}to obtain

S^t\+Δ​t\\hat\{S\}\_\{t\+\\Delta t\}\.

6:Simulate the Poisson process by sampling an event on

Xt\+Δ​tX\_\{t\+\\Delta t\}with probability

λθ​Δ​t\\lambda\_\{\\theta\}\\Delta t\.

7:If an event happens, insert the new feature value

snew∼ρθs\_\{\\text\{new\}\}\\sim\\rho\_\{\\theta\}into

S^t\+Δ​t\\hat\{S\}\_\{t\+\\Delta t\}to obtain

St\+Δ​tS\_\{t\+\\Delta t\}\.

8:Update

Yt\+Δ​t=\(Xt\+Δ​t,St\+Δ​t\)Y\_\{t\+\\Delta t\}=\(X\_\{t\+\\Delta t\},S\_\{t\+\\Delta t\}\)\.

9:Return:

Y1=\(X1,S1\)Y\_\{1\}=\(X\_\{1\},S\_\{1\}\)\.

Adapting to conditional generation is generally straightforward for GPFlow and does not affect the theoretical results\. In the multimodal setup, the condition partsY0=\(Xc,Sc\)Y\_\{0\}=\(X\_\{c\},S\_\{c\}\)are sampled from the clean conditioning data and remain unchanged along the conditional probability path\. During sampling, instead of starting from an empty list, we initialize from the provided conditions and their length\-dependent modality, mimicking the standard “inpainting” setup\. Depending on the task, we may or may not allow insertions between the conditions\. For example, in the motif scaffolding setup, insertion is not allowed inside a contiguous motif part, but is allowed between different motifs and at the beginning and end\.

## 4Theoretical Guarantees on Training Dynamics

In this section, we derive a generator\-based upper bound on the KL divergence between the ground\-truth and generated distributions and relate its terms to the exact likelihood components of our loss objective\. We consider general Markov processes as a superset of GPFlow\. We first introduce the*generator*as a useful tool for describing the time evolution of their probability densities, and then derive a bound on the KL divergence between the distributions generated by two Markov processes\. According toHolderrieth et al\. \[[23](https://arxiv.org/html/2607.09039#bib.bib23)\], under some regularity assumptions, a Markov process\{Xt\}t≥0\\\{X\_\{t\}\\\}\_\{t\\geq 0\}can be characterized with its generator𝒜t​f​\(Xt\)=limh→0\+\(𝔼​\[f​\(Xt\+h\)\|Xt\]−f​\(Xt\)\)/h\\mathcal\{A\}\_\{t\}f\(X\_\{t\}\)=\\lim\_\{h\\to 0^\{\+\}\}\(\\mathbb\{E\}\[f\(X\_\{t\+h\}\)\|X\_\{t\}\]\-f\(X\_\{t\}\)\)/h, which can be determined by the velocity fieldutu\_\{t\}, the diffusion coefficientΣt\\Sigma\_\{t\}, and the jump measureQtQ\_\{t\}as:

𝒜t​f​\(x\)=ut​\(x\)⋅∇f​\(x\)\+12​Tr​\(Σt​\(x\)​∇2f​\(x\)\)\+∫\[f​\(y\)−f​\(x\)\]​Qt​\(d​y\|x\)\.\\mathcal\{A\}\_\{t\}f\(x\)=u\_\{t\}\(x\)\\cdot\\nabla f\(x\)\+\\frac\{1\}\{2\}\\text\{Tr\}\\left\(\\Sigma\_\{t\}\(x\)\\nabla^\{2\}f\(x\)\\right\)\+\\int\\left\[f\(y\)\-f\(x\)\\right\]Q\_\{t\}\(\\mathrm\{d\}y\|x\)\.\(9\)If allXtX\_\{t\}’s have a probability density functionptp\_\{t\}, then𝔼x∼pt​\[𝒜t​f​\(x\)\]=∫p˙t​\(x\)​f​\(x\)​dx\\mathbb\{E\}\_\{x\\sim p\_\{t\}\}\[\\mathcal\{A\}\_\{t\}f\(x\)\]=\\int\\dot\{p\}\_\{t\}\(x\)f\(x\)\\,\\mathrm\{d\}xconnects the generator with the time evolution of the densityptp\_\{t\}\. In common diffusion models, we consider a Markov process whereΣt=σt2​I\\Sigma\_\{t\}=\\sigma\_\{t\}^\{2\}Ifor some scalar functionσt\\sigma\_\{t\}\.

### 4\.1KL Divergence Bound for Markov Processes

We now derive a generator\-based bound on the KL divergence between the distributions generated by two Markov processes, which generalizes the variable\-length result inCampbell et al\. \[[9](https://arxiv.org/html/2607.09039#bib.bib9)\]\. See Appendix[A\.6](https://arxiv.org/html/2607.09039#A1.SS6)for a detailed proof\.

\{theorem\}

\[store=klbound\] Let\{Xt\}t≥0\\\{X\_\{t\}\\\}\_\{t\\geq 0\}and\{X~t\}t≥0\\\{\\tilde\{X\}\_\{t\}\\\}\_\{t\\geq 0\}be two Markov processes with generators𝒜t\\mathcal\{A\}\_\{t\}and𝒜~t\\tilde\{\\mathcal\{A\}\}\_\{t\}, respectively\. Suppose𝒜t\\mathcal\{A\}\_\{t\}is determined by\(ut,σt2,Qt\)\(u\_\{t\},\\sigma\_\{t\}^\{2\},Q\_\{t\}\)and𝒜~t\\tilde\{\\mathcal\{A\}\}\_\{t\}is determined by\(u~t,σt2,Q~t\)\(\\tilde\{u\}\_\{t\},\\sigma\_\{t\}^\{2\},\\tilde\{Q\}\_\{t\}\), so they coincide in the diffusion coefficient\. Assume that both processes have the same initial distributionp0p\_\{0\}and that sufficient regularity conditions hold\. Then, the KL divergence betweenpTp\_\{T\}andp~T\\tilde\{p\}\_\{T\}can be bounded by the loss termsℒt​\(x\)\\mathcal\{L\}\_\{t\}\(x\)as

𝒟KL​\(pT∥p~T\)≤∫0T𝔼x∼pt​\[ℒt​\(x\)\]​dt,ℒt​\(x\)=‖ut​\(x\)−u~t​\(x\)‖222​σt2\+𝒟KL​\(Qt∥Q~t\),\\mathcal\{D\}\_\{\\text\{KL\}\}\(p\_\{T\}\\\|\\tilde\{p\}\_\{T\}\)\\leq\\int\_\{0\}^\{T\}\\mathbb\{E\}\_\{x\\sim p\_\{t\}\}\\left\[\\mathcal\{L\}\_\{t\}\(x\)\\right\]\\mathrm\{d\}t,\\quad\\mathcal\{L\}\_\{t\}\(x\)=\\frac\{\\\|u\_\{t\}\(x\)\-\\tilde\{u\}\_\{t\}\(x\)\\\|\_\{2\}^\{2\}\}\{2\\sigma\_\{t\}^\{2\}\}\+\\mathcal\{D\}\_\{\\text\{KL\}\}\(Q\_\{t\}\\\|\\tilde\{Q\}\_\{t\}\),\(10\)where the first drift\-diffusion term is omitted when no continuous component is present, and𝒟KL​\(Qt∥Q~t\)\\mathcal\{D\}\_\{\\text\{KL\}\}\(Q\_\{t\}\\\|\\tilde\{Q\}\_\{t\}\)denotes the generalized KL divergence between two positive measures:

𝒟KL​\(Q∥Q~\)=∫Q~​\(d​y\|x\)\+∫\(log⁡d​Q​\(y\|x\)d​Q~​\(y\|x\)−1\)​Q​\(d​y\|x\)\.\\displaystyle\\mathcal\{D\}\_\{\\text\{KL\}\}\(Q\\\|\\tilde\{Q\}\)=\\int\\tilde\{Q\}\(\\mathrm\{d\}y\|x\)\+\\int\\left\(\\log\\frac\{\\mathrm\{d\}Q\(y\|x\)\}\{\\mathrm\{d\}\\tilde\{Q\}\(y\|x\)\}\-1\\right\)Q\(\\mathrm\{d\}y\|x\)\.\(11\)

### 4\.2Generator View of GPFlow

We now discuss how to interpret GPFlow as a special case of a continuous\-time generator on the joint state space𝖸\\mathsf\{Y\}\. The marginal within\-length generatorℬt∗\\mathcal\{B\}\_\{t\}^\{\*\}that refines existing components constitutes the first drift\-diffusion term in the continuous case or the generalized KL term of the transition kernel for the finite Markov chain in the discrete case\[[17](https://arxiv.org/html/2607.09039#bib.bib17)\]\. On the other hand, the marginal jump measure across dimensions factorizes asQt∗=λt∗​ρt∗Q\_\{t\}^\{\*\}=\\lambda\_\{t\}^\{\*\}\\rho\_\{t\}^\{\*\}: at stateyy, an event occurs at rateλt∗​\(y\)\\lambda\_\{t\}^\{\*\}\(y\), after which the new component \(and insertion position, when applicable\) is sampled fromρt∗​\(d​m\|y\)\\rho\_\{t\}^\{\*\}\(\\mathop\{\}\\\!\\mathrm\{d\}m\|y\)\. The learned counterpart is the jump kernelQθ=λθ​ρθQ\_\{\\theta\}=\\lambda\_\{\\theta\}\\rho\_\{\\theta\}, factorized into rateλθ​\(y,t\)\\lambda\_\{\\theta\}\(y,t\)and sampling distributionρθ​\(d​m\|y,t\)\\rho\_\{\\theta\}\(\\mathop\{\}\\\!\\mathrm\{d\}m\|y,t\)\. See Appendix[A\.7](https://arxiv.org/html/2607.09039#A1.SS7)for the precise state\-space construction and the explicit generator definition\. One crucial observation is that the gradient of our NLL\-based rate loss objective coincides with the gradient of the expected generalized KL divergence:

∇θℒGP=∇θ𝔼t,Yt∼pt​\[𝒟KL​\(λt∗​\(Yt\)∥λθ​\(Yt,t\)\)\]\.\\nabla\_\{\\theta\}\\mathcal\{L\}\_\{\\text\{GP\}\}=\\nabla\_\{\\theta\}\\mathbb\{E\}\_\{t,Y\_\{t\}\\sim p\_\{t\}\}\\left\[\\mathcal\{D\}\_\{\\text\{KL\}\}\\bigl\(\\lambda\_\{t\}^\{\*\}\(Y\_\{t\}\)\\\|\\lambda\_\{\\theta\}\(Y\_\{t\},t\)\\bigr\)\\right\]\.\(12\)In Appendix[A\.7](https://arxiv.org/html/2607.09039#A1.SS7), we describe how the jump contribution decomposes exactly into a generalized KL for the rate and a rate\-weighted KL for the sampling distribution\. Therefore, Theorem[4\.1](https://arxiv.org/html/2607.09039#S4.SS1)is applicable and implies that the additive GPFlow objective in Eq\.[8](https://arxiv.org/html/2607.09039#S3.E8)minimizes the corresponding upper\-bound terms, thereby providing a finer\-grained distributional guarantee on the training dynamics\.

Additionally, our theoretical framework offers a natural explanation of why the generalized KL divergence is chosen over other Bregman divergences, such as MSE\. Importantly, existing works’ choice of the generalized KL divergence as the rate loss is purely empirical\. In contrast, our choice is derived directly from the exact point process NLL, thereby connecting variable\-length generative loss to the underlying, more principled stochastic process NLL minimization\.

## 5Experimental Results

We instantiate our GPFlow across five diverse protein design tasks across continuous, discrete, and Riemannian modalities\. Additional details on datasets, models, and pipelines can be found in Appendix[C](https://arxiv.org/html/2607.09039#A3)\. Additional ablation studies and experimental results are in Appendix[D](https://arxiv.org/html/2607.09039#A4)and[E](https://arxiv.org/html/2607.09039#A5)\.

### 5\.1Unconditional Protein Structure Generation

We first instantiate GPFlow for unconditional structure generation, in which the CA coordinates serve as the continuous length\-dependent modality\. We follow the setup in Proteina\[[18](https://arxiv.org/html/2607.09039#bib.bib18)\], a CA\-only protein generative model\. We build our order\-preserving GPFlow on the 60M Proteina model, adding additional heads to predict rates and reconstruct sampling distributions for the continuous coordinates, bringing the total to 65M trainable parameters\. The datasets also follow the Proteina setup, including PDB\[[5](https://arxiv.org/html/2607.09039#bib.bib5)\]and AFDB\[[25](https://arxiv.org/html/2607.09039#bib.bib25)\]\. For PDB, we retain only single\-chain entries with lengths between 50 and 256, yielding 27k data samples in total\. For AFDB, we follow Proteina’s clustering instructions to obtain 713k data samples\. We train two GPFlow variants on the two datasets from scratch\. For the Proteina baseline, we use its official 60M checkpoint trained on AFDB for inference and retrain a model on PDB using the same hyperparameters\.

Table 1:Unconditional protein structure design benchmark\. The best result for each metric is inbold\.\*A 60M Proteina is retrained on PDB\. †Official Proteina 60M checkpoint for inference\.

For the fixed\-length models, we follow the standard Proteina benchmark to sample 100 structures at each length in\{50,100,150,200,250\}\\\{50,100,150,200,250\\\}\. For GPFlow, we generate 500 structures without a length input; their lengths are generated from the learned rate\. We reportDesignabilityas the proportion of*designable*structures averaged across all generated samples\. A designable structure is defined as having a best self\-consistent RMSD \(scRMSD\) less than 2 Å to the 8 ProteinMPNN\[[15](https://arxiv.org/html/2607.09039#bib.bib15)\]redesigned and ESMFold\[[33](https://arxiv.org/html/2607.09039#bib.bib33)\]refolded candidates\. Additional diversity\-based metrics and secondary\-structure ratios are also reported\.Diversityis defined as the proportion of the designable clusters in all designable samples\.Noveltyis defined as the maximum TM\-score\[[52](https://arxiv.org/html/2607.09039#bib.bib52)\]via exhaustive search against the PDB or AFDB databases, averaged across all samples\. Therefore, a higher diversity score indicates more designable clusters, and a lower novelty score indicates more novel generations\.

![Refer to caption](https://arxiv.org/html/2607.09039v1/x2.png)Figure 1:Length distributions and representative examples of GPFlow\-generated structures\.The results are summarized in Table[1](https://arxiv.org/html/2607.09039#S5.T1), where GPFlow achieves a 3\.9%–18\.5% improvement in designability over the corresponding Proteina base model\. Figure[1](https://arxiv.org/html/2607.09039#S5.F1)serves as an intrinsic calibration check for the learned length marginal: GPFlow’s generated distribution follows the empirical training profile with a mean length of 153\.7, whereas the fixed\-length benchmark is 150\. Appendix[E](https://arxiv.org/html/2607.09039#A5)further reports designability across GPFlow’s generated length bins, providing more concrete evidence that our superior designability is not due to the unmatched length distribution but to its better capture of the joint distribution\. In terms of diversity, we note that GPFlow achieves a lower TM\-score than the corresponding Proteina model, indicating greater novelty, although both models fall behind other baselines\. The generally lower diversity of the Proteina\-family models, including GPFlow, is due to the inductive bias of low\-temperature sampling, further ablated in Appendix[D\.2](https://arxiv.org/html/2607.09039#A4.SS2)\. The motif scaffolding benchmark provides a more concrete evaluation of the diversity of GPFlow through the number of unique successes after clustering\.

### 5\.2Unconditional Protein Sequence Generation

![Refer to caption](https://arxiv.org/html/2607.09039v1/figs/gpflow_seq.png)Figure 2:Performance benchmark for unconditional protein sequence design\.\(a\)Foldability \(pLDDT\) across sequence length bins\.\(b\)Novelty measured by TM\-score to the nearest PDB structure\.\(c\)Diversity by FoldSeek clustering\.\(d\)Secondary structure distributions\.\(e\)Length distribution of GPFlow samples generated without length conditioning\. Error bands denote±\\pm1 SEM\.We instantiate GPFlow to generate variable\-length protein sequences, incorporating the amino acid sequence as the discrete length\-dependent modality\. We train GPFlow on 41M UniRef50\[[44](https://arxiv.org/html/2607.09039#bib.bib44)\]sequences of lengths 10\-1024 with a 642M\-parameter DiT\[[38](https://arxiv.org/html/2607.09039#bib.bib38)\]\. To improve long\-sequence generation, we followHavasi et al\. \[[20](https://arxiv.org/html/2607.09039#bib.bib20)\]to adopt a localized probability path as detailed in Appendix[C\.2\.5](https://arxiv.org/html/2607.09039#A3.SS2.SSS5)\. For inference\-time correction, we additionally learn a deletion process that can be coupled with the insertion process during sampling\. Three different paradigms of protein sequence generative models are compared\. For fixed\-length discrete diffusion over categorical amino\-acid tokens, we evaluate EvoDiff\-OADM\[[1](https://arxiv.org/html/2607.09039#bib.bib1)\]and DPLM\[[48](https://arxiv.org/html/2607.09039#bib.bib48)\]\. For masked language models with iterative decoding, we evaluate ESM3\[[21](https://arxiv.org/html/2607.09039#bib.bib21)\]restricted to its sequence track\. For variable\-length generation, we evaluate SCISOR\[[4](https://arxiv.org/html/2607.09039#bib.bib4)\], a deletion\-based discrete model, as a special case of EditFlow\[[20](https://arxiv.org/html/2607.09039#bib.bib20)\]\. While other baselines are trained on UniRef50, ESM3 is trained on a substantially larger corpus that subsumes UniRef50 and is included for reference\. To evaluate the generated samples, we fold the generated sequences with ESMFold\[[32](https://arxiv.org/html/2607.09039#bib.bib32)\]and quantify: \(a\)Foldabilityvia mean pLDDT; \(b\)Noveltyviapdb\-TM, the maximum TM\-score to any PDB structure found by FoldSeek search \(TM<0\.5\\text\{TM\}<0\.5\); \(c\)Diversityvia FoldSeek clustering atTM≥0\.5\\text\{TM\}\\geq 0\.5, reported as the proportion of unique clusters among all generations; \(d\)Secondary structurecomposition via DSSP \(helix/strand/coil fractions\)\[[26](https://arxiv.org/html/2607.09039#bib.bib26)\]; and \(e\)Length distribution, compared to UniRef50\.

Figure[2](https://arxiv.org/html/2607.09039#S5.F2)\(a\-d\) evaluates structural statistics when assessed by ESMFold\. Across all lengths, GPFlow is the closest match to UniRef50 in mean pLDDT, faithfully reproducing the structural characteristics rather than optimizing toward uniformly high\-confidence folds\. DPLM achieves an abnormally high pLDDT compared to the training data, with a mean pLDDT difference of 14\.40, whereas GPFlow achieves only 3\.17\. DPLM has the lowest structural novelty and diversity among the models, illustrating a clear trade\-off and off\-data behavior, whereas GPFlow maintains 17% higher diversity while keeping novelty closest to UniRef50\. In contrast, ESM3, EvoDiff, and SCISOR exhibit consistently lower pLDDT together with lowerpdb\-TM\. Finally, GPFlow best matches UniRef50 in secondary\-structure composition, including a near\-identical mean structural balance\. DPLM deviates from the UniRef50 profile in loop content, whereas ESM3, EvoDiff, and SCISOR exhibit large departures from the UniRef50 profile\.

Figure[2](https://arxiv.org/html/2607.09039#S5.F2)\(e\) illustrates the variable\-length formulation: GPFlow learns and samples from an emergent length prior that closely matches the training length\. Fixed\-length baselines define onlypθ​\(x\|L\)p\_\{\\theta\}\(x\|L\)and therefore have no marginal length distribution without an additional sampler overLL\. Selected generated sequences folded by ESMFold are presented in Figure[3](https://arxiv.org/html/2607.09039#S5.F3)with their lengths and pLDDTs\.

![Refer to caption](https://arxiv.org/html/2607.09039v1/figs/seq_sample.png)Figure 3:Representative unconditional sequence generations folded by ESMFold\.
### 5\.3Motif Scaffolding

Table 2:Number of unique successes \(1000 generations for each task\) on the RFDiffusion benchmark\[[49](https://arxiv.org/html/2607.09039#bib.bib49)\]\. Superscripts denote the number of motif segments\. Motif tasks with multiple lengths are indicated in parentheses\. Best results are shown inbold\.For conditional generation, we instantiate GPFlow for motif scaffolding, with the motif structure and/or sequence as the conditioning variable\. Starting from isolated motif segments, our model gradually grows a protein scaffold through inserting and updating linker residues between motif segments, without relying on predefined contig templates\. We consider both structure\-based and sequence\-based motif scaffolding tasks\. The latter is deferred to Appendix[E\.4](https://arxiv.org/html/2607.09039#A5.SS4)\.

For structure\-based motif scaffolding, we build our GPFlow upon the Proteina variant, which injects motif structure and residue information as conditioning signals and fixes the motif portion during both training and sampling\. We train it from scratch on the same AFDB dataset and adopt the motif training augmentation used in Proteina\. The evaluation criteria followWatson et al\. \[[49](https://arxiv.org/html/2607.09039#bib.bib49)\], which uses a combination of scRMSD, motifRMSD, pLDDT, and pAE thresholds detailed in Appendix[C\.1\.3](https://arxiv.org/html/2607.09039#A3.SS1.SSS3)\. All successes are clustered via FoldSeek to count unique successes rather than raw successes\.

![Refer to caption](https://arxiv.org/html/2607.09039v1/x3.png)Figure 4:GPFlow generates more diverse clusters for the same motif \(highlighted in red\)\.Table[2](https://arxiv.org/html/2607.09039#S5.T2)lists the unique success counts across 16 motif tasks from the RFDiffusion benchmark\[[49](https://arxiv.org/html/2607.09039#bib.bib49)\], where different contig templates corresponding to the same motif target are grouped into a single task for GPFlow\. Overall, GPFlow ranks first on 10 out of 16 motif tasks\. For easier tasks like 1BCF and 3IXT, GPFlow yields considerably more diverse generations, with variation in both secondary structures and scaffold lengths, whereas the base Proteina generations collapse into a single cluster\. \(Figure[4](https://arxiv.org/html/2607.09039#S5.F4)\)\. For more challenging tasks like 5WN9 and 5YUI, the raw success counts of GPFlow also dominate\. Noticeably, GPFlow can also generate successful scaffolds longer than 200 residues, whereas Proteina generations are restricted by the evaluated contig templates\. A finer\-grained analysis of the generation lengths in Appendix[E\.3](https://arxiv.org/html/2607.09039#A5.SS3)further confirms the diversity of our generations across different length bins\. Our evaluations show that the learned length\-varying process can produce successful scaffolds over a wider range than the predefined contig choices used by the fixed\-length baselines\.

### 5\.4Peptide Co\-Design

Finally, we extend GPFlow to peptide sequence\-structure co\-design\. Given a target receptor protein, the model jointly generates the all\-atom structure and amino\-acid sequence of a binding peptide, with the peptide length not predetermined\. We use the dataset and benchmark fromLi et al\. \[[29](https://arxiv.org/html/2607.09039#bib.bib29)\], in which the receptor structure and pocket positions are provided as conditioning context\. GPFlow is adapted from the PepFlow\[[29](https://arxiv.org/html/2607.09039#bib.bib29)\]architecture, in which per\-residue rotation, translation, residue type, and side\-chain torsion angles are learned jointly to model the all\-atom peptide structure\. Rotation and torsion angles are modeled using Riemannian flow matching, with an additional extension that learns inserted elements using the corresponding Riemannian norm\. The residue types are modeled using a simplex\-based continuous representation that embeds soft one\-hot logits in Euclidean geometry\.

![Refer to caption](https://arxiv.org/html/2607.09039v1/figs/pepdemo2.png)Figure 5:Samples of GPFlow\-generated peptides compared to ground truth and PepFlow\[[29](https://arxiv.org/html/2607.09039#bib.bib29)\]\.We compare against the three co\-design baselines: RFdiffusion\[[49](https://arxiv.org/html/2607.09039#bib.bib49)\]composed with ProteinMPNN\[[15](https://arxiv.org/html/2607.09039#bib.bib15)\]for sequence design, ProteinGenerator\[[35](https://arxiv.org/html/2607.09039#bib.bib35)\]that jointly samples backbone and sequence, and the base PepFlow model\. We report eight generative metrics: amino\-acid recovery \(AAR\),RMSD, secondary\-structure ratio \(SSR\), binding\-site recovery \(BSR\),Affinity,Stability,DesignabilityandDiversity\. The complete evaluation protocol and the length\-matched/unmatched breakdown are provided in Appendix[C\.3\.3](https://arxiv.org/html/2607.09039#A3.SS3.SSS3)\. As shown in Table[3](https://arxiv.org/html/2607.09039#S5.T3), GPFlow attains the best performance on 4 out of the 8 metrics and also surpasses the base PepFlow in terms of Designability and Diversity\. Notably, GPFlow is the only model that samples peptide length from a learned rate function rather than conditioning on the native length \(Appendix[E\.5](https://arxiv.org/html/2607.09039#A5.SS5)\)\. This matters forde novopeptide design, where fixed\-length baselines rely on a native\-length oracle unavailable in practice\. Figure[5](https://arxiv.org/html/2607.09039#S5.F5)highlights two qualitative behaviors of GPFlow with a better energy landscape\. In 1cwd\_P, the processed ground truth contains a gap induced by the removal of a non\-standard residue, and PepFlow inherits this gap through its native\-indexed generation\. In contrast, GPFlow generates a continuous peptide across this region\. In 2qlj\_E, the generated insertion extends more deeply into the binding site, suggesting greater sensitivity to the local pocket geometry\.

Table 3:Conditional sequence\-structure co\-design on peptide design\. Best results are shown inbold\.†are composite over the GT\-length matched/unmatched samples; see Appendix[C\.3\.3](https://arxiv.org/html/2607.09039#A3.SS3.SSS3)for details\.

## 6Conclusion & Limitations

In this work, we propose GPFlow, a unified framework for variable\-length generative modeling applicable to diverse protein design scenarios\. GPFlow uses a generalized Poisson process coupled to within\-length dynamics and covers multimodal settings with continuous, discrete, and mixed modalities\. Simulation\-free rate learning based on exact negative log\-likelihood minimization is proven via posterior marginalization, with an additional generator\-based KL bound on the training dynamics\. Our comprehensive experimental setup for GPFlow provides concrete evidence of its flexibility in unconditional and conditional generation without requiring a prescribed target length and demonstrates its superior performance in designability and distributional fitness across all tasks\.

One limitation of GPFlow is its sensitivity to scheduler choices, especially in multimodal setups where different modalities may interact\. Additional sampling techniques such asτ\\tau\-leaping \(Appendix[B\.5](https://arxiv.org/html/2607.09039#A2.SS5)\) and localized paths \(Appendix[C\.2\.5](https://arxiv.org/html/2607.09039#A3.SS2.SSS5)\) may also introduce subtle trade\-offs\. We are actively exploring empirical training and sampling techniques to develop a more comprehensive unified variable\-length generative framework for biological domains\.

## Acknowledgments

This research is supported in part by the Molecule Maker Lab Institute, an AI Research Institutes program supported by NSF under Award No\. 2505932, and the DOE Center for Advanced Bioenergy and Bioproducts Innovation \(U\.S\. Department of Energy, Office of Science, Biological and Environmental Research Program under Award Number DESC0018420\)\. Any opinions, findings, and conclusions or recommendations expressed in this publication are those of the author\(s\) and do not necessarily reflect the views of the U\.S\. Department of Energy\.

## References

- Alamdari et al\. \[2023\]Sarah Alamdari, Nitya Thakkar, Rianne van den Berg, Alex X\. Lu, Nicolo Fusi, Ava P\. Amini, and Kevin K\. Yang\.Protein generation with evolutionary diffusion: sequence is all you need, September 2023\.URL[https://www\.biorxiv\.org/content/10\.1101/2023\.09\.11\.556673v1](https://www.biorxiv.org/content/10.1101/2023.09.11.556673v1)\.Pages: 2023\.09\.11\.556673 Section: New Results\.
- Austin et al\. \[2021\]Jacob Austin, Daniel D Johnson, Jonathan Ho, Daniel Tarlow, and Rianne Van Den Berg\.Structured denoising diffusion models in discrete state\-spaces\.*Advances in neural information processing systems*, 34:17981–17993, 2021\.
- Baek et al\. \[2021\]Minkyung Baek, Frank DiMaio, Ivan Anishchenko, Justas Dauparas, Sergey Ovchinnikov, Gyu Rie Lee, Jue Wang, Qian Cong, Lisa N Kinch, R Dustin Schaeffer, et al\.Accurate prediction of protein structures and interactions using a three\-track neural network\.*Science*, 373\(6557\):871–876, 2021\.
- Baron et al\. \[2025\]Ethan Baron, Alan N\. Amin, Ruben Weitzman, Debora Marks, and Andrew Gordon Wilson\.A Diffusion Model to Shrink Proteins While Maintaining Their Function, November 2025\.URL[http://arxiv\.org/abs/2511\.07390](http://arxiv.org/abs/2511.07390)\.arXiv:2511\.07390 \[cs\] version: 1\.
- Berman et al\. \[2000\]Helen M Berman, John Westbrook, Zukang Feng, Gary Gilliland, Talapady N Bhat, Helge Weissig, Ilya N Shindyalov, and Philip E Bourne\.The protein data bank\.*Nucleic acids research*, 28\(1\):235–242, 2000\.
- Brémaud \[1975\]Pierre Brémaud\.An extension of watanabe’s theorem of characterization of poisson processes over the positive real half line\.*Journal of Applied Probability*, 12\(2\):396–399, 1975\.
- Brémaud \[1981\]Pierre Brémaud\.Point processes and queues\.*Springer*, 1981\.
- Butcher et al\. \[2025\]Jasper Butcher, Rohith Krishna, Raktim Mitra, Rafael I Brent, Yanjing Li, Nathaniel Corley, Paul T Kim, Jonathan Funk, Simon Mathis, Saman Salike, et al\.De novo design of all\-atom biomolecular interactions with rfdiffusion3\.*bioRxiv*, 2025\.
- Campbell et al\. \[2023\]Andrew Campbell, William Harvey, Christian Weilbach, Valentin De Bortoli, Thomas Rainforth, and Arnaud Doucet\.Trans\-dimensional generative modeling via jump diffusion models\.*Advances in Neural Information Processing Systems*, 36:42217–42257, 2023\.
- Chaudhury et al\. \[2010\]Sidhartha Chaudhury, Sergey Lyskov, and Jeffrey J Gray\.Pyrosetta: a script\-based interface for implementing molecular modeling algorithms using rosetta\.*Bioinformatics*, 26\(5\):689–691, 2010\.
- Chen and Lipman \[2024\]Ricky TQ Chen and Yaron Lipman\.Flow matching on general geometries\.In*International Conference on Learning Representations*, volume 2024, pages 47922–47945, 2024\.
- Cheng et al\. \[2024\]Chaoran Cheng, Jiahan Li, Jian Peng, and Ge Liu\.Categorical flow matching on statistical manifolds\.*Advances in Neural Information Processing Systems*, 37:54787–54819, 2024\.
- Cheng et al\. \[2025\]Chaoran Cheng, Jiahan Li, Jiajun Fan, and Ge Liu\.α\\alpha\-flow: A unified framework for continuous\-state discrete flow matching models\.*arXiv preprint arXiv:2504\.10283*, 2025\.
- Daley and Vere\-Jones \[2003\]Daryl J Daley and David Vere\-Jones\.*An introduction to the theory of point processes: volume I: elementary theory and methods*\.Springer, 2003\.
- Dauparas et al\. \[2022\]Justas Dauparas, Ivan Anishchenko, Nathaniel Bennett, Hua Bai, Robert J Ragotte, Lukas F Milles, Basile IM Wicky, Alexis Courbet, Rob J de Haas, Neville Bethel, et al\.Robust deep learning–based protein sequence design using proteinmpnn\.*Science*, 378\(6615\):49–56, 2022\.
- Dunn and Koes \[2024\]Ian Dunn and David Ryan Koes\.Mixed continuous and categorical flow matching for 3d de novo molecule generation\.*ArXiv*, pages arXiv–2404, 2024\.
- Gat et al\. \[2024\]Itai Gat, Tal Remez, Neta Shaul, Felix Kreuk, Ricky TQ Chen, Gabriel Synnaeve, Yossi Adi, and Yaron Lipman\.Discrete flow matching\.*Advances in Neural Information Processing Systems*, 37:133345–133385, 2024\.
- Geffner et al\. \[2025\]Tomas Geffner, Kieran Didi, Zuobai Zhang, Danny Reidenbach, Zhonglin Cao, Jason Yim, Mario Geiger, Christian Dallago, Emine Kucukbenli, Arash Vahdat, et al\.Proteina: Scaling flow\-based protein structure generative models\.*arXiv preprint arXiv:2503\.00710*, 2025\.
- Gillespie \[2001\]Daniel T Gillespie\.Approximate accelerated stochastic simulation of chemically reacting systems\.*The Journal of chemical physics*, 115\(4\):1716–1733, 2001\.
- Havasi et al\. \[2025\]Marton Havasi, Brian Karrer, Itai Gat, and Ricky TQ Chen\.Edit flows: Flow matching with edit operations\.*arXiv preprint arXiv:2506\.09018*, 2025\.
- Hayes et al\. \[2025\]Thomas Hayes, Roshan Rao, Halil Akin, Nicholas J\. Sofroniew, Deniz Oktay, Zeming Lin, Robert Verkuil, Vincent Q\. Tran, Jonathan Deaton, Marius Wiggert, Rohil Badkundri, Irhum Shafkat, Jun Gong, Alexander Derry, Raul S\. Molina, Neil Thomas, Yousuf A\. Khan, Chetan Mishra, Carolyn Kim, Liam J\. Bartie, Matthew Nemeth, Patrick D\. Hsu, Tom Sercu, Salvatore Candido, and Alexander Rives\.Simulating 500 million years of evolution with a language model\.*Science*, 387\(6736\):850–858, February 2025\.doi:10\.1126/science\.ads0018\.URL[https://www\.science\.org/doi/10\.1126/science\.ads0018](https://www.science.org/doi/10.1126/science.ads0018)\.Publisher: American Association for the Advancement of Science\.
- Ho et al\. \[2020\]Jonathan Ho, Ajay Jain, and Pieter Abbeel\.Denoising diffusion probabilistic models\.*Advances in neural information processing systems*, 33:6840–6851, 2020\.
- Holderrieth et al\. \[2024\]Peter Holderrieth, Marton Havasi, Jason Yim, Neta Shaul, Itai Gat, Tommi Jaakkola, Brian Karrer, Ricky TQ Chen, and Yaron Lipman\.Generator matching: Generative modeling with arbitrary markov processes\.*arXiv preprint arXiv:2410\.20587*, 2024\.
- Hoogeboom et al\. \[2021\]Emiel Hoogeboom, Didrik Nielsen, Priyank Jaini, Patrick Forré, and Max Welling\.Argmax flows and multinomial diffusion: Learning categorical distributions\.*Advances in neural information processing systems*, 34:12454–12465, 2021\.
- Jumper et al\. \[2021\]John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Žídek, Anna Potapenko, et al\.Highly accurate protein structure prediction with alphafold\.*nature*, 596\(7873\):583–589, 2021\.
- Kabsch and Sander \[1983\]Wolfgang Kabsch and Christian Sander\.Dictionary of protein secondary structure: Pattern recognition of hydrogen\-bonded and geometrical features\.*Biopolymers*, 22\(12\):2577–2637, 1983\.doi:10\.1002/bip\.360221211\.
- Kingma and Ba \[2017\]Diederik P\. Kingma and Jimmy Ba\.Adam: A Method for Stochastic Optimization, January 2017\.URL[http://arxiv\.org/abs/1412\.6980](http://arxiv.org/abs/1412.6980)\.arXiv:1412\.6980 \[cs\]\.
- Kunzmann and Hamacher \[2018\]Patrick Kunzmann and Kay Hamacher\.Biotite: a unifying open source computational biology framework in Python\.*BMC Bioinformatics*, 19\(1\):346, 2018\.doi:10\.1186/s12859\-018\-2367\-z\.
- Li et al\. \[2024\]Jiahan Li, Chaoran Cheng, Zuofan Wu, Ruihan Guo, Shitong Luo, Zhizhou Ren, Jian Peng, and Jianzhu Ma\.Full\-atom peptide design based on multi\-modal flow matching\.*arXiv preprint arXiv:2406\.00735*, 2024\.
- Li \[2007\]Tiejun Li\.Analysis of explicit tau\-leaping schemes for simulating chemically reacting systems\.*Multiscale Modeling & Simulation*, 6\(2\):417–436, 2007\.
- Lin et al\. \[2024\]Yeqing Lin, Minji Lee, Zhao Zhang, and Mohammed AlQuraishi\.Out of many, one: Designing and scaffolding proteins at the scale of the structural universe with genie 2\.*arXiv preprint arXiv:2405\.15489*, 2024\.
- Lin et al\. \[2023a\]Zeming Lin, Halil Akin, Roshan Rao, Brian Hie, Zhongkai Zhu, Wenting Lu, Nikita Smetanin, Robert Verkuil, Ori Kabeli, Yaniv Shmueli, Allan dos Santos Costa, Maryam Fazel\-Zarandi, Tom Sercu, Salvatore Candido, and Alexander Rives\.Evolutionary\-scale prediction of atomic\-level protein structure with a language model\.*Science*, 379\(6637\):1123–1130, 2023a\.doi:10\.1126/science\.ade2574\.
- Lin et al\. \[2023b\]Zeming Lin, Halil Akin, Roshan Rao, Brian Hie, Zhongkai Zhu, Wenting Lu, Nikita Smetanin, Robert Verkuil, Ori Kabeli, Yaniv Shmueli, et al\.Evolutionary\-scale prediction of atomic\-level protein structure with a language model\.*Science*, 379\(6637\):1123–1130, 2023b\.
- Lipman et al\. \[2022\]Yaron Lipman, Ricky TQ Chen, Heli Ben\-Hamu, Maximilian Nickel, and Matt Le\.Flow matching for generative modeling\.*arXiv preprint arXiv:2210\.02747*, 2022\.
- Lisanza et al\. \[2025\]Sidney Lyayuga Lisanza, Jacob Merle Gershon, Samuel WK Tipps, Jeremiah Nelson Sims, Lucas Arnoldt, Samuel J Hendel, Miriam K Simma, Ge Liu, Muna Yase, Hongwei Wu, et al\.Multistate and functional protein design using rosettafold sequence space diffusion\.*Nature biotechnology*, 43\(8\):1288–1298, 2025\.
- London et al\. \[2011\]Nir London, Barak Raveh, Eyal Cohen, Guy Fathi, and Ora Schueler\-Furman\.Rosetta flexpepdock web server—high resolution modeling of peptide–protein interactions\.*Nucleic acids research*, 39\(suppl\_2\):W249–W253, 2011\.
- Mariani et al\. \[2013\]Valerio Mariani, Marco Biasini, Alessandro Barbato, and Torsten Schwede\.lDDT: A local superposition\-free score for comparing protein structures and models using distance difference tests\.*Bioinformatics*, 29\(21\):2722–2728, 2013\.doi:10\.1093/bioinformatics/btt473\.
- Peebles and Xie \[2023\]William Peebles and Saining Xie\.Scalable Diffusion Models with Transformers, March 2023\.URL[http://arxiv\.org/abs/2212\.09748](http://arxiv.org/abs/2212.09748)\.arXiv:2212\.09748 \[cs\]\.
- Raveh et al\. \[2011\]Barak Raveh, Nir London, Lior Zimmerman, and Ora Schueler\-Furman\.Rosetta flexpepdock ab\-initio: simultaneous folding, docking and refinement of peptides onto their receptors\.*PloS one*, 6\(4\):e18934, 2011\.
- Ren et al\. \[2025\]Yinuo Ren, Haoxuan Chen, Yuchen Zhu, Wei Guo, Yongxin Chen, Grant M Rotskoff, Molei Tao, and Lexing Ying\.Fast solvers for discrete diffusion models: Theory and applications of high\-order algorithms\.*arXiv preprint arXiv:2502\.00234*, 2025\.
- Sahoo et al\. \[2024\]Subham Sahoo, Marianne Arriola, Yair Schiff, Aaron Gokaslan, Edgar Marroquin, Justin Chiu, Alexander Rush, and Volodymyr Kuleshov\.Simple and effective masked diffusion language models\.*Advances in Neural Information Processing Systems*, 37:130136–130184, 2024\.
- Song and Ermon \[2019\]Yang Song and Stefano Ermon\.Generative modeling by estimating gradients of the data distribution\.*Advances in neural information processing systems*, 32, 2019\.
- Steinegger and Söding \[2017\]Martin Steinegger and Johannes Söding\.Mmseqs2 enables sensitive protein sequence searching for the analysis of massive data sets\.*Nature biotechnology*, 35\(11\):1026–1028, 2017\.
- Suzek et al\. \[2007\]Baris E\. Suzek, Hongzhan Huang, Peter McGarvey, Raja Mazumder, and Cathy H\. Wu\.UniRef: comprehensive and non\-redundant UniProt reference clusters\.*Bioinformatics*, 23\(10\):1282–1288, May 2007\.ISSN 1367\-4803\.doi:10\.1093/bioinformatics/btm098\.URL[https://doi\.org/10\.1093/bioinformatics/btm098](https://doi.org/10.1093/bioinformatics/btm098)\.
- The UniProt Consortium \[2023\]The UniProt Consortium\.UniProt: the Universal Protein Knowledgebase in 2023\.*Nucleic Acids Research*, 51\(D1\):D523–D531, January 2023\.ISSN 0305\-1048\.doi:10\.1093/nar/gkac1052\.URL[https://doi\.org/10\.1093/nar/gkac1052](https://doi.org/10.1093/nar/gkac1052)\.
- Van Kempen et al\. \[2024\]Michel Van Kempen, Stephanie S Kim, Charlotte Tumescheit, Milot Mirdita, Jeongjae Lee, Cameron LM Gilchrist, Johannes Söding, and Martin Steinegger\.Fast and accurate protein structure search with foldseek\.*Nature biotechnology*, 42\(2\):243–246, 2024\.
- Wang et al\. \[2025\]Lei Wang, Xudong Li, Han Zhang, Jinyi Wang, Dingkang Jiang, Zhidong Xue, and Yan Wang\.A comprehensive review of protein language models\.*arXiv preprint arXiv:2502\.06881*, 2025\.
- Wang et al\. \[2024\]Xinyou Wang, Zaixiang Zheng, Fei Ye, Dongyu Xue, Shujian Huang, and Quanquan Gu\.Diffusion Language Models Are Versatile Protein Learners, October 2024\.URL[http://arxiv\.org/abs/2402\.18567](http://arxiv.org/abs/2402.18567)\.arXiv:2402\.18567 \[cs\]\.
- Watson et al\. \[2023\]Joseph L Watson, David Juergens, Nathaniel R Bennett, Brian L Trippe, Jason Yim, Helen E Eisenach, Woody Ahern, Andrew J Borst, Robert J Ragotte, Lukas F Milles, et al\.De novo design of protein structure and function with rfdiffusion\.*Nature*, 620\(7976\):1089–1100, 2023\.
- Yim et al\. \[2023a\]Jason Yim, Andrew Campbell, Andrew YK Foong, Michael Gastegger, José Jiménez\-Luna, Sarah Lewis, Victor Garcia Satorras, Bastiaan S Veeling, Regina Barzilay, Tommi Jaakkola, et al\.Fast protein backbone generation with se \(3\) flow matching\.*arXiv preprint arXiv:2310\.05297*, 2023a\.
- Yim et al\. \[2023b\]Jason Yim, Brian L Trippe, Valentin De Bortoli, Emile Mathieu, Arnaud Doucet, Regina Barzilay, and Tommi Jaakkola\.Se \(3\) diffusion model with application to protein backbone generation\.*arXiv preprint arXiv:2302\.02277*, 2023b\.
- Zhang and Skolnick \[2004\]Yang Zhang and Jeffrey Skolnick\.Scoring function for automated assessment of protein structure template quality\.*Proteins: Structure, Function, and Bioinformatics*, 57\(4\):702–710, 2004\.doi:10\.1002/prot\.20264\.
- Zhang and Skolnick \[2005\]Yang Zhang and Jeffrey Skolnick\.Tm\-align: a protein structure alignment algorithm based on the tm\-score\.*Nucleic acids research*, 33\(7\):2302–2309, 2005\.

Supplementary Material

###### Contents

1. [1Introduction](https://arxiv.org/html/2607.09039#S1)
2. [2Related Work](https://arxiv.org/html/2607.09039#S2)
3. [3Generalized Poisson Flow](https://arxiv.org/html/2607.09039#S3)1. [3\.1Length\-Only Generalized Poisson Flow](https://arxiv.org/html/2607.09039#S3.SS1) 2. [3\.2Multimodal Generalized Poisson Flow](https://arxiv.org/html/2607.09039#S3.SS2)
4. [4Theoretical Guarantees on Training Dynamics](https://arxiv.org/html/2607.09039#S4)1. [4\.1KL Divergence Bound for Markov Processes](https://arxiv.org/html/2607.09039#S4.SS1) 2. [4\.2Generator View of GPFlow](https://arxiv.org/html/2607.09039#S4.SS2)
5. [5Experimental Results](https://arxiv.org/html/2607.09039#S5)1. [5\.1Unconditional Protein Structure Generation](https://arxiv.org/html/2607.09039#S5.SS1) 2. [5\.2Unconditional Protein Sequence Generation](https://arxiv.org/html/2607.09039#S5.SS2) 3. [5\.3Motif Scaffolding](https://arxiv.org/html/2607.09039#S5.SS3) 4. [5\.4Peptide Co\-Design](https://arxiv.org/html/2607.09039#S5.SS4)
6. [6Conclusion & Limitations](https://arxiv.org/html/2607.09039#S6)
7. [References](https://arxiv.org/html/2607.09039#bib)
8. [ATheoretical Details](https://arxiv.org/html/2607.09039#A1)1. [A\.1Proof for Theorem3\.1\(Conditional Rate\)](https://arxiv.org/html/2607.09039#A1.SS1) 2. [A\.2Proof for Theorem1\(Marginal Rate\)](https://arxiv.org/html/2607.09039#A1.SS2) 3. [A\.3PropositionA\.3\(Poisson NLL\)](https://arxiv.org/html/2607.09039#A1.SS3) 4. [A\.4Proof for Theorem3\.2\(Multimodal GPFlow\)](https://arxiv.org/html/2607.09039#A1.SS4) 5. [A\.5TheoremA\.5\(Order\-Preserving GPFlow\)](https://arxiv.org/html/2607.09039#A1.SS5) 6. [A\.6Proof for Theorem4\.1\(KL Bound\)](https://arxiv.org/html/2607.09039#A1.SS6) 7. [A\.7Generator for GPFlow](https://arxiv.org/html/2607.09039#A1.SS7) 8. [A\.8Connection to EditFlow](https://arxiv.org/html/2607.09039#A1.SS8)
9. [BAlgorithmic Details](https://arxiv.org/html/2607.09039#A2)1. [B\.1Order\-Preserving Generalized Poisson Flow](https://arxiv.org/html/2607.09039#A2.SS1) 2. [B\.2Training and Sampling Pseudocode](https://arxiv.org/html/2607.09039#A2.SS2) 3. [B\.3Flow Parameterization](https://arxiv.org/html/2607.09039#A2.SS3) 4. [B\.4Sampling Distribution Parameterization](https://arxiv.org/html/2607.09039#A2.SS4) 5. [B\.5τ\\tau\-Leaping](https://arxiv.org/html/2607.09039#A2.SS5)
10. [CExperiment Details](https://arxiv.org/html/2607.09039#A3)1. [C\.1Protein Structure Generation](https://arxiv.org/html/2607.09039#A3.SS1) 2. [C\.2Protein Sequence Generation](https://arxiv.org/html/2607.09039#A3.SS2) 3. [C\.3Peptide Co\-Design](https://arxiv.org/html/2607.09039#A3.SS3)
11. [DAblation Studies](https://arxiv.org/html/2607.09039#A4)1. [D\.1Insertion Scheduler and Sampling Step Impact](https://arxiv.org/html/2607.09039#A4.SS1) 2. [D\.2Designability\-Diversity Trade\-Off](https://arxiv.org/html/2607.09039#A4.SS2)
12. [EAdditional Results](https://arxiv.org/html/2607.09039#A5)1. [E\.1Unconditional Protein Generation](https://arxiv.org/html/2607.09039#A5.SS1) 2. [E\.2Sequence\-Based Length Control](https://arxiv.org/html/2607.09039#A5.SS2) 3. [E\.3Structure\-Based Motif Scaffolding](https://arxiv.org/html/2607.09039#A5.SS3) 4. [E\.4Sequence\-Based Motif Scaffolding](https://arxiv.org/html/2607.09039#A5.SS4) 5. [E\.5Peptide Co\-Design](https://arxiv.org/html/2607.09039#A5.SS5)
13. [FBroader Impacts](https://arxiv.org/html/2607.09039#A6)

## Appendix ATheoretical Details

In this section, we provide proofs for the major theorems established in the main context\. We also discuss the technical details of interpreting GPFlow as a generator and its connections to existing models such as EditFlow\.

### A\.1Proof for Theorem[3\.1](https://arxiv.org/html/2607.09039#S3.SS1)\(Conditional Rate\)

\\getkeytheorem

forwardrate

###### Proof\.

ConsiderYt=L−XtY\_\{t\}=L\-X\_\{t\}as a pure\-death process with death rateμt​\(Yt\)=Yt​ht\\mu\_\{t\}\(Y\_\{t\}\)=Y\_\{t\}h\_\{t\}\. Letpt​\(k\):=Pr⁡\(Yt=k\)p\_\{t\}\(k\):=\\Pr\(Y\_\{t\}=k\)\. The Kolmogorov forward equation reads:

p˙t​\(k\)\\displaystyle\\dot\{p\}\_\{t\}\(k\)=μt​\(k\+1\)​pt​\(k\+1\)−μt​\(k\)​pt​\(k\)=ht​\[\(k\+1\)​pt​\(k\+1\)−k​pt​\(k\)\]\.\\displaystyle=\\mu\_\{t\}\(k\+1\)p\_\{t\}\(k\+1\)\-\\mu\_\{t\}\(k\)p\_\{t\}\(k\)=h\_\{t\}\\left\[\(k\+1\)p\_\{t\}\(k\+1\)\-kp\_\{t\}\(k\)\\right\]\.\(13\)Consider the probability generating function \(PGF\)G​\(z,t\):=𝔼​\[zYt\]=∑k=0Lpt​\(k\)​zkG\(z,t\):=\\mathbb\{E\}\[z^\{Y\_\{t\}\}\]=\\sum\_\{k=0\}^\{L\}p\_\{t\}\(k\)z^\{k\}\. We have

∂G∂t\\displaystyle\\frac\{\\partial G\}\{\\partial t\}=∑k=0Lp˙t​\(k\)​zk=ht​∑k=0L\[\(k\+1\)​pt​\(k\+1\)−k​pt​\(k\)\]​zk\\displaystyle=\\sum\_\{k=0\}^\{L\}\\dot\{p\}\_\{t\}\(k\)z^\{k\}=h\_\{t\}\\sum\_\{k=0\}^\{L\}\\left\[\(k\+1\)p\_\{t\}\(k\+1\)\-kp\_\{t\}\(k\)\\right\]z^\{k\}\(14\)=ht​\(∂G∂z−z​∂G∂z\)=ht​\(1−z\)​∂G∂z\.\\displaystyle=h\_\{t\}\\left\(\\frac\{\\partial G\}\{\\partial z\}\-z\\frac\{\\partial G\}\{\\partial z\}\\right\)=h\_\{t\}\(1\-z\)\\frac\{\\partial G\}\{\\partial z\}\.\(15\)One can verify that the function

G​\(z,t\)=\(κt\+z​\(1−κt\)\)LG\(z,t\)=\\left\(\\kappa\_\{t\}\+z\(1\-\\kappa\_\{t\}\)\\right\)^\{L\}\(16\)satisfies the above PDE with the initial conditionY0=LY\_\{0\}=L\. AsG​\(z,t\)G\(z,t\)is the standard PGF for the binomial distributionB​\(L,1−κt\)B\(L,1\-\\kappa\_\{t\}\), we haveYt∼B​\(L,1−κt\)Y\_\{t\}\\sim B\(L,1\-\\kappa\_\{t\}\)andXt∼B​\(L,κt\)X\_\{t\}\\sim B\(L,\\kappa\_\{t\}\)fort<τκt<\\tau\_\{\\kappa\}\. Ast→τκ−t\\to\\tau\_\{\\kappa\}^\{\-\},κt→1−\\kappa\_\{t\}\\to 1^\{\-\}and𝔼​\[Yt\]=L​\(1−κt\)→0\\mathbb\{E\}\[Y\_\{t\}\]=L\(1\-\\kappa\_\{t\}\)\\to 0\.

Moreover,Yt=L−XtY\_\{t\}=L\-X\_\{t\}is nonnegative and nonincreasing along every pure\-birth path, so its pathwise limit exists; the vanishing expectation implies that this limit is zero almost surely\. HenceXt→LX\_\{t\}\\to Lalmost surely\. DefiningLLas absorbing and setting the rate to zero fort≥τκt\\geq\\tau\_\{\\kappa\}yields the stated result without evaluating the hazard after saturation\. ∎

The general definition of the scheduler above allows us to use additional scheduler families without compromising theoretical validity\. In particular, the early\-insertion scheduleκt=min⁡\{1,t/τ\}\\kappa\_\{t\}=\\min\\\{1,t/\\tau\\\}hasht=1/\(τ−t\)h\_\{t\}=1/\(\\tau\-t\)fort<τt<\\tau; att=τt=\\tau, the process has reachedLLalmost surely, after which the state is absorbing, and the rate is defined to be zero\. Thus, the expressionκ˙t/\(1−κt\)\\dot\{\\kappa\}\_\{t\}/\(1\-\\kappa\_\{t\}\)is never evaluated in the saturated regime\. In the losses, the target insertion contribution is also set to zero fort≥τκt\\geq\\tau\_\{\\kappa\}\. Such a scheduler was already used in prior work, such as TDDM\[[9](https://arxiv.org/html/2607.09039#bib.bib9)\], but it was never rigorously formulated\.

### A\.2Proof for Theorem[1](https://arxiv.org/html/2607.09039#S3.E1)\(Marginal Rate\)

###### Proof\.

The Kolmogorov forward equation reads:

p˙t​\(k\|y\)=λt​\(k−1\|y\)​pt​\(k−1\|y\)−λt​\(k\|y\)​pt​\(k\|y\)\.\\dot\{p\}\_\{t\}\(k\|y\)=\\lambda\_\{t\}\(k\-1\|y\)p\_\{t\}\(k\-1\|y\)\-\\lambda\_\{t\}\(k\|y\)p\_\{t\}\(k\|y\)\.\(17\)For the marginalpt​\(k\):=∑yq​\(y\)​pt​\(k\|y\)p\_\{t\}\(k\):=\\sum\_\{y\}q\(y\)p\_\{t\}\(k\|y\), differentiating with respect to time, we obtain:

p˙t​\(k\)=∑yq​\(y\)​p˙t​\(k\|y\)\.\\dot\{p\}\_\{t\}\(k\)=\\sum\_\{y\}q\(y\)\\dot\{p\}\_\{t\}\(k\|y\)\.\(18\)Substituting the Kolmogorov equation, we get

p˙t​\(k\)=∑yq​\(y\)​λt​\(k−1\|y\)​pt​\(k−1\|y\)⏟event−∑yq​\(y\)​λt​\(k\|y\)​pt​\(k\|y\)⏟survival\.\\dot\{p\}\_\{t\}\(k\)=\\underbrace\{\\sum\_\{y\}q\(y\)\\lambda\_\{t\}\(k\-1\|y\)p\_\{t\}\(k\-1\|y\)\}\_\{\\text\{event\}\}\-\\underbrace\{\\sum\_\{y\}q\(y\)\\lambda\_\{t\}\(k\|y\)p\_\{t\}\(k\|y\)\}\_\{\\text\{survival\}\}\.\(19\)For everykkwithpt​\(k\)\>0p\_\{t\}\(k\)\>0, Bayes’ rule gives the posterior probability stated in Theorem[1](https://arxiv.org/html/2607.09039#S3.E1):

p​\(X1=y\|Xt=k\)=pt​\(k\|y\)​q​\(y\)pt​\(k\)\.p\(X\_\{1\}=y\|X\_\{t\}=k\)=\\frac\{p\_\{t\}\(k\|y\)q\(y\)\}\{p\_\{t\}\(k\)\}\.\(20\)Therefore, the marginal rate conditioned on the explicit current stateXt=kX\_\{t\}=kis

λt∗​\(k\):=𝔼X1∼p\(⋅\|Xt=k\)​\[λt​\(k\|X1\)\]=∑ypt​\(k\|y\)​q​\(y\)pt​\(k\)​λt​\(k\|y\)\.\\lambda\_\{t\}^\{\*\}\(k\):=\\mathbb\{E\}\_\{X\_\{1\}\\sim p\(\\cdot\|X\_\{t\}=k\)\}\[\\lambda\_\{t\}\(k\|X\_\{1\}\)\]=\\sum\_\{y\}\\frac\{p\_\{t\}\(k\|y\)q\(y\)\}\{p\_\{t\}\(k\)\}\\lambda\_\{t\}\(k\|y\)\.\(21\)Therefore,

λt∗​\(k\)​pt​\(k\)=∑yq​\(y\)​pt​\(k\|y\)​λt​\(k\|y\),\\lambda^\{\*\}\_\{t\}\(k\)p\_\{t\}\(k\)=\\sum\_\{y\}q\(y\)p\_\{t\}\(k\|y\)\\lambda\_\{t\}\(k\|y\),\(22\)which is the survival term in Eq\.[19](https://arxiv.org/html/2607.09039#A1.E19)\. Similarly, we have

λt∗​\(k−1\)​pt​\(k−1\)=∑yq​\(y\)​pt​\(k−1\|y\)​λt​\(k−1\|y\),\\lambda^\{\*\}\_\{t\}\(k\-1\)p\_\{t\}\(k\-1\)=\\sum\_\{y\}q\(y\)p\_\{t\}\(k\-1\|y\)\\lambda\_\{t\}\(k\-1\|y\),\(23\)which is the event term in Eq\.[19](https://arxiv.org/html/2607.09039#A1.E19)\. Combining both results:

p˙t​\(k\)=λt∗​\(k−1\)​pt​\(k−1\)−λt∗​\(k\)​pt​\(k\),\\dot\{p\}\_\{t\}\(k\)=\\lambda^\{\*\}\_\{t\}\(k\-1\)p\_\{t\}\(k\-1\)\-\\lambda^\{\*\}\_\{t\}\(k\)p\_\{t\}\(k\),\(24\)which is precisely the Kolmogorov forward equation for a generalized Poisson process driven by the rateλt∗​\(Xt\)\\lambda^\{\*\}\_\{t\}\(X\_\{t\}\)\. Sincep1​\(k\|y\)=𝟙​\{k=y\}p\_\{1\}\(k\|y\)=\\mathds\{1\}\\\{k=y\\\}, the terminal marginal isp1​\(k\)=q​\(k\)p\_\{1\}\(k\)=q\(k\)\. ∎

### A\.3Proposition[A\.3](https://arxiv.org/html/2607.09039#A1.SS3)\(Poisson NLL\)

\{proposition\}

\[store=ppnll\] For a target generalized Poisson process with rate functionλt∗\\lambda^\{\*\}\_\{t\}, consider a realization on\[0,1\]\[0,1\]with event times\{tk\}k=1L\\\{t\_\{k\}\\\}\_\{k=1\}^\{L\}, where0<tk≤10<t\_\{k\}\\leq 1\. Its negative log\-likelihood \(NLL\) under an estimated rateλt\\lambda\_\{t\}is:

ℒGP=∫01λt​d​t−∑k=1Llog⁡λtk\.\\mathcal\{L\}\_\{\\text\{GP\}\}=\\int\_\{0\}^\{1\}\\lambda\_\{t\}\\mathop\{\}\\\!\\mathrm\{d\}t\-\\sum\_\{k=1\}^\{L\}\\log\\lambda\_\{t\_\{k\}\}\.\(25\)
###### Proof\.

At each time steptt, the probability of an event isλt​d​t\\lambda\_\{t\}\\mathop\{\}\\\!\\mathrm\{d\}t, whereas the probability of survival is1−λt​d​t1\-\\lambda\_\{t\}\\mathop\{\}\\\!\\mathrm\{d\}t\. Therefore, we have

ℒGP\\displaystyle\\mathcal\{L\}\_\{\\text\{GP\}\}=−log⁡\(∏t∈no event\(1−λt​d​t\)⋅∏k=1Lλtk​d​t\)\\displaystyle=\-\\log\\left\(\\prod\_\{t\\in\{\\text\{no event\}\}\}\(1\-\\lambda\_\{t\}\\mathrm\{d\}t\)\\cdot\\prod\_\{k=1\}^\{L\}\\lambda\_\{t\_\{k\}\}\\mathrm\{d\}t\\right\)\(26\)=−log⁡\(∏t=01\(1−λt​d​t\)⋅∏k=1Lλtk\)\\displaystyle=\-\\log\\left\(\\prod\_\{t=0\}^\{1\}\(1\-\\lambda\_\{t\}\\mathrm\{d\}t\)\\cdot\\prod\_\{k=1\}^\{L\}\\lambda\_\{t\_\{k\}\}\\right\)\(27\)=∫01λt​dt−∑k=1Llog⁡λtk,\\displaystyle=\\int\_\{0\}^\{1\}\\lambda\_\{t\}\\mathrm\{d\}t\-\\sum\_\{k=1\}^\{L\}\\log\\lambda\_\{t\_\{k\}\},\(28\)which concludes the proof\. ∎

The identity used in Section[3\.1](https://arxiv.org/html/2607.09039#S3.SS1)to rewrite the event term of the Poisson NLL follows from the integration \(smoothing\) formula for the stochastic intensity of a point process\[[7](https://arxiv.org/html/2607.09039#bib.bib7)\]\. LetNtN\_\{t\}be a simple point process with event times\{tk\}\\\{t\_\{k\}\\\}, adapted to a filtration\{ℱt\}\\\{\\mathcal\{F\}\_\{t\}\\\}, that admits anℱt\\mathcal\{F\}\_\{t\}\-stochastic intensityλt∗\\lambda^\{\*\}\_\{t\}, i\.e\.,Nt−∫0tλs∗​d​sN\_\{t\}\-\\int\_\{0\}^\{t\}\\lambda^\{\*\}\_\{s\}\\mathop\{\}\\\!\\mathrm\{d\}sis anℱt\\mathcal\{F\}\_\{t\}\-martingale\. Then, for every nonnegativeℱt\\mathcal\{F\}\_\{t\}\-predictable processHtH\_\{t\}:

𝔼​\[∑kHtk\]=𝔼​\[∫01Ht​dNt\]=𝔼​\[∫01Ht​λt∗​d​t\],\\mathbb\{E\}\\left\[\\sum\_\{k\}H\_\{t\_\{k\}\}\\right\]=\\mathbb\{E\}\\left\[\\int\_\{0\}^\{1\}H\_\{t\}\\mathrm\{d\}N\_\{t\}\\right\]=\\mathbb\{E\}\\left\[\\int\_\{0\}^\{1\}H\_\{t\}\\lambda^\{\*\}\_\{t\}\\mathop\{\}\\\!\\mathrm\{d\}t\\right\],\(29\)a property that in fact characterizes the stochastic intensityλt∗\\lambda^\{\*\}\_\{t\}; the deterministic\-integrand case recovers Campbell’s theorem\[[14](https://arxiv.org/html/2607.09039#bib.bib14)\]\. By Theorem[1](https://arxiv.org/html/2607.09039#S3.E1), the marginal rateλt∗\\lambda^\{\*\}\_\{t\}is such a stochastic intensity \(the process conditioned onX1=LX\_\{1\}=Lis an inhomogeneous pure\-birth process with predictable compensator\[[6](https://arxiv.org/html/2607.09039#bib.bib6)\]\)\. Choosing the predictable integrandHt=−log⁡λt​\(Xt−\)H\_\{t\}=\-\\log\\lambda\_\{t\}\(X\_\{t^\{\-\}\}\)with the learned rateλt\\lambda\_\{t\}then yields:

𝔼\{tk\},L​\[−∑klog⁡λtk\]=𝔼​\[−∫01λt∗​log⁡λt​d​t\]=𝔼t​\[−λt∗​log⁡λt\],\\mathbb\{E\}\_\{\\\{t\_\{k\}\\\},L\}\\left\[\-\\sum\_\{k\}\\log\\lambda\_\{t\_\{k\}\}\\right\]=\\mathbb\{E\}\\left\[\-\\int\_\{0\}^\{1\}\\lambda^\{\*\}\_\{t\}\\log\\lambda\_\{t\}\\mathop\{\}\\\!\\mathrm\{d\}t\\right\]=\\mathbb\{E\}\_\{t\}\[\-\\lambda^\{\*\}\_\{t\}\\log\\lambda\_\{t\}\],\(30\)which is the identity in the main text\. In this way, the NLL loss objective is fully simulation\-free\.

### A\.4Proof for Theorem[3\.2](https://arxiv.org/html/2607.09039#S3.SS2)\(Multimodal GPFlow\)

We first describe a measure\-theoretic characterization of the variable\-length state space𝖸\\mathsf\{Y\}that ensures that defining a probability path is mathematically well\-posed\. Let\(𝖲,𝒮,ν\)\(\\mathsf\{S\},\\mathcal\{S\},\\nu\)be a component space with aσ\\sigma\-finite reference measure\. We use Lebesgue measure for Euclidean components, counting measure for categorical components, Riemannian volume for manifold\-valued components, and the corresponding product measure for mixed components\. Define

𝖸:=⨆k≥0\{k\}×𝖲k,μ:=∑k≥0δk⊗ν⊗k\.\\mathsf\{Y\}:=\\bigsqcup\_\{k\\geq 0\}\\\{k\\\}\\times\\mathsf\{S\}^\{k\},\\qquad\\mu:=\\sum\_\{k\\geq 0\}\\delta\_\{k\}\\otimes\\nu^\{\\otimes k\}\.\(31\)Thus, a probability law on variable\-length objects can be represented by a density or mass function with respect to the single reference measureμ\\mu\.

Fory=\(k,s\)∈𝖸y=\(k,s\)\\in\\mathsf\{Y\}, let𝖬​\(y\)\\mathsf\{M\}\(y\)be the insertion\-variable space and letI​\(y,m\)∈\{k\+1\}×𝖲k\+1I\(y,m\)\\in\\\{k\+1\\\}\\times\\mathsf\{S\}^\{k\+1\}be the measurable insertion map\. In the order\-agnostic case,mmcontains only the new component value; in the ordered case, it also contains an insertion slot\. For a clean targetY1=y1Y\_\{1\}=y\_\{1\}, letρt​\(d​m\|y,y1\)\\rho\_\{t\}\(\\mathop\{\}\\\!\\mathrm\{d\}m\|y,y\_\{1\}\)be a probability kernel on𝖬​\(y\)\\mathsf\{M\}\(y\)\. The associated conditional jump measure is

Qt​\(A\|y,y1\):=λt​\(y\|y1\)​∫𝖬​\(y\)𝟙​\{I​\(y,m\)∈A\}​ρt​\(d​m\|y,y1\)\.Q\_\{t\}\(A\|y,y\_\{1\}\):=\\lambda\_\{t\}\(y\|y\_\{1\}\)\\int\_\{\\mathsf\{M\}\(y\)\}\\mathds\{1\}\\\{I\(y,m\)\\in A\\\}\\rho\_\{t\}\(\\mathop\{\}\\\!\\mathrm\{d\}m\|y,y\_\{1\}\)\.\(32\)This equation makes the coupling explicit: an insertion value is drawn fromρt\(⋅\|y,y1\)\\rho\_\{t\}\(\\cdot\|y,y\_\{1\}\)only when an event governed byλt​\(y\|y1\)\\lambda\_\{t\}\(y\|y\_\{1\}\)occurs\. The total mass ofQt\(⋅\|y,y1\)Q\_\{t\}\(\\cdot\|y,y\_\{1\}\)is exactlyλt​\(y\|y1\)\\lambda\_\{t\}\(y\|y\_\{1\}\)\.

###### Proof\.

This proof follows the outline in transdimensional jump diffusion\[[9](https://arxiv.org/html/2607.09039#bib.bib9)\], while extending it to allow the inserted component to be continuous, discrete, Riemannian, or mixed\. Letℬty1\\mathcal\{B\}\_\{t\}^\{y\_\{1\}\}denote the conditional within\-length generator\. For every bounded test functionffin the domain of the generator, the conditional path satisfies the weak forward equation:

dd​t​∫𝖸f​\(y\)​pt​\(d​y\|y1\)=∫𝖸\[ℬty1​f​\(y\)\+∫𝖸\(f​\(y′\)−f​\(y\)\)​Qt​\(d​y′\|y,y1\)\]​pt​\(d​y\|y1\)\.\\frac\{\\mathop\{\}\\\!\\mathrm\{d\}\}\{\\mathop\{\}\\\!\\mathrm\{d\}t\}\\int\_\{\\mathsf\{Y\}\}f\(y\)p\_\{t\}\(\\mathop\{\}\\\!\\mathrm\{d\}y\|y\_\{1\}\)=\\int\_\{\\mathsf\{Y\}\}\\left\[\\mathcal\{B\}\_\{t\}^\{y\_\{1\}\}f\(y\)\+\\int\_\{\\mathsf\{Y\}\}\(f\(y^\{\\prime\}\)\-f\(y\)\)Q\_\{t\}\(\\mathop\{\}\\\!\\mathrm\{d\}y^\{\\prime\}\|y,y\_\{1\}\)\\right\]p\_\{t\}\(\\mathop\{\}\\\!\\mathrm\{d\}y\|y\_\{1\}\)\.\(33\)For brevity in the proof, denote the posterior law written explicitly in Theorem[3\.2](https://arxiv.org/html/2607.09039#S3.SS2)by

πt​\(d​y1\|y\):=p​\(d​y1\|Yt=y\)\.\\pi\_\{t\}\(\\mathop\{\}\\\!\\mathrm\{d\}y\_\{1\}\|y\):=p\(\\mathop\{\}\\\!\\mathrm\{d\}y\_\{1\}\|Y\_\{t\}=y\)\.\(34\)Equivalently, it is characterized by the disintegration identityq​\(d​y1\)​pt​\(d​y\|y1\)=pt​\(d​y\)​πt​\(d​y1\|y\)q\(\\mathop\{\}\\\!\\mathrm\{d\}y\_\{1\}\)p\_\{t\}\(\\mathop\{\}\\\!\\mathrm\{d\}y\|y\_\{1\}\)=p\_\{t\}\(\\mathop\{\}\\\!\\mathrm\{d\}y\)\\pi\_\{t\}\(\\mathop\{\}\\\!\\mathrm\{d\}y\_\{1\}\|y\)\. Thus, the three posterior expectations in the theorem are:

ℬt∗​f​\(y\)\\displaystyle\\mathcal\{B\}\_\{t\}^\{\*\}f\(y\)=∫ℬty1​f​\(y\)​πt​\(d​y1\|y\),\\displaystyle=\\int\\mathcal\{B\}\_\{t\}^\{y\_\{1\}\}f\(y\)\\pi\_\{t\}\(\\mathop\{\}\\\!\\mathrm\{d\}y\_\{1\}\|y\),\(35\)λt∗​\(y\)\\displaystyle\\lambda\_\{t\}^\{\*\}\(y\)=∫λt​\(y\|y1\)​πt​\(d​y1\|y\),\\displaystyle=\\int\\lambda\_\{t\}\(y\|y\_\{1\}\)\\pi\_\{t\}\(\\mathop\{\}\\\!\\mathrm\{d\}y\_\{1\}\|y\),\(36\)ρt∗​\(d​m\|y\)\\displaystyle\\rho\_\{t\}^\{\*\}\(\\mathop\{\}\\\!\\mathrm\{d\}m\|y\)=∫λt​\(y\|y1\)​ρt​\(d​m\|y,y1\)​πt​\(d​y1\|y\)∫λt​\(y\|y1\)​πt​\(d​y1\|y\)\.\\displaystyle=\\frac\{\\int\\lambda\_\{t\}\(y\|y\_\{1\}\)\\rho\_\{t\}\(\\mathop\{\}\\\!\\mathrm\{d\}m\|y,y\_\{1\}\)\\pi\_\{t\}\(\\mathop\{\}\\\!\\mathrm\{d\}y\_\{1\}\|y\)\}\{\\int\\lambda\_\{t\}\(y\|y\_\{1\}\)\\pi\_\{t\}\(\\mathop\{\}\\\!\\mathrm\{d\}y\_\{1\}\|y\)\}\.\(37\)Integrating Eq\.[33](https://arxiv.org/html/2607.09039#A1.E33)overq​\(d​y1\)q\(\\mathop\{\}\\\!\\mathrm\{d\}y\_\{1\}\)and applying the disintegration identity then yields the marginal jump measure

Qt∗​\(A\|y\):=∫Qt​\(A\|y,y1\)​πt​\(d​y1\|y\)=λt∗​\(y\)​∫𝖬​\(y\)𝟙​\{I​\(y,m\)∈A\}​ρt∗​\(d​m\|y\),\\displaystyle Q\_\{t\}^\{\*\}\(A\|y\)=\\int Q\_\{t\}\(A\|y,y\_\{1\}\)\\pi\_\{t\}\(\\mathop\{\}\\\!\\mathrm\{d\}y\_\{1\}\|y\)=\\lambda\_\{t\}^\{\*\}\(y\)\\int\_\{\\mathsf\{M\}\(y\)\}\\mathds\{1\}\\\{I\(y,m\)\\in A\\\}\\rho\_\{t\}^\{\*\}\(\\mathop\{\}\\\!\\mathrm\{d\}m\|y\),\(38\)whereλt∗\\lambda\_\{t\}^\{\*\}andρt∗\\rho\_\{t\}^\{\*\}are precisely those in Theorem[3\.2](https://arxiv.org/html/2607.09039#S3.SS2)\. Therefore,

dd​t​∫f​\(y\)​pt​\(d​y\)=∫\[ℬt∗​f​\(y\)\+∫\(f​\(y′\)−f​\(y\)\)​Qt∗​\(d​y′\|y\)\]​pt​\(d​y\),\\frac\{\\mathop\{\}\\\!\\mathrm\{d\}\}\{\\mathop\{\}\\\!\\mathrm\{d\}t\}\\int f\(y\)p\_\{t\}\(\\mathop\{\}\\\!\\mathrm\{d\}y\)=\\int\\left\[\\mathcal\{B\}\_\{t\}^\{\*\}f\(y\)\+\\int\(f\(y^\{\\prime\}\)\-f\(y\)\)Q\_\{t\}^\{\*\}\(\\mathop\{\}\\\!\\mathrm\{d\}y^\{\\prime\}\|y\)\\right\]p\_\{t\}\(\\mathop\{\}\\\!\\mathrm\{d\}y\),\(39\)which is the weak forward equation of the claimed marginal process\. Sincep1\(⋅\|y1\)=δy1p\_\{1\}\(\\cdot\|y\_\{1\}\)=\\delta\_\{y\_\{1\}\}forqq\-almost everyy1y\_\{1\}, its endpoint isp1=qp\_\{1\}=q\. ∎

For intuition, the corresponding forward equation can be written, as an identity of measures, as

∂tpt=ℬt∗†pt−λt∗pt\+∫𝖸pt\(dy¯\)Qt∗\(⋅\|y¯\)\.\\partial\_\{t\}p\_\{t\}=\\mathcal\{B\}\_\{t\}^\{\*\\dagger\}p\_\{t\}\-\\lambda\_\{t\}^\{\*\}p\_\{t\}\+\\int\_\{\\mathsf\{Y\}\}p\_\{t\}\(\\mathop\{\}\\\!\\mathrm\{d\}\\bar\{y\}\)Q\_\{t\}^\{\*\}\(\\cdot\|\\bar\{y\}\)\.\(40\)The second term is outgoing probability mass, whereas the third term is inflow from predecessor states; only the outgoing insertion measureQt∗=λt∗​ρt∗Q\_\{t\}^\{\*\}=\\lambda\_\{t\}^\{\*\}\\rho\_\{t\}^\{\*\}belongs to the generator at the current state\. When𝖲\\mathsf\{S\}is continuous,ℬt∗†​pt=−∇⋅\(vt∗​pt\)\\mathcal\{B\}\_\{t\}^\{\*\\dagger\}p\_\{t\}=\-\\nabla\\cdot\(v\_\{t\}^\{\*\}p\_\{t\}\)\(or its diffusion/Riemannian analog\), and integrals over inserted values are ordinary integrals\. When𝖲\\mathsf\{S\}is discrete, the same display is a master equation, all integrals over𝖲\\mathsf\{S\}become sums, andℬt∗\\mathcal\{B\}\_\{t\}^\{\*\}is a within\-length CTMC generator\. Hence, the proof applies to both cases without treating a categorical component as if it had a Euclidean density\.

### A\.5Theorem[A\.5](https://arxiv.org/html/2607.09039#A1.SS5)\(Order\-Preserving GPFlow\)

In the order\-preserving specialization,𝖬​\(y\)=\{0,…,k\}×𝖲\\mathsf\{M\}\(y\)=\\\{0,\\dots,k\\\}\\times\\mathsf\{S\}andm=\(i,a\)m=\(i,a\)\. The general insertion map becomesIi​\(y,a\):=I​\(y,m\)=I​\(y,\(i,a\)\)I\_\{i\}\(y,a\):=I\(y,m\)=I\(y,\(i,a\)\), whereIiI\_\{i\}places componentaain slotii\. Thus, the slot is part of the insertion variable whose rate\-weighted distribution is marginalized; the insertion map itself remains deterministic\. A detailed construction is provided in Appendix[B\.1](https://arxiv.org/html/2607.09039#A2.SS1)\.

\{theorem\}

\[store=order\] At a stateYt=y=\(k,s\)Y\_\{t\}=y=\(k,s\), let the insertion variable bem=\(i,a\)m=\(i,a\), wherei∈\{0,…,k\}i\\in\\\{0,\\dots,k\\\}is an insertion slot anda∈𝖲a\\in\\mathsf\{S\}is the new component\. Using the posterior lawp​\(d​y1\|Yt=y\)p\(\\mathop\{\}\\\!\\mathrm\{d\}y\_\{1\}\|Y\_\{t\}=y\)from Theorem[3\.2](https://arxiv.org/html/2607.09039#S3.SS2), define for each slot

\(λti\)∗​\(y\)\\displaystyle\(\\lambda\_\{t\}^\{i\}\)^\{\*\}\(y\):=𝔼Y1∼p\(⋅\|Yt=y\)​\[λti​\(y\|Y1\)\],\\displaystyle:=\\mathbb\{E\}\_\{Y\_\{1\}\\sim p\(\\cdot\|Y\_\{t\}=y\)\}\[\\lambda\_\{t\}^\{i\}\(y\|Y\_\{1\}\)\],\(41\)\(ρti\)∗​\(d​a\|y\)\\displaystyle\(\\rho\_\{t\}^\{i\}\)^\{\*\}\(\\mathop\{\}\\\!\\mathrm\{d\}a\|y\):=𝔼Y1∼p\(⋅\|Yt=y\)​\[λti​\(y\|Y1\)​ρti​\(d​a\|y,Y1\)\]𝔼Y1∼p\(⋅\|Yt=y\)​\[λti​\(y\|Y1\)\]\.\\displaystyle:=\\frac\{\\mathbb\{E\}\_\{Y\_\{1\}\\sim p\(\\cdot\|Y\_\{t\}=y\)\}\[\\lambda\_\{t\}^\{i\}\(y\|Y\_\{1\}\)\\rho\_\{t\}^\{i\}\(\\mathop\{\}\\\!\\mathrm\{d\}a\|y,Y\_\{1\}\)\]\}\{\\mathbb\{E\}\_\{Y\_\{1\}\\sim p\(\\cdot\|Y\_\{t\}=y\)\}\[\\lambda\_\{t\}^\{i\}\(y\|Y\_\{1\}\)\]\}\.\(42\)Together with the marginal within\-length generatorℬt∗\\mathcal\{B\}\_\{t\}^\{\*\}from Theorem[3\.2](https://arxiv.org/html/2607.09039#S3.SS2), the insertion jump measure

Qt∗​\(d​y′\|y\)=∑i=0k\(λti\)∗​\(y\)​∫𝖲δIi​\(y,a\)​\(d​y′\)​\(ρti\)∗​\(d​a\|y\)Q\_\{t\}^\{\*\}\(\\mathop\{\}\\\!\\mathrm\{d\}y^\{\\prime\}\|y\)=\\sum\_\{i=0\}^\{k\}\(\\lambda\_\{t\}^\{i\}\)^\{\*\}\(y\)\\int\_\{\\mathsf\{S\}\}\\delta\_\{I\_\{i\}\(y,a\)\}\(\\mathop\{\}\\\!\\mathrm\{d\}y^\{\\prime\}\)\(\\rho\_\{t\}^\{i\}\)^\{\*\}\(\\mathop\{\}\\\!\\mathrm\{d\}a\|y\)\(43\)generates the ordered target distributionY1=\(X1,S1\)∼qY\_\{1\}=\(X\_\{1\},S\_\{1\}\)\\sim q, whereS1=\(S11,…,S1X1\)S\_\{1\}=\(S\_\{1\}^\{1\},\\dots,S\_\{1\}^\{X\_\{1\}\}\)\. As before,\(ρti\)∗\(\\rho\_\{t\}^\{i\}\)^\{\*\}may be chosen arbitrarily if\(λti\)∗​\(y\)=0\(\\lambda\_\{t\}^\{i\}\)^\{\*\}\(y\)=0\.

###### Proof\.

The order\-preserving process is the specialization of Theorem[3\.2](https://arxiv.org/html/2607.09039#S3.SS2)in which the insertion variable ism=\(i,a\)m=\(i,a\)andI​\(y,\(i,a\)\)=Ii​\(y,a\)I\(y,\(i,a\)\)=I\_\{i\}\(y,a\)\. Its outgoing insertion jump measure is therefore:

Qt∗​\(d​y′\|y\)=∑i=0k\(λti\)∗​\(y\)​∫𝖲δIi​\(y,a\)​\(d​y′\)​\(ρti\)∗​\(d​a\|y\)\.Q\_\{t\}^\{\*\}\(\\mathop\{\}\\\!\\mathrm\{d\}y^\{\\prime\}\|y\)=\\sum\_\{i=0\}^\{k\}\(\\lambda\_\{t\}^\{i\}\)^\{\*\}\(y\)\\int\_\{\\mathsf\{S\}\}\\delta\_\{I\_\{i\}\(y,a\)\}\(\\mathop\{\}\\\!\\mathrm\{d\}y^\{\\prime\}\)\(\\rho\_\{t\}^\{i\}\)^\{\*\}\(\\mathop\{\}\\\!\\mathrm\{d\}a\|y\)\.\(44\)For a clean target of lengthLLand current lengthkk, the thinning path partitions theL−kL\-kremaining target components into insertion binsΩ0,…,Ωk\\Omega\_\{0\},\\dots,\\Omega\_\{k\}\(see Appendix[B\.1](https://arxiv.org/html/2607.09039#A2.SS1)for construction\)\. The conditional slot rate is\|Ωi\|​ht\|\\Omega\_\{i\}\|h\_\{t\}, hence the total conditional rate is:

∑i=0kλti​\(k\|L\)=\(L−k\)​ht,\\sum\_\{i=0\}^\{k\}\\lambda\_\{t\}^\{i\}\(k\|L\)=\(L\-k\)h\_\{t\},\(45\)which recovers the conditional rate in Theorem[3\.1](https://arxiv.org/html/2607.09039#S3.SS1)\. Posterior marginalization of each slot\-specific insertion measure gives exactly the rates and insertion kernels in Theorem[A\.5](https://arxiv.org/html/2607.09039#A1.SS5)\. Therefore, the terminal distribution is the ordered target distributionqq\. ∎

The general weak\-forward argument from Appendix[A\.4](https://arxiv.org/html/2607.09039#A1.SS4)then yields

∂tpt=ℬt∗†pt−pt∑i=0k\(λti\)∗\+∫𝖸pt\(dy¯\)Qt∗\(⋅\|y¯\),\\partial\_\{t\}p\_\{t\}=\\mathcal\{B\}\_\{t\}^\{\*\\dagger\}p\_\{t\}\-p\_\{t\}\\sum\_\{i=0\}^\{k\}\(\\lambda\_\{t\}^\{i\}\)^\{\*\}\+\\int\_\{\\mathsf\{Y\}\}p\_\{t\}\(\\mathop\{\}\\\!\\mathrm\{d\}\\bar\{y\}\)Q\_\{t\}^\{\*\}\(\\cdot\|\\bar\{y\}\),\(46\)and the insertion mapsIiI\_\{i\}preserve the clean\-target ordering by construction\.

### A\.6Proof for Theorem[4\.1](https://arxiv.org/html/2607.09039#S4.SS1)\(KL Bound\)

###### Proof\.

We first discuss one sufficient set of regularity conditions\. These conditions ensure that the Radon\-Nikodym derivative and every expression below are well\-defined\.

Both processes start from the same distributionp0p\_\{0\}\. For everyt∈\[0,T\]t\\in\[0,T\],ptp\_\{t\}andp~t\\tilde\{p\}\_\{t\}are strictly positive densities with respect to a commonσ\\sigma\-finite reference measure; they are differentiable in time and twice differentiable in every continuous coordinate\. The densities, their derivatives, and all generator terms are integrable so that differentiation may be exchanged with integration, the generator identities may be applied, and the continuous\-coordinate integration\-by\-parts boundary terms vanish\. On every continuous component, the two processes have the same measurable diffusion coefficientσt2​I\\sigma\_\{t\}^\{2\}Iwithσt\>0\\sigma\_\{t\}\>0, and the time integral of𝔼pt​\[‖ut−u~t‖2/σt2\]\\mathbb\{E\}\_\{p\_\{t\}\}\[\\\|u\_\{t\}\-\\tilde\{u\}\_\{t\}\\\|^\{2\}/\\sigma\_\{t\}^\{2\}\]is finite\. The jump measures are finite kernels,Qt\(⋅\|x\)Q\_\{t\}\(\\cdot\|x\)is absolutely continuous with respect toQ~t\(⋅\|x\)\\tilde\{Q\}\_\{t\}\(\\cdot\|x\)ford​t​pt​\(d​x\)\\mathop\{\}\\\!\\mathrm\{d\}t\\,p\_\{t\}\(\\mathop\{\}\\\!\\mathrm\{d\}x\)\-almost every\(t,x\)\(t,x\), and the generalized KL term is measurable and integrable over time andptp\_\{t\}\. If no continuous component is present, the regularity conditions are unnecessary\.

The proof uses the generator identity together with completing the square for the continuous component and the nonnegativity ofz​log⁡z−z\+1z\\log z\-z\+1for the jump component\. For readability, we suppress the time subscript onQtQ\_\{t\}andQ~t\\tilde\{Q\}\_\{t\}below\. Consider the time derivative of the KL divergence:

dd​t​𝒟KL​\(pt∥p~t\)\\displaystyle\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}\\mathcal\{D\}\_\{\\rm KL\}\(p\_\{t\}\\\|\\tilde\{p\}\_\{t\}\)=dd​t​∫pt​\(x\)​log⁡pt​\(x\)p~t​\(x\)​d​x\\displaystyle=\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}\\int p\_\{t\}\(x\)\\log\\frac\{p\_\{t\}\(x\)\}\{\\tilde\{p\}\_\{t\}\(x\)\}\\,\\mathrm\{d\}x\(47\)=∫p˙t​\(x\)​log⁡pt​\(x\)p~t​\(x\)​d​x\+∫pt​\(x\)​p˙t​\(x\)pt​\(x\)​dx⏟=0−∫d​p~td​t​\(x\)​pt​\(x\)p~t​\(x\)​dx\\displaystyle=\\int\\dot\{p\}\_\{t\}\(x\)\\log\\frac\{p\_\{t\}\(x\)\}\{\\tilde\{p\}\_\{t\}\(x\)\}\\,\\mathrm\{d\}x\+\\underset\{=0\}\{\\underbrace\{\\int p\_\{t\}\(x\)\\frac\{\\dot\{p\}\_\{t\}\(x\)\}\{p\_\{t\}\(x\)\}\\,\\mathrm\{d\}x\}\}\-\\int\\frac\{\\mathrm\{d\}\\tilde\{p\}\_\{t\}\}\{\\mathrm\{d\}t\}\(x\)\\frac\{p\_\{t\}\(x\)\}\{\\tilde\{p\}\_\{t\}\(x\)\}\\,\\mathrm\{d\}x\(48\)=𝔼x∼pt​\[𝒜t​\(log⁡ptp~t\)​\(x\)\]−𝔼x∼p~t​\[𝒜~t​\(ptp~t\)​\(x\)\]\\displaystyle=\\mathbb\{E\}\_\{x\\sim p\_\{t\}\}\\left\[\\mathcal\{A\}\_\{t\}\\left\(\\log\\frac\{p\_\{t\}\}\{\\tilde\{p\}\_\{t\}\}\\right\)\(x\)\\right\]\-\\mathbb\{E\}\_\{x\\sim\\tilde\{p\}\_\{t\}\}\\left\[\\tilde\{\\mathcal\{A\}\}\_\{t\}\\left\(\\frac\{p\_\{t\}\}\{\\tilde\{p\}\_\{t\}\}\\right\)\(x\)\\right\]\(49\)=𝔼x∼pt​\[𝒜t​\(log⁡ptp~t\)​\(x\)−p~t​\(x\)pt​\(x\)​𝒜~t​\(ptp~t\)​\(x\)\]\.\\displaystyle=\\mathbb\{E\}\_\{x\\sim p\_\{t\}\}\\left\[\\mathcal\{A\}\_\{t\}\\left\(\\log\\frac\{p\_\{t\}\}\{\\tilde\{p\}\_\{t\}\}\\right\)\(x\)\-\\frac\{\\tilde\{p\}\_\{t\}\(x\)\}\{p\_\{t\}\(x\)\}\\tilde\{\\mathcal\{A\}\}\_\{t\}\\left\(\\frac\{p\_\{t\}\}\{\\tilde\{p\}\_\{t\}\}\\right\)\(x\)\\right\]\.\(50\)
Below, for simplicity, we denotef=pt/p~tf=p\_\{t\}/\\tilde\{p\}\_\{t\}\. Then

dd​t​𝒟KL​\(pt∥p~t\)\\displaystyle\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}\\mathcal\{D\}\_\{\\rm KL\}\(p\_\{t\}\\\|\\tilde\{p\}\_\{t\}\)\(51\)=dd​t​∫pt​\(x\)​log⁡pt​\(x\)p~t​\(x\)​d​x=𝔼x∼pt​\[𝒜t​\(log⁡f\)​\(x\)−1f​\(x\)​𝒜~t​f​\(x\)\]\\displaystyle=\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}\\int p\_\{t\}\(x\)\\log\\frac\{p\_\{t\}\(x\)\}\{\\tilde\{p\}\_\{t\}\(x\)\}\\,\\mathrm\{d\}x=\\mathbb\{E\}\_\{x\\sim p\_\{t\}\}\\left\[\\mathcal\{A\}\_\{t\}\(\\log f\)\(x\)\-\\frac\{1\}\{f\(x\)\}\\tilde\{\\mathcal\{A\}\}\_\{t\}f\(x\)\\right\]\(52\)=𝔼x∼pt\[ut​\(x\)⋅∇log⁡f​\(x\)\+12​σt2​∇⋅∇log⁡f​\(x\)\+∫\[log⁡f​\(y\)−log⁡f​\(x\)\]​Q​\(d​y\|x\)−1f​\(x\)\(u~t\(x\)⋅∇f\(x\)\+12σt2∇⋅∇f\(x\)\+∫\[f\(y\)−f\(x\)\]Q~\(dy\|x\)\)\]\.\\displaystyle=\\begin\{aligned\} \\mathbb\{E\}\_\{x\\sim p\_\{t\}\}\\Bigg\[&u\_\{t\}\(x\)\\cdot\\nabla\\log f\(x\)\+\\frac\{1\}\{2\}\\sigma\_\{t\}^\{2\}\\nabla\\cdot\\nabla\\log f\(x\)\+\\int\[\\log f\(y\)\-\\log f\(x\)\]Q\(\\mathrm\{d\}y\|x\)\\\\ &\-\\frac\{1\}\{f\(x\)\}\\left\(\\tilde\{u\}\_\{t\}\(x\)\\cdot\\nabla f\(x\)\+\\frac\{1\}\{2\}\\sigma\_\{t\}^\{2\}\\nabla\\cdot\\nabla f\(x\)\+\\int\[f\(y\)\-f\(x\)\]\\tilde\{Q\}\(\\mathrm\{d\}y\|x\)\\right\)\\Bigg\]\.\\end\{aligned\}\(53\)Substitute∇f​\(x\)/f​\(x\)=∇log⁡f​\(x\)\\nabla f\(x\)/f\(x\)=\\nabla\\log f\(x\)and∇⋅∇f​\(x\)/f​\(x\)=∇⋅∇log⁡f​\(x\)\+‖∇log⁡f​\(x\)‖2\\nabla\\cdot\\nabla f\(x\)/f\(x\)=\\nabla\\cdot\\nabla\\log f\(x\)\+\\\|\\nabla\\log f\(x\)\\\|^\{2\}in Eq\.[53](https://arxiv.org/html/2607.09039#A1.E53)to get:

dd​t​𝒟KL​\(pt∥p~t\)\\displaystyle\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}\\mathcal\{D\}\_\{\\rm KL\}\(p\_\{t\}\\\|\\tilde\{p\}\_\{t\}\)\(54\)=𝔼x∼pt\[ut​\(x\)⋅∇log⁡f​\(x\)\+12​σt2​∇⋅∇log⁡f​\(x\)\+∫\[log⁡f​\(y\)−log⁡f​\(x\)\]​Q​\(d​y\|x\)−u~t\(x\)⋅∇f​\(x\)f​\(x\)−12σt2∇⋅∇f​\(x\)f​\(x\)−∫\(f​\(y\)f​\(x\)−1\)Q~\(dy\|x\)\]\\displaystyle=\\begin\{aligned\} \\mathbb\{E\}\_\{x\\sim p\_\{t\}\}\\Bigg\[&u\_\{t\}\(x\)\\cdot\\nabla\\log f\(x\)\+\\frac\{1\}\{2\}\\sigma\_\{t\}^\{2\}\\nabla\\cdot\\nabla\\log f\(x\)\+\\int\[\\log f\(y\)\-\\log f\(x\)\]Q\(\\mathrm\{d\}y\|x\)\\\\ &\-\\tilde\{u\}\_\{t\}\(x\)\\cdot\\frac\{\\nabla f\(x\)\}\{f\(x\)\}\-\\frac\{1\}\{2\}\\sigma\_\{t\}^\{2\}\\frac\{\\nabla\\cdot\\nabla f\(x\)\}\{f\(x\)\}\-\\int\\left\(\\frac\{f\(y\)\}\{f\(x\)\}\-1\\right\)\\tilde\{Q\}\(\\mathrm\{d\}y\|x\)\\Bigg\]\\end\{aligned\}\(55\)=𝔼x∼pt\[\(ut​\(x\)−u~t​\(x\)\)⋅∇log⁡f​\(x\)−12​σt2​‖∇log⁡f​\(x\)‖2\+∫\(logf​\(y\)f​\(x\)\)Q\(dy\|x\)\+∫\(1−f​\(y\)f​\(x\)\)Q~\(dy\|x\)\]\.\\displaystyle=\\begin\{aligned\} \\mathbb\{E\}\_\{x\\sim p\_\{t\}\}\\Bigg\[&\(u\_\{t\}\(x\)\-\\tilde\{u\}\_\{t\}\(x\)\)\\cdot\\nabla\\log f\(x\)\-\\frac\{1\}\{2\}\\sigma\_\{t\}^\{2\}\\\|\\nabla\\log f\(x\)\\\|^\{2\}\\\\ &\+\\int\\left\(\\log\\frac\{f\(y\)\}\{f\(x\)\}\\right\)Q\(\\mathrm\{d\}y\|x\)\+\\int\\left\(1\-\\frac\{f\(y\)\}\{f\(x\)\}\\right\)\\tilde\{Q\}\(\\mathrm\{d\}y\|x\)\\Bigg\]\.\\end\{aligned\}\(56\)
Note that

\(ut​\(x\)−u~t​\(x\)\)⋅∇log⁡f​\(x\)−σt22​‖∇log⁡f​\(x\)‖2\\displaystyle\(u\_\{t\}\(x\)\-\\tilde\{u\}\_\{t\}\(x\)\)\\cdot\\nabla\\log f\(x\)\-\\frac\{\\sigma\_\{t\}^\{2\}\}\{2\}\\\|\\nabla\\log f\(x\)\\\|^\{2\}\(57\)=‖ut​\(x\)−u~t​\(x\)‖22​σt2−σt22​‖∇log⁡f​\(x\)−ut​\(x\)−u~t​\(x\)σt2‖2≤‖ut​\(x\)−u~t​\(x\)‖22​σt2,\\displaystyle=\\frac\{\\\|u\_\{t\}\(x\)\-\\tilde\{u\}\_\{t\}\(x\)\\\|^\{2\}\}\{2\\sigma\_\{t\}^\{2\}\}\-\\frac\{\\sigma\_\{t\}^\{2\}\}\{2\}\\left\\\|\\nabla\\log f\(x\)\-\\frac\{u\_\{t\}\(x\)\-\\tilde\{u\}\_\{t\}\(x\)\}\{\\sigma\_\{t\}^\{2\}\}\\right\\\|^\{2\}\\leq\\frac\{\\\|u\_\{t\}\(x\)\-\\tilde\{u\}\_\{t\}\(x\)\\\|^\{2\}\}\{2\\sigma\_\{t\}^\{2\}\},\(58\)For the jump terms, definea​\(y\):=f​\(y\)/f​\(x\)a\(y\):=f\(y\)/f\(x\)andr​\(y\):=d​Q​\(y\|x\)/d​Q~​\(y\|x\)r\(y\):=\\mathrm\{d\}Q\(y\|x\)/\\mathrm\{d\}\\tilde\{Q\}\(y\|x\)\. The pointwise inequalityr​log⁡a\+1−a≤r​log⁡r−r\+1r\\log a\+1\-a\\leq r\\log r\-r\+1holds because the difference between the right\- and left\-hand sides isa​\[\(r/a\)​log⁡\(r/a\)−r/a\+1\]≥0a\[\(r/a\)\\log\(r/a\)\-r/a\+1\]\\geq 0\. Consequently,

∫log⁡f​\(y\)f​\(x\)​Q​\(d​y\|x\)\+∫\(1−f​\(y\)f​\(x\)\)​Q~​\(d​y\|x\)\\displaystyle\\int\\log\\frac\{f\(y\)\}\{f\(x\)\}Q\(\\mathrm\{d\}y\|x\)\+\\int\\left\(1\-\\frac\{f\(y\)\}\{f\(x\)\}\\right\)\\tilde\{Q\}\(\\mathrm\{d\}y\|x\)\(59\)≤∫\(rlogr−r\+1\)Q~\(dy\|x\)=𝒟KL\(Q\(⋅\|x\)∥Q~\(⋅\|x\)\)\.\\displaystyle\\qquad\\leq\\int\\left\(r\\log r\-r\+1\\right\)\\tilde\{Q\}\(\\mathrm\{d\}y\|x\)=\\mathcal\{D\}\_\{\\text\{KL\}\}\(Q\(\\cdot\|x\)\\\|\\tilde\{Q\}\(\\cdot\|x\)\)\.\(60\)Combining the continuous and jump bounds, we conclude:

dd​t​𝒟KL​\(pt∥p~t\)\\displaystyle\\frac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}\\mathcal\{D\}\_\{\\rm KL\}\(p\_\{t\}\\\|\\tilde\{p\}\_\{t\}\)≤𝔼x∼pt​\[‖ut​\(x\)−u~t​\(x\)‖22​σt2\+∫\(log⁡d​Q​\(y\|x\)d​Q~​\(y\|x\)−1\)​Q​\(d​y\|x\)\+∫Q~​\(d​y\|x\)\],\\displaystyle\\leq\\mathbb\{E\}\_\{x\\sim p\_\{t\}\}\\left\[\\frac\{\\\|u\_\{t\}\(x\)\-\\tilde\{u\}\_\{t\}\(x\)\\\|^\{2\}\}\{2\\sigma\_\{t\}^\{2\}\}\+\\int\\left\(\\log\\frac\{\\mathrm\{d\}Q\(y\|x\)\}\{\\mathrm\{d\}\\tilde\{Q\}\(y\|x\)\}\-1\\right\)Q\(\\mathrm\{d\}y\|x\)\+\\int\\tilde\{Q\}\(\\mathrm\{d\}y\|x\)\\right\],\(61\)which completes the proof by integration\. If no continuous component is present, the completion\-of\-the\-square step and its resulting drift\-diffusion term are absent, leaving only the jump\-measure term\. ∎

### A\.7Generator for GPFlow

Table 4:Generators for each variant of GPFlow\.The continuity equations in Eq\.[24](https://arxiv.org/html/2607.09039#A1.E24),[40](https://arxiv.org/html/2607.09039#A1.E40), and[46](https://arxiv.org/html/2607.09039#A1.E46)describe the evolution of the three GPFlow variants\. Table[4](https://arxiv.org/html/2607.09039#A1.T4)lists the corresponding*outgoing*generators\. It is important to distinguish the jump measure from its appearance in the forward equation\. At a current stateyy, the generator contains one outgoing insertion measure:

Qt​\(d​y′\|y\)=λt​\(y\)​∫δI​\(y,m\)​\(d​y′\)​ρt​\(d​m\|y\)\.Q\_\{t\}\(\\mathop\{\}\\\!\\mathrm\{d\}y^\{\\prime\}\|y\)=\\lambda\_\{t\}\(y\)\\int\\delta\_\{I\(y,m\)\}\(\\mathop\{\}\\\!\\mathrm\{d\}y^\{\\prime\}\)\\rho\_\{t\}\(\\mathop\{\}\\\!\\mathrm\{d\}m\|y\)\.\(62\)Its total mass isλt​\(y\)\\lambda\_\{t\}\(y\)\. Consequently, the adjoint forward equation contains an outgoing term−λt​\(y\)​pt​\(y\)\-\\lambda\_\{t\}\(y\)p\_\{t\}\(y\)and an incoming term obtained by integratingpt​\(y¯\)​Qt​\(d​y\|y¯\)p\_\{t\}\(\\bar\{y\}\)Q\_\{t\}\(\\mathop\{\}\\\!\\mathrm\{d\}y\|\\bar\{y\}\)over predecessor statesy¯\\bar\{y\}\. The predecessor inflow is not a second jump measure atyy\. For length\-only GPFlow, Eq\.[62](https://arxiv.org/html/2607.09039#A1.E62)reduces toQt​\(d​k′\|k\)=λt​\(k\)​δk\+1​\(d​k′\)Q\_\{t\}\(\\mathop\{\}\\\!\\mathrm\{d\}k^\{\\prime\}\|k\)=\\lambda\_\{t\}\(k\)\\delta\_\{k\+1\}\(\\mathop\{\}\\\!\\mathrm\{d\}k^\{\\prime\}\)\.

This also clarifies the relation to a general generator\-matching formulation\. A general jump measure may move arbitrary parts of an existing continuous state\. In GPFlow, its support is restricted to length\-increasing successorsI​\(y,m\)I\(y,m\): the jump adds a new component sampled fromρt\\rho\_\{t\}, while existing components evolve between jumps through the within\-length generatorℬt\\mathcal\{B\}\_\{t\}\. Thus,λt\\lambda\_\{t\}andρt\\rho\_\{t\}are always coupled as the rate and normalized insertion kernel of the same event\.

LetQt=λt​ρtQ\_\{t\}=\\lambda\_\{t\}\\rho\_\{t\}andQ~t=λ~t​ρ~t\\tilde\{Q\}\_\{t\}=\\tilde\{\\lambda\}\_\{t\}\\tilde\{\\rho\}\_\{t\}, suppressing the deterministic insertion map for clarity\. Becauseρt\\rho\_\{t\}andρ~t\\tilde\{\\rho\}\_\{t\}are probability kernels, the generalized KL divergence between the positive jump measures decomposes exactly as:

𝒟KL​\(Qt∥Q~t\)\\displaystyle\\mathcal\{D\}\_\{\\text\{KL\}\}\(Q\_\{t\}\\\|\\tilde\{Q\}\_\{t\}\)=λ~t−λt\+λt​log⁡λtλ~t\+λt​∫ρt​\(d​m\)​log⁡d​ρtd​ρ~t​\(m\)\\displaystyle=\\tilde\{\\lambda\}\_\{t\}\-\\lambda\_\{t\}\+\\lambda\_\{t\}\\log\\frac\{\\lambda\_\{t\}\}\{\\tilde\{\\lambda\}\_\{t\}\}\+\\lambda\_\{t\}\\int\\rho\_\{t\}\(\\mathop\{\}\\\!\\mathrm\{d\}m\)\\log\\frac\{\\mathop\{\}\\\!\\mathrm\{d\}\\rho\_\{t\}\}\{\\mathop\{\}\\\!\\mathrm\{d\}\\tilde\{\\rho\}\_\{t\}\}\(m\)\(63\)=𝒟KL​\(λt∥λ~t\)\+λt​𝒟KL​\(ρt∥ρ~t\)\.\\displaystyle=\\mathcal\{D\}\_\{\\text\{KL\}\}\(\\lambda\_\{t\}\\\|\\tilde\{\\lambda\}\_\{t\}\)\+\\lambda\_\{t\}\\mathcal\{D\}\_\{\\text\{KL\}\}\(\\rho\_\{t\}\\\|\\tilde\{\\rho\}\_\{t\}\)\.\(64\)There is therefore no double counting: the Poisson rate NLL supplies the first term, and the insertion\-kernel NLL supplies the second\. With exact NLL parameterizations, their sum is the jump\-measure term in Theorem[4\.1](https://arxiv.org/html/2607.09039#S4.SS1)\. The mean\-regression approximation used for continuous insertion values is a practical surrogate\.

### A\.8Connection to EditFlow

In this section, we demonstrate the connection between GPFlow, when instantiated only on the discrete modality, and existing work such as EditFlow\[[20](https://arxiv.org/html/2607.09039#bib.bib20)\]\. We first note that GPFlow can adopt the “deletion” operation by extending the generalized Poisson process to an inhomogeneous birth\-death process\. For simplicity, we now consider EditFlow with only insertion and substitution\. Recall that the EditFlow is a variable\-length discrete flow matching model whose loss is defined as:

ℒedit=𝔼t,St,Xt\[∑k=1Xtut\(⋅\|St\)−∑k=1neditκ˙t1−κtlogut\(S1\|St\)\],\\mathcal\{L\}\_\{\\text\{edit\}\}=\\mathbb\{E\}\_\{t,S\_\{t\},X\_\{t\}\}\\left\[\\sum\_\{k=1\}^\{X\_\{t\}\}u\_\{t\}\(\\cdot\|S\_\{t\}\)\-\\sum\_\{k=1\}^\{n\_\{\\text\{edit\}\}\}\\frac\{\\dot\{\\kappa\}\_\{t\}\}\{1\-\\kappa\_\{t\}\}\\log u\_\{t\}\(S\_\{1\}\|S\_\{t\}\)\\right\],\(65\)wherenedit≥X1−Xtn\_\{\\text\{edit\}\}\\geq X\_\{1\}\-X\_\{t\}is the number of edit operations, including insertions \(nins=X1−Xtn\_\{\\text\{ins\}\}=X\_\{1\}\-X\_\{t\}\) and substitutions \(nsub=nedit−ninsn\_\{\\text\{sub\}\}=n\_\{\\text\{edit\}\}\-n\_\{\\text\{ins\}\}\)\. EditFlow predicts the token\-wise mixed\-type “vector field”uθ​\(St,t\)u\_\{\\theta\}\(S\_\{t\},t\), which is parameterized as:

uθins​\(St,t\)\\displaystyle u^\{\\text\{ins\}\}\_\{\\theta\}\(S\_\{t\},t\):=λθins​\(St,t\)​Qθins​\(s\|St,t\),\\displaystyle:=\\lambda^\{\\text\{ins\}\}\_\{\\theta\}\(S\_\{t\},t\)Q\_\{\\theta\}^\{\\text\{ins\}\}\(s\|S\_\{t\},t\),\(66\)uθsub​\(St,t\)\\displaystyle u^\{\\text\{sub\}\}\_\{\\theta\}\(S\_\{t\},t\):=λθsub​\(St,t\)​Qθsub​\(s,s′\|St,t\),\\displaystyle:=\\lambda^\{\\text\{sub\}\}\_\{\\theta\}\(S\_\{t\},t\)Q\_\{\\theta\}^\{\\text\{sub\}\}\(s,s^\{\\prime\}\|S\_\{t\},t\),\(67\)whereλθ\\lambda\_\{\\theta\}is the parameterized insertion/substitution rate,QθinsQ\_\{\\theta\}^\{\\text\{ins\}\}is the Markov chain transition matrix representation of the probability of inserting a new tokenss, andQθsubQ\_\{\\theta\}^\{\\text\{sub\}\}is the probability of substituting the old tokens′s^\{\\prime\}with a new tokenss\. We can first decompose the loss into the insertion and substitution parts as:

ℒedit=ℒins\+ℒsub=𝔼t,St,Xt​\[∑k=1Xtutins−∑k=1ninsκ˙t1−κt​log⁡utins​\(s1\)\]\+𝔼t,St,Xt​\[∑k=1Xtutsub−∑k=1nsubκ˙t1−κt​log⁡utsub​\(s1,s′\)\]\.\\mathcal\{L\}\_\{\\text\{edit\}\}=\\mathcal\{L\}\_\{\\text\{ins\}\}\+\\mathcal\{L\}\_\{\\text\{sub\}\}=\\;\\begin\{aligned\} &\\mathbb\{E\}\_\{t,S\_\{t\},X\_\{t\}\}\\left\[\\sum\_\{k=1\}^\{X\_\{t\}\}u^\{\\text\{ins\}\}\_\{t\}\-\\sum\_\{k=1\}^\{n\_\{\\text\{ins\}\}\}\\frac\{\\dot\{\\kappa\}\_\{t\}\}\{1\-\\kappa\_\{t\}\}\\log u^\{\\text\{ins\}\}\_\{t\}\(s\_\{1\}\)\\right\]\\\\ &\+\\mathbb\{E\}\_\{t,S\_\{t\},X\_\{t\}\}\\left\[\\sum\_\{k=1\}^\{X\_\{t\}\}u^\{\\text\{sub\}\}\_\{t\}\-\\sum\_\{k=1\}^\{n\_\{\\text\{sub\}\}\}\\frac\{\\dot\{\\kappa\}\_\{t\}\}\{1\-\\kappa\_\{t\}\}\\log u^\{\\text\{sub\}\}\_\{t\}\(s\_\{1\},s^\{\\prime\}\)\\right\]\.\\end\{aligned\}\(68\)We first note that for bothQθQ\_\{\\theta\}, the distribution property requires∑sQ​\(s\|⋅\)=1\\sum\_\{s\}Q\(s\|\\cdot\)=1\. In this way, the first term in Eq\.[65](https://arxiv.org/html/2607.09039#A1.E65)reduces to∑λθ\\sum\\lambda\_\{\\theta\}, the total rate described in the order\-preserving GPFlow in Appendix[B\.1](https://arxiv.org/html/2607.09039#A2.SS1)\. Usinglog⁡uθ=log⁡λθ\+log⁡Qθ\\log u\_\{\\theta\}=\\log\\lambda\_\{\\theta\}\+\\log Q\_\{\\theta\}, we have:

ℒ=∑k=1Xtλ−∑k=1nκ˙t1−κt​log⁡λ⏟GP−∑k=1nκ˙t1−κt​log⁡Q⏟reconstruction\.\\mathcal\{L\}=\\underbrace\{\\sum\_\{k=1\}^\{X\_\{t\}\}\\lambda\-\\sum\_\{k=1\}^\{n\}\\frac\{\\dot\{\\kappa\}\_\{t\}\}\{1\-\\kappa\_\{t\}\}\\log\\lambda\}\_\{\\text\{GP\}\}\-\\underbrace\{\\sum\_\{k=1\}^\{n\}\\frac\{\\dot\{\\kappa\}\_\{t\}\}\{1\-\\kappa\_\{t\}\}\\log Q\}\_\{\\text\{reconstruction\}\}\.\(69\)For the insertion loss, it immediately becomes clear that the first two terms together coincide withℒGP\\mathcal\{L\}\_\{\\text\{GP\}\}in Eq\.[4](https://arxiv.org/html/2607.09039#S3.E4)asnins=X1−Xtn\_\{\\text\{ins\}\}=X\_\{1\}\-X\_\{t\}, and the last terms coincide with the cross\-entropy reconstruction lossℒrec\\mathcal\{L\}\_\{\\text\{rec\}\}in Eq\.[91](https://arxiv.org/html/2607.09039#A2.E91)using the formulation inGat et al\. \[[17](https://arxiv.org/html/2607.09039#bib.bib17)\]\. For the substitution loss, since substitutions do not alter sequence lengths, the rate component can be effectively coalesced into a single logit componentQ^θ\\hat\{Q\}\_\{\\theta\}\. Specifically, the survival of a substitution event is equivalent to a substitution withs=s′s=s^\{\\prime\}\. In this way, we may assume the event always happens while compensating with a non\-zeroQ^θ​\(s,s\):=λθsub\>0\\hat\{Q\}\_\{\\theta\}\(s,s\):=\\lambda\_\{\\theta\}^\{\\text\{sub\}\}\>0to allow staying in the same state\. Therefore, throwing away the unnecessary rate part, the substitution loss becomes:

ℒsub=−∑k=1nκ˙t1−κt​log⁡Q^,\\mathcal\{L\}\_\{\\text\{sub\}\}=\-\\sum\_\{k=1\}^\{n\}\\frac\{\\dot\{\\kappa\}\_\{t\}\}\{1\-\\kappa\_\{t\}\}\\log\\hat\{Q\},\(70\)which is the standard loss for training a discrete flow matching model inGat et al\. \[[17](https://arxiv.org/html/2607.09039#bib.bib17)\], serving asℒFM\\mathcal\{L\}\_\{\\text\{FM\}\}to refine the discrete modality in our GPFlow formulation\. Combining the losses together, we getℒedit=ℒins\+ℒsub=ℒGP\+ℒrec\+ℒFM\\mathcal\{L\}\_\{\\text\{edit\}\}=\\mathcal\{L\}\_\{\\text\{ins\}\}\+\\mathcal\{L\}\_\{\\text\{sub\}\}=\\mathcal\{L\}\_\{\\text\{GP\}\}\+\\mathcal\{L\}\_\{\\text\{rec\}\}\+\\mathcal\{L\}\_\{\\text\{FM\}\}, which is but a special case of our GPFlow loss in Eq\.[8](https://arxiv.org/html/2607.09039#S3.E8)on a single discrete modality\.

## Appendix BAlgorithmic Details

![Refer to caption](https://arxiv.org/html/2607.09039v1/figs/train_sample.png)Figure 6:Training procedure \(left\) and one sampling step \(right\) for three variants of GPFlow:Length\-Only\(Section[3\.1](https://arxiv.org/html/2607.09039#S3.SS1)\),Multimodal Order\-Agnostic\(Section[3\.2](https://arxiv.org/html/2607.09039#S3.SS2)\), andMultimodal Order\-Preserving\(Appendix[B\.1](https://arxiv.org/html/2607.09039#A2.SS1)\)\.In this section, we provide additional details on the algorithms, model design, and general techniques we used to improve our GPFlow\. Explicit formulas for the losses are also provided\.

### B\.1Order\-Preserving Generalized Poisson Flow

For multimodal GPFlow, a natural setting arises when the ordering of a sequence matters in practice, as in protein design, where amino acids have a natural ordering\. In the main context, we treat all Poisson events as an unordered set and model the counting process, which is suitable for order\-agnostic generative tasks such as molecule generation\. We now discuss how to extend GPFlow to ordered sequential data in a natural way\.

For an ordered clean stateY1=\(X1,S1\)Y\_\{1\}=\(X\_\{1\},S\_\{1\}\)withX1=LX\_\{1\}=LandS1=\(S11,…,S1L\)S\_\{1\}=\(S^\{1\}\_\{1\},\\dots,S\_\{1\}^\{L\}\), consider a current stateYt=y=\(k,s\)Y\_\{t\}=y=\(k,s\)whosekkcomponents correspond to retained target indicesr1<⋯<rkr\_\{1\}<\\cdots<r\_\{k\}\. Setr0=0r\_\{0\}=0andrk\+1=L\+1r\_\{k\+1\}=L\+1\. TheL−kL\-kunretained target indices are partitioned among thek\+1k\+1insertion slots by:

Ωi:=\{ℓ∈\{1,…,L\}∖\{r1,…,rk\}:ri<ℓ<ri\+1\},0≤i≤k\.\\Omega\_\{i\}:=\\\{\\ell\\in\\\{1,\\dots,L\\\}\\setminus\\\{r\_\{1\},\\dots,r\_\{k\}\\\}:r\_\{i\}<\\ell<r\_\{i\+1\}\\\},\\quad 0\\leq i\\leq k\.\(71\)Thus,Ωi\\Omega\_\{i\}contains exactly the upcoming components that belong between theii\-th and\(i\+1\)\(i\+1\)\-st retained components, withi=0i=0andi=ki=krepresenting the beginning and end\. For each unretained indexℓ\\ell, letaℓ∈𝖲a\_\{\\ell\}\\in\\mathsf\{S\}denote its component value along the conditional path at timett\.

Consistent with Appendix[A\.4](https://arxiv.org/html/2607.09039#A1.SS4), define the insertion\-variable space atyyas:

𝖬​\(y\):=\{0,…,k\}×𝖲,m=\(i,a\),\\mathsf\{M\}\(y\):=\\\{0,\\dots,k\\\}\\times\\mathsf\{S\},\\quad m=\(i,a\),\(72\)whereiiis the insertion slot andaais the new component value\. Fors=\(s1,…,sk\)s=\(s^\{1\},\\dots,s^\{k\}\), the general insertion map specializes to:

I\(y,\(i,a\)\)=:Ii\(y,a\):=\(k\+1,\(s1,…,si,a,si\+1,…,sk\)\),I\(y,\(i,a\)\)=:I\_\{i\}\(y,a\):=\\left\(k\+1,\(s^\{1\},\\dots,s^\{i\},a,s^\{i\+1\},\\dots,s^\{k\}\)\\right\),\(73\)with the natural empty\-prefix or empty\-suffix convention ati=0i=0ori=ki=k\.

For the conditional thinning path indexed byY1Y\_\{1\}, slotiihas rate

λti​\(y\|Y1\)=\|Ωi\|​ht,λt​\(y\|Y1\)=∑i=0kλti​\(y\|Y1\)=\(L−k\)​ht,\\lambda\_\{t\}^\{i\}\(y\|Y\_\{1\}\)=\|\\Omega\_\{i\}\|h\_\{t\},\\quad\\lambda\_\{t\}\(y\|Y\_\{1\}\)=\\sum\_\{i=0\}^\{k\}\\lambda\_\{t\}^\{i\}\(y\|Y\_\{1\}\)=\(L\-k\)h\_\{t\},\(74\)andρti​\(d​a\|y,Y1\)\\rho\_\{t\}^\{i\}\(\\mathop\{\}\\\!\\mathrm\{d\}a\|y,Y\_\{1\}\)is the normalized sampling distribution over\{aℓ:ℓ∈Ωi\}\\\{a\_\{\\ell\}:\\ell\\in\\Omega\_\{i\}\\\}\. Whenλt​\(y\|Y1\)\>0\\lambda\_\{t\}\(y\|Y\_\{1\}\)\>0, these slot\-specific quantities define the joint insertion\-variable distribution:

ρt​\(i,d​a\|y,Y1\):=λti​\(y\|Y1\)λt​\(y\|Y1\)​ρti​\(d​a\|y,Y1\)\.\\rho\_\{t\}\(i,\\mathop\{\}\\\!\\mathrm\{d\}a\|y,Y\_\{1\}\):=\\frac\{\\lambda\_\{t\}^\{i\}\(y\|Y\_\{1\}\)\}\{\\lambda\_\{t\}\(y\|Y\_\{1\}\)\}\\rho\_\{t\}^\{i\}\(\\mathop\{\}\\\!\\mathrm\{d\}a\|y,Y\_\{1\}\)\.\(75\)Pushing this distribution through Eq\.[73](https://arxiv.org/html/2607.09039#A2.E73)gives the conditional outgoing jump measure:

Qt​\(A\|y,Y1\)=∑i=0kλti​\(y\|Y1\)​∫𝖲𝟙​\{I​\(y,\(i,a\)\)∈A\}​ρti​\(d​a\|y,Y1\),Q\_\{t\}\(A\|y,Y\_\{1\}\)=\\sum\_\{i=0\}^\{k\}\\lambda\_\{t\}^\{i\}\(y\|Y\_\{1\}\)\\int\_\{\\mathsf\{S\}\}\\mathds\{1\}\\\{I\(y,\(i,a\)\)\\in A\\\}\\rho\_\{t\}^\{i\}\(\\mathop\{\}\\\!\\mathrm\{d\}a\|y,Y\_\{1\}\),\(76\)which is the order\-preserving specialization of Eq\.[32](https://arxiv.org/html/2607.09039#A1.E32)\. Theorem[A\.5](https://arxiv.org/html/2607.09039#A1.SS5)posterior\-marginalizes these slot\-specific rates and distributions\.

During learned sampling, the same factorization first simulates an event with total rateλθ​\(y,t\)=∑i=0kλθi​\(y,t\)\\lambda\_\{\\theta\}\(y,t\)=\\sum\_\{i=0\}^\{k\}\\lambda\_\{\\theta\}^\{i\}\(y,t\), then chooses slotiiwith probabilityλθi​\(y,t\)/λθ​\(y,t\)\\lambda\_\{\\theta\}^\{i\}\(y,t\)/\\lambda\_\{\\theta\}\(y,t\), samplesa∼ρθi\(⋅\|y,t\)a\\sim\\rho\_\{\\theta\}^\{i\}\(\\cdot\|y,t\), and updates the state throughI​\(y,\(i,a\)\)=Ii​\(y,a\)I\(y,\(i,a\)\)=I\_\{i\}\(y,a\)\. This process generates the target order\-preserving distribution, as guaranteed by Theorem[A\.5](https://arxiv.org/html/2607.09039#A1.SS5)\. In this way, the generalized Poisson NLL loss in Eq\.[4](https://arxiv.org/html/2607.09039#S3.E4)is adapted as the summation overk\+1k\+1bins:

ℒGP=𝔼t,Y1∼q,Yt∼pt\(⋅\|Y1\)​\[∑i=0Xt\(λti−ht​\|Ωi\|​log⁡λti\)\],\\mathcal\{L\}\_\{\\text\{GP\}\}=\\mathbb\{E\}\_\{t,Y\_\{1\}\\sim q,Y\_\{t\}\\sim p\_\{t\}\(\\cdot\|Y\_\{1\}\)\}\\left\[\\sum\_\{i=0\}^\{X\_\{t\}\}\\left\(\\lambda^\{i\}\_\{t\}\-h\_\{t\}\|\\Omega\_\{i\}\|\\log\\lambda^\{i\}\_\{t\}\\right\)\\right\],\(77\)and the reconstruction loss of the sampling distributions in Eq\.[7](https://arxiv.org/html/2607.09039#S3.E7)is similarly adapted as:

ℒrec=𝔼t,Y1∼q,Yt∼pt\(⋅\|Y1\)​\[−ht​∑i=0Xt∑ℓ∈Ωilog⁡ρti​\(aℓ\|Yt\)\]\.\\mathcal\{L\}\_\{\\text\{rec\}\}=\\mathbb\{E\}\_\{t,Y\_\{1\}\\sim q,Y\_\{t\}\\sim p\_\{t\}\(\\cdot\|Y\_\{1\}\)\}\\left\[\-h\_\{t\}\\sum\_\{i=0\}^\{X\_\{t\}\}\\sum\_\{\\ell\\in\\Omega\_\{i\}\}\\log\\rho\_\{t\}^\{i\}\(a\_\{\\ell\}\|Y\_\{t\}\)\\right\]\.\(78\)We outline the training and sampling procedures for the order\-preserving GPFlow in Algorithm[2](https://arxiv.org/html/2607.09039#alg2)and[3](https://arxiv.org/html/2607.09039#alg3), respectively, and adopt it in our experimental setup for protein design\. A pictorial comparison of the training and sampling stages of our GPFlow variants is shown in Figure[6](https://arxiv.org/html/2607.09039#A2.F6)across different scenarios\. We also note that the previous discussions for both order\-agnostic and order\-preserving GPFlow can be easily generalized to an arbitrary number of length\-dependent modalitiesS=\{S\(m\)\}1≤m≤M=\{S\(m\),i\}1≤m≤M,1≤i≤XS=\\\{S^\{\(m\)\}\\\}\_\{1\\leq m\\leq M\}=\\\{S^\{\(m\),i\}\\\}\_\{1\\leq m\\leq M,1\\leq i\\leq X\}\. For simplicity and clarity of notation, we will omit the modality superscripts\.

Algorithm 2Training Order\-Preserving GPFlow1:whilenot convergeddo

2:Sample

t∼\[0,1\];Y1=\(X1,S1\)∼q​\(X,S\)t\\sim\[0,1\];Y\_\{1\}=\(X\_\{1\},S\_\{1\}\)\\sim q\(X,S\)\.

3:Sample

St∼pt​\(S\)S\_\{t\}\\sim p\_\{t\}\(S\)for each

Sti,1≤i≤X1S\_\{t\}^\{i\},1\\leq i\\leq X\_\{1\}\.

4:For each position

ℓ\\ell, delete

StℓS\_\{t\}^\{\\ell\}with probability

1−κt1\-\\kappa\_\{t\}to obtain

S^t\\hat\{S\}\_\{t\}and

Xt=\|S^t\|,Yt=\(Xt,S^t\)X\_\{t\}=\|\\hat\{S\}\_\{t\}\|,Y\_\{t\}=\(X\_\{t\},\\hat\{S\}\_\{t\}\)\.

5:Predict the rates

λθi​\(Yt,t\)\\lambda^\{i\}\_\{\\theta\}\(Y\_\{t\},t\), the length\-dependent modality vector field

vθ​\(Yt,t\)v\_\{\\theta\}\(Y\_\{t\},t\), and the sampling distributions

ρθi​\(Yt,t\),0≤i≤Xt\\rho^\{i\}\_\{\\theta\}\(Y\_\{t\},t\),0\\leq i\\leq X\_\{t\}\.

6:Optimize the loss

ℒ\\mathcal\{L\}in Eq\.[8](https://arxiv.org/html/2607.09039#S3.E8)\.

Algorithm 3Sampling from Order\-Preserving GPFlow \(Euler\)1:

Y0=\(X0=0,S0=∅\)Y\_\{0\}=\(X\_\{0\}=0,S\_\{0\}=\\emptyset\)\.

2:for

n←1,2,…,Nn\\leftarrow 1,2,\\dots,Ndo

3:

t←tn,Δ​t←tn\+1−tnt\\leftarrow t\_\{n\},\\Delta t\\leftarrow t\_\{n\+1\}\-t\_\{n\}, where

tN\+1:=1t\_\{N\+1\}:=1\.

4:Predict the rates

λθi​\(Yt,t\)\\lambda^\{i\}\_\{\\theta\}\(Y\_\{t\},t\), the vector field

vθ​\(Yt,t\)v\_\{\\theta\}\(Y\_\{t\},t\), and the sampling distributions

ρθi​\(Yt,t\)\\rho\_\{\\theta\}^\{i\}\(Y\_\{t\},t\)for

0≤i≤Xt0\\leq i\\leq X\_\{t\}\.

5:Advance the generative flow for each

StS\_\{t\}with

vθv\_\{\\theta\}to obtain

S^t\+Δ​t\\hat\{S\}\_\{t\+\\Delta t\}\.

6:Set

Y^t\+Δ​t:=\(Xt,S^t\+Δ​t\)\\hat\{Y\}\_\{t\+\\Delta t\}:=\(X\_\{t\},\\hat\{S\}\_\{t\+\\Delta t\}\)and

λθ:=∑i=0Xtλθi​\(Yt,t\)\\lambda\_\{\\theta\}:=\\sum\_\{i=0\}^\{X\_\{t\}\}\\lambda^\{i\}\_\{\\theta\}\(Y\_\{t\},t\)\.

7:Sample an insertion event with probability

λθ​Δ​t\\lambda\_\{\\theta\}\\Delta t\.

8:If an event happens, choose slot

iiwith probability

λθi/λθ\\lambda\_\{\\theta\}^\{i\}/\\lambda\_\{\\theta\}, sample

a∼ρθi\(⋅\|Yt,t\)a\\sim\\rho\_\{\\theta\}^\{i\}\(\\cdot\|Y\_\{t\},t\), and set

Yt\+Δ​t=Ii​\(Y^t\+Δ​t,a\)Y\_\{t\+\\Delta t\}=I\_\{i\}\(\\hat\{Y\}\_\{t\+\\Delta t\},a\); otherwise set

Yt\+Δ​t=Y^t\+Δ​tY\_\{t\+\\Delta t\}=\\hat\{Y\}\_\{t\+\\Delta t\}\.

9:Return:

Y1=\(X1,S1\)Y\_\{1\}=\(X\_\{1\},S\_\{1\}\)\.

### B\.2Training and Sampling Pseudocode

We provide additional pseudo\-code for training and sampling length\-only \(Algorithm[5](https://arxiv.org/html/2607.09039#alg5),[6](https://arxiv.org/html/2607.09039#alg6)\), order\-agnostic \(Algorithm[4](https://arxiv.org/html/2607.09039#alg4)\), and conditional order\-preserving \(Algorithm[7](https://arxiv.org/html/2607.09039#alg7),[8](https://arxiv.org/html/2607.09039#alg8)\) GPFlow\. The three unconditional variants are summarized in Figure[6](https://arxiv.org/html/2607.09039#A2.F6), which highlights the differences in the training objectives and sampling predictions\.

Algorithm 4Training Multimodal GPFlow1:whilenot convergeddo

2:Sample

t∼\[0,1\];Y1=\(X1,S1\)∼q​\(X,S\)t\\sim\[0,1\];Y\_\{1\}=\(X\_\{1\},S\_\{1\}\)\\sim q\(X,S\)\.

3:Sample

St∼pt​\(S\)S\_\{t\}\\sim p\_\{t\}\(S\)for each

Sti,1≤i≤X1S\_\{t\}^\{i\},1\\leq i\\leq X\_\{1\}\.

4:For each element

Sti∈StS\_\{t\}^\{i\}\\in S\_\{t\}, delete it with probability

1−κt1\-\\kappa\_\{t\}to obtain

S^t\\hat\{S\}\_\{t\}and

Xt=\|S^t\|,Yt=\(Xt,S^t\)X\_\{t\}=\|\\hat\{S\}\_\{t\}\|,Y\_\{t\}=\(X\_\{t\},\\hat\{S\}\_\{t\}\)\.

5:Predict the rate

λθ​\(Yt,t\)\\lambda\_\{\\theta\}\(Y\_\{t\},t\), the length\-dependent modality vector field

vθ​\(Yt,t\)v\_\{\\theta\}\(Y\_\{t\},t\), and the sampling distribution

ρθ​\(Yt,t\)\\rho\_\{\\theta\}\(Y\_\{t\},t\)\.

6:Optimize the loss

ℒ\\mathcal\{L\}in Eq\.[8](https://arxiv.org/html/2607.09039#S3.E8)\.

Algorithm 5Training Length\-Only GPFlow1:whilenot convergeddo

2:Sample

t∼\[0,1\];X1∼q​\(X\)t\\sim\[0,1\];X\_\{1\}\\sim q\(X\)\.

3:Sample

Xt∼B​\(X1,κt\)X\_\{t\}\\sim B\(X\_\{1\},\\kappa\_\{t\}\)\.

4:Predict the rate

λθ​\(Xt,t\)\\lambda\_\{\\theta\}\(X\_\{t\},t\)and optimize the loss

ℒPP\\mathcal\{L\}\_\{\\text\{PP\}\}in Eq\.[4](https://arxiv.org/html/2607.09039#S3.E4)\.

Algorithm 6Sampling from Length\-Only GPFlow \(Euler\)1:

X0=0X\_\{0\}=0\.

2:for

i←1,2,…,Ni\\leftarrow 1,2,\\dots,Ndo

3:

t←ti,Δ​t←ti\+1−tit\\leftarrow t\_\{i\},\\Delta t\\leftarrow t\_\{i\+1\}\-t\_\{i\}, where

tN\+1:=1t\_\{N\+1\}:=1\.

4:Predict the rate

λθ​\(Xt,t\)\\lambda\_\{\\theta\}\(X\_\{t\},t\)\.

5:Sample an event with a probability of

λθ​Δ​t\\lambda\_\{\\theta\}\\Delta t\.

6:If an event happens,

Xt\+Δ​t=Xt\+1X\_\{t\+\\Delta t\}=X\_\{t\}\+1; else

Xt\+Δ​t=XtX\_\{t\+\\Delta t\}=X\_\{t\}\.

7:Return:

X1X\_\{1\}\.

Algorithm 7Training Conditional Order\-Preserving GPFlow1:whilenot convergeddo

2:Sample

t∼\[0,1\];Y1=\(X1,S1\)∼q​\(X,S\)t\\sim\[0,1\];Y\_\{1\}=\(X\_\{1\},S\_\{1\}\)\\sim q\(X,S\)\.

3:Sample condition

Y0=\(Xc,Sc\)∼qc​\(X,S\|X1,S1\)Y\_\{0\}=\(X\_\{c\},S\_\{c\}\)\\sim q\_\{c\}\(X,S\|X\_\{1\},S\_\{1\}\)with clean values\.

4:Sample

St∼pt​\(S\)S\_\{t\}\\sim p\_\{t\}\(S\)for each non\-conditioning position

ii,

Sti,1≤i≤X1S\_\{t\}^\{i\},1\\leq i\\leq X\_\{1\}\.

5:For each non\-conditioning position

ii, delete

StiS\_\{t\}^\{i\}with probability

1−κt1\-\\kappa\_\{t\}to obtain

S^t\\hat\{S\}\_\{t\}and

Xt=\|S^t\|,Yt=\(Xt,S^t\)X\_\{t\}=\|\\hat\{S\}\_\{t\}\|,Y\_\{t\}=\(X\_\{t\},\\hat\{S\}\_\{t\}\)\.

6:Predict the rates

λθi​\(Yt,t\)\\lambda^\{i\}\_\{\\theta\}\(Y\_\{t\},t\), the length\-dependent modality vector field

vθ​\(Yt,t\)v\_\{\\theta\}\(Y\_\{t\},t\), and the sampling distributions

ρθi​\(Yt,t\),0≤i≤Xt\\rho^\{i\}\_\{\\theta\}\(Y\_\{t\},t\),0\\leq i\\leq X\_\{t\}\.

7:Optimize the loss

ℒ\\mathcal\{L\}in Eq\.[8](https://arxiv.org/html/2607.09039#S3.E8)\.

Algorithm 8Sampling from Conditional Order\-Preserving GPFlow \(Euler\)1:Input:

Y0=\(Xc,Sc\)Y\_\{0\}=\(X\_\{c\},S\_\{c\}\)\.

2:for

i←1,2,…,Ni\\leftarrow 1,2,\\dots,Ndo

3:

t←ti,Δ​t←ti\+1−tit\\leftarrow t\_\{i\},\\Delta t\\leftarrow t\_\{i\+1\}\-t\_\{i\}, where

tN\+1:=1t\_\{N\+1\}:=1\.

4:Predict the rates

λθi​\(Yt,t\)\\lambda^\{i\}\_\{\\theta\}\(Y\_\{t\},t\), the vector field

vθ​\(Yt,t\)v\_\{\\theta\}\(Y\_\{t\},t\), and the sampling distributions

ρθi​\(Yt,t\)\\rho\_\{\\theta\}^\{i\}\(Y\_\{t\},t\)for

0≤i≤Xt0\\leq i\\leq X\_\{t\}\.

5:Advance the generative flow for each

StS\_\{t\}with

vθv\_\{\\theta\}to obtain

S^t\+Δ​t\\hat\{S\}\_\{t\+\\Delta t\}for each non\-conditioning position

ii\.

6:Zero out the rates where insertions are not allowed \(e\.g\., inside a contiguous condition sequence\)\.

7:Simulate the Poisson process by sampling an event on

Xt\+Δ​tX\_\{t\+\\Delta t\}with a probability of

λθ​Δ​t\\lambda\_\{\\theta\}\\Delta twhere

λθ=∑iλθi\\lambda\_\{\\theta\}=\\sum\_\{i\}\\lambda^\{i\}\_\{\\theta\}\.

8:If an event happens, choose the insertion position

iiwith probabilities proportional to

\{λi\}\\\{\\lambda\_\{i\}\\\}and insert the new feature value

si∼ρθis\_\{i\}\\sim\\rho^\{i\}\_\{\\theta\}into

S^t\+Δ​t\\hat\{S\}\_\{t\+\\Delta t\}to obtain

St\+Δ​tS\_\{t\+\\Delta t\}\.

9:Update

Yt\+Δ​t=\(Xt\+Δ​t,St\+Δ​t\)Y\_\{t\+\\Delta t\}=\(X\_\{t\+\\Delta t\},S\_\{t\+\\Delta t\}\)\.

10:Return:

Y1=\(X1,S1\)Y\_\{1\}=\(X\_\{1\},S\_\{1\}\)\.

In all sampling algorithms, the sequence of discretization time steps0≤t1<t2<⋯<tN<tN\+1=10\\leq t\_\{1\}<t\_\{2\}<\\cdots<t\_\{N\}<t\_\{N\+1\}=1is defined by an inference\-time scheduler, and more advanced solvers \(e\.g\., the second\-order Heun solver\) can be readily applied\.

### B\.3Flow Parameterization

##### Continuous Euclidean Modality

Existing works have explored fixed\-length diffusion or flow\-based generative models for both continuous and discrete modalities, whose loss can be adapted as ourℒFM\\mathcal\{L\}\_\{\\text\{FM\}\}for refining the existing length\-dependent elements\. For example, consider a flow matching model with the optimal\-transport path\[[34](https://arxiv.org/html/2607.09039#bib.bib34)\]st=\(1−t\)​s0\+t​s1s\_\{t\}=\(1\-t\)s\_\{0\}\+ts\_\{1\}, wheres0∼𝒩​\(0,σ02\)s\_\{0\}\\sim\\mathcal\{N\}\(0,\\sigma\_\{0\}^\{2\}\)\. The flow matching loss is

ℒFM=𝔼t,s0∼p0,s1∼q​\[‖vθ​\(st,t\)−\(s1−s0\)‖2\],\\mathcal\{L\}\_\{\\text\{FM\}\}=\\mathbb\{E\}\_\{t,s\_\{0\}\\sim p\_\{0\},s\_\{1\}\\sim q\}\[\\\|v\_\{\\theta\}\(s\_\{t\},t\)\-\(s\_\{1\}\-s\_\{0\}\)\\\|^\{2\}\],\(79\)wherevθ​\(st,t\)v\_\{\\theta\}\(s\_\{t\},t\)is the vector field predictor\. We follow Proteina in using this loss to refine the continuous CA coordinates in our structure design model\. For sampling from such a continuous vector field, a simple Euler update reads:

s←s\+vθ​Δ​t\.s\\leftarrow s\+v\_\{\\theta\}\\Delta t\.\(80\)We may also employ more complex sampling algorithms that incorporate higher\-order information or stochasticity \(e\.g\., SDE sampling\)\. For example, one sampling step in Proteina\[[18](https://arxiv.org/html/2607.09039#bib.bib18)\]reads:

s←s\+\(vθ\+gt​sθ\)​Δ​t\+2​gt​γ​Δ​t​ε\.s\\leftarrow s\+\(v\_\{\\theta\}\+g\_\{t\}s\_\{\\theta\}\)\\Delta t\+\\sqrt\{2g\_\{t\}\\gamma\\Delta t\}\\,\\varepsilon\.\(81\)wheresθ​\(st,t\):=\(t​vθ​\(st,t\)−st\)/\(1−t\)s\_\{\\theta\}\(s\_\{t\},t\):=\(tv\_\{\\theta\}\(s\_\{t\},t\)\-s\_\{t\}\)/\(1\-t\)is the score function,gtg\_\{t\}is the score scheduler,0<γ<10<\\gamma<1is the noise scale, andε∼𝒩​\(0,1\)\\varepsilon\\sim\\mathcal\{N\}\(0,1\)\. We also adopt this SDE sampling for the coordinates\.

##### Continuous Riemannian Modality

For Riemannian modalities such as rotations \(which lie in the 3D rotation groupSO​\(3\)\\mathrm\{SO\}\(3\)\) and torsion angles \(which lie in the flat torus𝕋\\mathbb\{T\}\) in the peptide co\-design task, the standard Riemannian flow matching loss\[[11](https://arxiv.org/html/2607.09039#bib.bib11)\]is applied:

ℒFM=𝔼t,s0∼p0,s1∼q\[∥vθ\(st,t\)−ut\(st\|s1\)∥g2\],\\mathcal\{L\}\_\{\\text\{FM\}\}=\\mathbb\{E\}\_\{t,s\_\{0\}\\sim p\_\{0\},s\_\{1\}\\sim q\}\[\\\|v\_\{\\theta\}\(s\_\{t\},t\)\-u\_\{t\}\(s\_\{t\}\|s\_\{1\}\)\\\|\_\{g\}^\{2\}\],\(82\)where geodesic interpolation is used as the noising processst:=exps0⁡\(κt​logs0⁡s1\)s\_\{t\}:=\\exp\_\{s\_\{0\}\}\(\\kappa\_\{t\}\\log\_\{s\_\{0\}\}s\_\{1\}\)with a differentiable schedulerκt\\kappa\_\{t\}\. Theexp\\expdenotes the exponential map,log\\logthe logarithm map, and∥⋅∥g\\\|\\cdot\\\|\_\{g\}the Riemannian norm\. The conditional vector fieldutu\_\{t\}can be calculated in closed form as

ut​\(st\|s1\)=κ˙t1−κt​logst⁡s1\.u\_\{t\}\(s\_\{t\}\|s\_\{1\}\)=\\frac\{\\dot\{\\kappa\}\_\{t\}\}\{1\-\\kappa\_\{t\}\}\\log\_\{s\_\{t\}\}s\_\{1\}\.\(83\)
Sampling from Riemannian flow matching follows the Euler discretization as:

s←exps⁡\(vθ​Δ​t\)\.s\\leftarrow\\exp\_\{s\}\(v\_\{\\theta\}\\Delta t\)\.\(84\)
The mathematical expression for the Riemannian operations onSO​\(3\)\\mathrm\{SO\}\(3\)and𝕋\\mathbb\{T\}is known in closed form\. We refer interested readers to, e\.g\., Appendix A ofLi et al\. \[[29](https://arxiv.org/html/2607.09039#bib.bib29)\]for full details\.

##### Discrete Modality

For continuous\-state discrete flow matching models\[[13](https://arxiv.org/html/2607.09039#bib.bib13)\]that embed the discrete token into a continuous modality, the previous discussion of continuous\-modality losses applies directly\[[24](https://arxiv.org/html/2607.09039#bib.bib24),[12](https://arxiv.org/html/2607.09039#bib.bib12),[16](https://arxiv.org/html/2607.09039#bib.bib16)\]\. For example, the simplex\-based GPFlow variant for peptide design adopts the soft logitss1=K​\(2​δc1−1\)s\_\{1\}=K\(2\\delta\_\{c\_\{1\}\}\-1\), wherec1c\_\{1\}is the corresponding amino acid type andK=5K=5\.s0s\_\{0\}is sampled from𝒩​\(0,K​I\)\\mathcal\{N\}\(0,KI\), and Euclidean geometry is adopted\. FollowingLi et al\. \[[29](https://arxiv.org/html/2607.09039#bib.bib29)\], the cross\-entropy loss is used instead of the vector field matching loss:

ℒFM=𝔼t,s0∼p0,s1∼q​\[CE​\(fθ​\(st,t\),c1\)\],\\mathcal\{L\}\_\{\\text\{FM\}\}=\\mathbb\{E\}\_\{t,s\_\{0\}\\sim p\_\{0\},s\_\{1\}\\sim q\}\[\\mathrm\{CE\}\(f\_\{\\theta\}\(s\_\{t\},t\),c\_\{1\}\)\],\(85\)wherefθ​\(st,t\)f\_\{\\theta\}\(s\_\{t\},t\)is the logit predictor \(probability denoiser\)\.

For the discrete modality that relies on a continuous\-time Markov chain, the intermediate noisysts\_\{t\}is a discrete token rather than continuous logits\. The discrete flow matching loss\[[17](https://arxiv.org/html/2607.09039#bib.bib17),[41](https://arxiv.org/html/2607.09039#bib.bib41)\]reads:

ℒFM=𝔼t,s0∼p0,s1∼q​\[∑st≠s1CE​\(fθ​\(st,t\),s1\)\],\\mathcal\{L\}\_\{\\text\{FM\}\}=\\mathbb\{E\}\_\{t,s\_\{0\}\\sim p\_\{0\},s\_\{1\}\\sim q\}\\left\[\\sum\_\{s\_\{t\}\\neq s\_\{1\}\}\\mathrm\{CE\}\(f\_\{\\theta\}\(s\_\{t\},t\),s\_\{1\}\)\\right\],\(86\)where the noisysts\_\{t\}is sampled from some pre\-defined probability path on the noisy logitsxtx\_\{t\}:

st:=\{∼Cat​\(xt\),w\.p\.​1−t,s1,w\.p\.​t\.s\_\{t\}:=\\begin\{cases\}\\sim\\mathrm\{Cat\}\(x\_\{t\}\),&\\text\{w\.p\. \}1\-t,\\\\ s\_\{1\},&\\text\{w\.p\. \}t\.\\end\{cases\}\(87\)A simple yet effective practice follows mask diffusion\[[41](https://arxiv.org/html/2607.09039#bib.bib41)\]: each noisy token starts with a special mask tokenmmwiths0≡δms\_\{0\}\\equiv\\delta\_\{m\}, and the FM loss is computed over all mask tokens\. During sampling, a single mask diffusion update reads:

s←\{∼Cat​\(fθ​\(st,t\)\),if​st=m​and w\.p\.​Δ​t/\(1−t\),st,else\.s\\leftarrow\\begin\{cases\}\\sim\\mathrm\{Cat\}\(f\_\{\\theta\}\(s\_\{t\},t\)\),&\\text\{if \}s\_\{t\}=m\\text\{ and w\.p\. \}\\Delta t/\(1\-t\),\\\\ s\_\{t\},&\\text\{else\}\.\\end\{cases\}\(88\)Notably, we may consider a special case in which the insertion element type is fixed to be the mask token, leaving no reconstruction lossℒrec=0\\mathcal\{L\}\_\{\\text\{rec\}\}=0for the discrete modality\. In this case, the model completely relies on the non\-zero mask diffusion lossℒFM\\mathcal\{L\}\_\{\\text\{FM\}\}to determine the final discrete type\.

### B\.4Sampling Distribution Parameterization

The marginal sampling distribution in Theorem[3\.2](https://arxiv.org/html/2607.09039#S3.SS2)and[A\.5](https://arxiv.org/html/2607.09039#A1.SS5)is a mixture distribution of the forward noisy length\-dependent modality from which sampling is feasible \(when conditioned on the clean target\)\. Therefore, we can rely on another generative model for modeling such a distribution, whose general NLL loss form is described in Eq\.[7](https://arxiv.org/html/2607.09039#S3.E7)and[78](https://arxiv.org/html/2607.09039#A2.E78)\.

##### Continuous Euclidean Modality

We first consider the continuous modality case where directly regressing on the sampling marginalρt\\rho\_\{t\}is often intractable\. We may rely on another flow matching head or any generative model to approximateρt\\rho\_\{t\}parametrically, using the current stateStS\_\{t\}and the noisy element to be inserted as inputs\. While such a nested flow approach is theoretically preferred, it requires simulating an additional full flow sampling for each Poisson event, thereby substantially slowing the sampling process\. In practice, we found that using the MSE loss to regress the mean behavior ofStS\_\{t\}already performs well enough with the existing later denoising steps\. The simplified loss can be written as:

ℒrec=𝔼Y1,Yt​\[ht​∑k=1X1−Xt\(fθ​\(Yt\)−Stk\)2\]\.\\mathcal\{L\}\_\{\\text\{rec\}\}=\\mathbb\{E\}\_\{Y\_\{1\},Y\_\{t\}\}\\left\[h\_\{t\}\\sum\_\{k=1\}^\{X\_\{1\}\-X\_\{t\}\}\(f\_\{\\theta\}\(Y\_\{t\}\)\-S\_\{t\}^\{k\}\)^\{2\}\\right\]\.\(89\)

##### Continuous Riemannian Modality

The simplified reconstruction loss can be easily extended to Riemannian generative modeling \(e\.g\., rotations and torsion angles in peptide design\), which takes:

ℒrec=𝔼Y1,Yt​\[ht​∑k=1X1−Xtdg2​\(fθ​\(Yt\),Stk\)\],\\mathcal\{L\}\_\{\\text\{rec\}\}=\\mathbb\{E\}\_\{Y\_\{1\},Y\_\{t\}\}\\left\[h\_\{t\}\\sum\_\{k=1\}^\{X\_\{1\}\-X\_\{t\}\}d\_\{g\}^\{2\}\(f\_\{\\theta\}\(Y\_\{t\}\),S\_\{t\}^\{k\}\)\\right\],\(90\)wheredg2​\(x,y\)=‖logx⁡y‖g2d\_\{g\}^\{2\}\(x,y\)=\\\|\\log\_\{x\}y\\\|\_\{g\}^\{2\}is the squared geodesic distance\.

##### Discrete Modality

For Markov\-chain\-based discrete flow matching models,Gat et al\. \[[17](https://arxiv.org/html/2607.09039#bib.bib17)\]has established a connection between the discrete vector field and the probability flow over categorical data, where the sampling marginal can be obtained in closed form\. Following this, the KL\-divergence \(cross\-entropy loss\) is used to train the discrete flow, and the reconstruction loss can be calculated as a summation of weighted cross\-entropy losses as:

ℒrec=𝔼Y1,Yt​\[ht​∑k=1X1−XtCE​\(fθ​\(Yt\),Stk\)\],\\mathcal\{L\}\_\{\\text\{rec\}\}=\\mathbb\{E\}\_\{Y\_\{1\},Y\_\{t\}\}\\left\[h\_\{t\}\\sum\_\{k=1\}^\{X\_\{1\}\-X\_\{t\}\}\\mathrm\{CE\}\(f\_\{\\theta\}\(Y\_\{t\}\),S\_\{t\}^\{k\}\)\\right\],\(91\)wherefθf\_\{\\theta\}is the logit predictor\. Thus, unlike the continuous case, the discrete case is always simulation\-free, as the learned logits already encode the marginalized discrete sampling distribution\.

In contrast to the flow\-oriented loss discussed in the previous subsection, we may consider another special case in which the element type is determined upon insertion, yielding no refinement loss withℒFM=0\\mathcal\{L\}\_\{\\text\{FM\}\}=0\. In this approach, the sampling distribution serves a dual role in determining the discrete types without further refinement\. This parameterization and sampling scheme are adopted for our sequence design model\. Nonetheless, one may also allow token refinement relying on the discrete diffusion\[[2](https://arxiv.org/html/2607.09039#bib.bib2),[41](https://arxiv.org/html/2607.09039#bib.bib41)\]or discrete flow matching\[[17](https://arxiv.org/html/2607.09039#bib.bib17)\]lossℒFM\\mathcal\{L\}\_\{\\text\{FM\}\}to further refine the discrete modality\.

For all time\-dependent modalities, the training noise schedulers and inference time sampling schedulers for the flow and reconstruction parts can be either coupled with the rate scheduler or separate\. We typically use different schedulers for the rate and length\-dependent modality to allow greater flexibility during sampling\. In this way, time steps across modalities are not shared during training, and separate time embeddings are fed into the model\.

### B\.5τ\\tau\-Leaping

One practical limitation of existing rate\-based variable\-length generative models is their slower sampling speed compared to standard fixed\-length flow sampling, because we must simulate the inhomogeneous Poisson process at each time step, which is more sensitive\. For example, EditFlow\[[20](https://arxiv.org/html/2607.09039#bib.bib20)\]required 5,000\-10,000 sampling steps to achieve reasonable performance, which is significantly larger than the average sequence length and the typical number of steps used in fixed\-length diffusion or flow matching models\. To speed up the sampling process, we consider simulating multiple Poisson events within a small time stepτ\\tauby leveraging the*τ\\tau\-leaping*method\[[19](https://arxiv.org/html/2607.09039#bib.bib19)\]\. Specifically, instead of sampling one event with a success probability ofλ​τ\\lambda\\tau, the increment is now sampled from a Poisson distribution with a rate ofλ​τ\\lambda\\tau:

Xs−Xt∼𝒫​\(λ​\(Xt,t\)​τ\),τ=s−t\.X\_\{s\}\-X\_\{t\}\\sim\\mathcal\{P\}\\left\(\\lambda\(X\_\{t\},t\)\\tau\\right\),\\quad\\tau=s\-t\.\(92\)Note that we havePr⁡\(Xs−Xt=0\)=exp⁡\(−λ​τ\)\\Pr\(X\_\{s\}\-X\_\{t\}=0\)=\\exp\{\(\-\\lambda\\tau\)\}, which is close to1−λ​τ1\-\\lambda\\tauwhenλ​τ\\lambda\\tauis small, recovering the standard simulation procedure\. Nonetheless, whenλ​τ\\lambda\\tauis large, theτ\\tau\-leaping method can significantly reduce the number of sampling steps while still maintaining a good approximation to the original process\. As an example, the ground\-truth marginal rate on the PDB dataset att=0t=0is the average sequence length, which is approximately 166\. With a standard 400\-step sampling setup in Proteina, the initial rate is approximately 0\.4, at which point the approximation will fail\.

In addition, we also allow for a more general form ofτ\\tau\-leaping, where simultaneous insertions can happen at multiple places independently according to their own rates and sampling distributions\[[30](https://arxiv.org/html/2607.09039#bib.bib30)\]\. Such a multi\-place insertion better approximates the original process than the Bernoulli approximation\. Nonetheless, we note thatτ\\tau\-leaping relies on the assumption that a single event does not significantly affect predictions, which does not necessarily hold even for discrete diffusion with a fixed time step\[[40](https://arxiv.org/html/2607.09039#bib.bib40)\]\. In our scenario, where length is also variable, such an approximation error may be more pronounced\. Despite this, we empirically find that both forms ofτ\\tau\-leaping improve performance and efficiency in the 400\-step generation setup\.

## Appendix CExperiment Details

In this section, we detail the datasets, task\-specific architectures, and evaluation pipelines\. The two structure\-based generation tasks are built on the Proteina base model \(with different feature variants\) and use the same datasets; the two sequence\-design tasks share a separate backbone family, and peptide design is described in its own subsection\. We therefore organize this section by dataset/model architecture\.

Proteins are naturally ordered according to their amino acid sequence\. We always use the order\-preserving GPFlow for generative modeling in all tasks\. Table[5](https://arxiv.org/html/2607.09039#A3.T5)summarizes the per\-task modeling choices, including the modality geometry, the flow and reconstruction losses, the insertion scheduler, and the base model architecture\.

Table 5:Per\-task modeling choices for GPFlow\. All tasks use the order\-preserving variant\.### C\.1Protein Structure Generation

#### C\.1\.1Data Specification

The datasets we use to train the model include PDB\[[5](https://arxiv.org/html/2607.09039#bib.bib5)\], the experimentally validated protein structure database, and AFDB\[[25](https://arxiv.org/html/2607.09039#bib.bib25)\], the AlphaFold2\-predicted structure database that contains more structures\. For the PDB dataset, we use the Proteina preprocessing script to filter entries that have a single chain, contain no non\-canonical amino acids, and have lengths between 50 and 256\. The data are then partitioned using a 50% sequence\-similarity threshold to prevent information leakage within the same cluster, yielding 6822 distinct clusters, for a total of 27k training structures\. During training, we follow Proteina’s approach to uniformly sample one cluster and then a member of that cluster\.

For AFDB, we follow Proteina to filter and cluster the original 21M AFDB structures using sequence\-based MMseqs2\[[43](https://arxiv.org/html/2607.09039#bib.bib43)\]and structure\-based FoldSeek\[[46](https://arxiv.org/html/2607.09039#bib.bib46)\], resulting in 713k structures\. Unlike PDB, we use the standard uniform sampling strategy\. As AFDB contains much longer proteins, each structure is dynamically cropped to a maximum sequence length of 512 to avoid memory issues during training\.

#### C\.1\.2Model Specification

##### Architecture\.

We build each structure model on top of its corresponding 60M Proteina\[[18](https://arxiv.org/html/2607.09039#bib.bib18)\]architecture and add heads to the final hidden representation to learn per\-position rates and sampling distributions, thereby increasing the total number of trainable parameters to 65 M\. In both unconditional generation and motif scaffolding, the only length\-dependent variable is the continuous CA coordinates\.

For motif scaffolding, we adapt Proteina’s dedicated motif architecture variant, which includes additional information on the motif mask, coordinates, and pairwise distances\. This architecture is not checkpoint\-compatible with the unconditional Proteina model, so neither method transfers an unconditional checkpoint into the motif experiment\. All residue indices are renumbered to avoid exposing their original positions\. We additionally inject sequence embeddings for the motif residues \(if available\), with other non\-motif residues set to unknown \(X\)\. We adopt a hard constraint in an inpainting fashion: the motif coordinates and types are never corrupted or modified and incur no loss during either training or sampling\.

##### Training\.

During training, we generally follow the original setup and hyperparameters in Proteina, where the time step is sampled fromt∼0\.98​Beta​\(1\.9,1\.0\)\+0\.02​U​\[0,1\]t\\sim 0\.98\\mathrm\{Beta\}\(1\.9,1\.0\)\+0\.02U\[0,1\]\. For the length modality, we adopt the early\-completion principle used by TDDM\[[9](https://arxiv.org/html/2607.09039#bib.bib9)\]and setκt:=min⁡\{1,t/0\.6\}\\kappa\_\{t\}:=\\min\\\{1,t/0\.6\\\}\. The length timet∼U​\[0,1\]t\\sim U\[0,1\]is sampled independently from the continuous time and is fed to the model as an additional input\. As defined in Theorem[3\.1](https://arxiv.org/html/2607.09039#S3.SS1), the process is absorbed onceκt=1\\kappa\_\{t\}=1: fort≥0\.6t\\geq 0\.6, all target elements are present, the insertion rate and insertion\-loss contribution are set to zero, and only the within\-length refinement process remains active\. The OT\-path\-based flow matching loss in Eq\.[79](https://arxiv.org/html/2607.09039#A2.E79), the NLL loss for the Poisson process in Eq\.[4](https://arxiv.org/html/2607.09039#S3.E4), and the mean reconstruction loss in Eq\.[89](https://arxiv.org/html/2607.09039#A2.E89)are used with weightswFM=wrec=1w\_\{\\text\{FM\}\}=w\_\{\\text\{rec\}\}=1\. For motif scaffolding, we use the conditional generation setup, in which the motif CA coordinates remain unchanged and are never deleted during the forward noising process\. The other losses are identical\.

For unconditional generation, two separate 65M models are trained: one with a total batch size of 128 and 300k iterations on PDB, and another with a total batch size of 32 and 1\.6M iterations on AFDB\. For motif scaffolding, both the 65M GPFlow model and the corresponding Proteina baseline use the AFDB dataset and the same motif augmentation\. Our model is trained for 1M iterations with a total batch size of 192\. We followLin et al\. \[[31](https://arxiv.org/html/2607.09039#bib.bib31)\]to randomly sample 0\-4 contiguous parts \(contigs\) as motifs during training\. All models are trained with the Adam optimizer with a learning rate of10−410^\{\-4\}\. An exponential moving average \(EMA\) of the model parameters is maintained with a decay of 0\.999 and used for all evaluations\. All models were trained on 8 NVIDIA A100 GPUs, taking approximately 2 days on PDB and 5 days on AFDB\.

##### Sampling\.

During sampling, 400 sampling steps are used following the Proteina setup\. For the continuous modality, we also follow Proteina to utilize SDE sampling described by Eq\.[81](https://arxiv.org/html/2607.09039#A2.E81)withgt=1/tg\_\{t\}=1/t,γ=0\.35\\gamma=0\.35, and an exponential inference scheduler\. For the length modality, we useκt:=min⁡\{1,t/0\.3\}\\kappa\_\{t\}:=\\min\\\{1,t/0\.3\\\}with a constant rate scaling factor of 0\.7\. Consequently, insertions are completed byt=0\.3t=0\.3, and the remaining trajectory refines the fixed set of generated coordinates\.

For motif scaffolding, motif contigs are provided as the model’s initial input\. Insertions are allowed between motif contigs and at the beginning and the end, but not within the contiguous motif\. The motif parts are centered and remain unchanged during the refinement for the continuous coordinates\. We track the lengths of each contig \(motifs and scaffolds\) to determine the new motif positions for the next input in the iterative sampling process\. The same sampling hyperparameters as in the unconditional setup are used, except for a rate scaling factor of 0\.6 to compensate for the longer length distribution in AFDB\.

We use theτ\\tau\-leaping technique described in Appendix[B\.5](https://arxiv.org/html/2607.09039#A2.SS5)to enable high\-quality sampling with the 400\-step sampling setup\. Simultaneous insertion at multiple places is allowed, and multiple insertions at the same place are also allowed, with the count sampled from the corresponding Poisson distribution\. In contrast to Proteina, classifier\-free guidance \(CFG\) and self\-conditioning are not used\.

#### C\.1\.3Evaluation Pipeline

In our evaluations of structure quality and diversity, the following metrics are used:

Designability\.For each protein structure, we first use ProteinMPNN\[[15](https://arxiv.org/html/2607.09039#bib.bib15)\]to design 8 candidate sequences, which are then folded by ESMFold\[[33](https://arxiv.org/html/2607.09039#bib.bib33)\]to give the refolded structures\. The self\-consistency root\-mean\-square deviation \(scRMSD\), defined as the RMSD between the generated and refolded structures, is used as the primary indicator of structural fidelity\. A*designable*protein structure in unconditional protein generation is defined as having a minimum scRMSD<<2Å among the 8 refolded candidates\. The scRMSD is always calculated on the CA coordinates due to Proteina’s CA\-only architecture\.*Designability*is defined as the proportion of designable protein structures in all generations\. The designability score assesses the quality of structural generation, since natural protein structures and scaffolds should be designable\.

Scaffold Designability\.For motif scaffolding, the following additional metrics are calculated:

- •motifRMSD, which is the RMSD between the structure of the refolded motif region and the ground\-truth motif structure specified as the condition\. We extract the motif residues in each refolded structure to calculate motifRMSD\. motifRMSD is calculated over the four backbone atoms \(N, C, CA, O\) instead of CA only\.
- •pLDDT\(predicted local distance difference test score\)\[[37](https://arxiv.org/html/2607.09039#bib.bib37)\], which estimates structural fidelity on a 0\-100 scale given the residue sequence\. pLDDT is calculated for the ProteinMPNN\-redesigned sequences\.
- •pAE\(predicted aligned error\), which measures the alignment error of residue pairs between the generated and refolded structures\.

Additionally, the amino acid sequences in the motif template are fixed and will not be redesigned by ProteinMPNN unless specified as unknown \(X\) in the template\. This ensures that motif information, such as enzymatic residues, is preserved during sequence redesign\. FollowingWatson et al\. \[[49](https://arxiv.org/html/2607.09039#bib.bib49)\], a scaffold is considered*successful*if at least one refolded structure \(out of 8 candidates\) satisfies scRMSD≤\\leq2Å, motifRMSD≤\\leq1Å, pLDDT≥\\geq70, and pAE≤\\leq5Å\. Successful designs are clustered as described below, and each resulting cluster counts as one*unique success*\.

Diversity\.*Diversity*is defined as the proportion of designable clusters with respect to the total number of designable samples\. All predicted structures are clustered with FoldSeek\[[46](https://arxiv.org/html/2607.09039#bib.bib46)\]at a specified TM\-score threshold\[[52](https://arxiv.org/html/2607.09039#bib.bib52)\], TM\-score is a length\-independent measure of global fold similarity\. We calculate the diversity using the following FoldSeek\[[46](https://arxiv.org/html/2607.09039#bib.bib46)\]command:

```
foldseek easy-cluster <path_samples> <path_tmp>/res <path_tmp> \
  --alignment-type 1 --cov-mode 0 --min-seq-id 0 \
  --tmscore-threshold <threshold>
```

where the threshold is 0\.5 for unconditional generation and 0\.6 for motif scaffolding\.

Novelty\.We assess the structural novelty score via exhaustive search against a structure database with FoldSeek\.*Novelty*is defined as the maximum TM\-score across all hits for each generated structure, averaged across all generations\. We calculate the novelty score using the following FoldSeek command:

```
foldseek easy-search <path_samples> <path_db> out.tsv <path_tmp> \
  --alignment-type 1 --exhaustive-search --tmscore-threshold 0.0 \
  --max-seqs 10000000000 --format-output query,target,alntmscore,lddt
```

where`path\_db`is the path prefix to either the full PDB structure database or the filtered 713k AFDB subset, resulting in the two score variants in Table[1](https://arxiv.org/html/2607.09039#S5.T1)\. The output TSV file contains results from the exhaustive search against the designated database, which is then processed to extract the highest TM\-score\.

Secondary structure composition\.Per\-residue three\-state secondary structure is assigned with DSSP\[[26](https://arxiv.org/html/2607.09039#bib.bib26)\]via the Biotite library\[[28](https://arxiv.org/html/2607.09039#bib.bib28)\]\. We report the average per\-protein helix \(α\\alpha\) and strand \(β\\beta\) fractions in the generated structures, and the coil fraction is given by1−α−β1\-\\alpha\-\\beta\.

For unconditional protein design, we calculate all metrics over 500 generations per model\. For GPFlow, the 500 lengths are determined by the learned rate\. Following the published fixed\-length benchmark, each baseline receives 100 requests at each length in\{50,100,150,200,250\}\\\{50,100,150,200,250\\\}\. We note that sampling baseline lengths from the empirical PDB marginal would instead constitute an oracle\-conditioned diagnostic because the fixed\-length models do not learnp​\(L\)p\(L\)\. We therefore use the standard grid for the primary end\-to\-end comparison and report GPFlow’s binned designability separately in Appendix[E](https://arxiv.org/html/2607.09039#A5)\.

For motif scaffolding, we compute the metrics on 1000 generated structures for each motif target, using FoldSeek’s clustering utility to identify the number of unique successes described above\. For motif tasks 5TRV, 6E6R, 6EXZ, and 7MRX, the original RFDiffusion benchmark provides contig templates across 3 different length ranges\. Each contig template specifies the minimum and maximum lengths of each linker segment and of the entire backbone\. For fixed\-length generative models like Proteina, given a contig template, the model first randomly samples the lengths of each linker segment to satisfy the contig constraints, then fixes the positions of these linker residues and redesigns their structures\. In this way, the length distribution in Figure[9](https://arxiv.org/html/2607.09039#A5.F9)exhibits 3 modes corresponding to the long, medium, and short length ranges\. For the baselines, we use the results from the best contig configuration on tasks 5TRV, 6E6R, 6EXZ, and 7MRX\. In contrast, GPFlow does not rely on predefined contig templates or length\-range constraints\. Instead, we sample 1000 backbones conditioned solely on the motif segments, yielding a continuous length distribution without the 3 distinct modes\.

#### C\.1\.4Baseline Information

For protein structure design, we compare GPFlow against the following baselines:

Proteina\[[18](https://arxiv.org/html/2607.09039#bib.bib18)\]is a CA\-only structure design model based on flow matching\. For unconditional generation on PDB, a 60M model is retrained on the same PDB dataset using the same training hyperparameters as our GPFlow for 300k iterations, with a total batch size of 128\. For AFDB, the official 60M unconditional Proteina checkpoint is used\. For motif scaffolding, we use the separate 60M Proteina motif architecture with the same AFDB training set and motif augmentation as our motif GPFlow\. Because the unconditional and motif architectures expose different inputs, their checkpoints are not interchangeable; each experiment therefore uses its corresponding Proteina base model\. For all Proteina models, we use exactly the same sampling hyperparameters as GPFlow: 400 sampling steps, a noise scale of 0\.35, and an exponential scheduler\. No self\-conditioning is used, as suggested in the default sampling config\.

FrameDiff\[[51](https://arxiv.org/html/2607.09039#bib.bib51)\]is a Riemannian diffusion model that learns a rotation \(orientation\) and a translation \(position\) for each amino acid in a protein\.FrameFlow\[[50](https://arxiv.org/html/2607.09039#bib.bib50)\]further extends it by incorporating Riemannian flow matching, with improved training stability and performance reported in the original work\. The results are directly taken from the Proteina paper, which runs inference using the provided checkpoint following the recommended hyperparameters\.

RFDiffusion\[[49](https://arxiv.org/html/2607.09039#bib.bib49)\]is a diffusion\-based protein backbone generative model that leverages the structure prediction power of RoseTTAFold\[[3](https://arxiv.org/html/2607.09039#bib.bib3)\]\. RFDiffusion also supports motif scaffolding by incorporating motifs as conditions\. The unconditional generation results for RFDiffusion are taken directly from the Proteina paper, whereas the motif scaffolding results are taken directly from the original paper\.

Genie2\[[31](https://arxiv.org/html/2607.09039#bib.bib31)\]is a diffusion\-based generative model with a specialized multi\-motif framework that designs co\-occurring motifs\. The motif scaffolding results are taken directly from the Proteina paper, which follows the recommended sampling hyperparameters in the original paper\.

### C\.2Protein Sequence Generation

#### C\.2\.1Data Specification

We train our sequence\-generation GPFlow on the UniRef50 database\[[44](https://arxiv.org/html/2607.09039#bib.bib44)\], containing 41,887,589 representative protein sequences obtained by clustering UniProt\[[45](https://arxiv.org/html/2607.09039#bib.bib45)\]sequences at 50% sequence identity\. We follow the data processing in EvoDiff\[[33](https://arxiv.org/html/2607.09039#bib.bib33)\], which results in a dataset split into 41,546,293 training, 82,929 validation, and 48,941 test sequences\. We further filter training sequences to retain only those with lengths between 10 and 1024 residues, yielding 40,471,430 training sequences with a mean length of 247\.7 residues\. Each sequence is tokenized at the single\-residue level over a 21\-character alphabet comprising the 20 standard amino acids plus an unknown token X \(ACDEFGHIKLMNPQRSTVWYX\)\. No subsampling of long sequences is applied; all sequences within the length range are used as\-is\. For sequence\-based motif scaffolding, we finetune on a subset of UniRef50 filtered to sequences with AlphaFold2\[[25](https://arxiv.org/html/2607.09039#bib.bib25)\]pLDDT≥95\\geq 95and lengths between 10 and 200, yielding 480,143 sequences\.

#### C\.2\.2Model Specification

##### Architecture\.

We use the Diffusion Transformer \(DiT\)\[[38](https://arxiv.org/html/2607.09039#bib.bib38)\]architecture for both the insertion and deletion models\. Each model consists of 14 transformer blocks with a hidden dimension of 1600, 25 attention heads, and an MLP expansion ratio of 6\.0×\\times, totaling 642M parameters\. Dropout is set to 0\.1\. Length conditioning is performed via a continuous sinusoidal length embedding over the range\[10,1024\]\[10,1024\]\.

##### Training\.

Both models are trained with AdamW\[[27](https://arxiv.org/html/2607.09039#bib.bib27)\]\(β1=0\.9\\beta\_\{1\}=0\.9,β2=0\.95\\beta\_\{2\}=0\.95, no weight decay\) and a learning rate of3×10−43\\times 10^\{\-4\}with 2,000 warmup steps followed by a cosine annealing schedule \(ηmin=10−5\\eta\_\{\\min\}=10^\{\-5\}\)\. Gradient clipping is applied at a norm of 5\.0\. Dynamic batching is used with a budget of 80,000 tokens per batch with approximate random\-length sampling\. We use gradient accumulation over 4 micro\-batches, yielding an effective batch size of 320,000 tokens\. An EMA of the model parameters is maintained with a decay of 0\.9999 and is used for all evaluations\. Both the insertion and deletion models are trained for 150k gradient steps\.

To enable classifier\-free guidance during sampling, the length conditionccis randomly dropped with probability 0\.2 during training, so that a single model learns both the conditional edit rateutθ​\(x\|xt,c\)u\_\{t\}^\{\\theta\}\(x\|x\_\{t\},c\)and the unconditional rateutθ​\(x\|xt\)u\_\{t\}^\{\\theta\}\(x\|x\_\{t\}\)\. At sampling time, these are combined as described in Eq\.[93](https://arxiv.org/html/2607.09039#A3.E93)and[94](https://arxiv.org/html/2607.09039#A3.E94)\. All training was carried out on 4 NVIDIA RTX PRO 6000 GPUs, taking approximately 3 days\.

##### Sampling\.

Generation proceeds by simulating the learned CTMC from the empty sequence \(t=0t=0\) to the data distribution \(t=1t=1\) over a uniform time grid with step sizeτ=1/T\\tau=1/T\(Euler solver\)\. Each step independently samples the number of insertions at each positioniifrom a Poisson distribution with intensityλt,iins​\(Xt\)​τ\\lambda\_\{t,i\}^\{\\mathrm\{ins\}\}\(X\_\{t\}\)\\tau, and draws inserted token identities fromQt,iins\(⋅\|Xt\)Q\_\{t,i\}^\{\\mathrm\{ins\}\}\(\\cdot\|X\_\{t\}\)\. This first\-order approximation corresponds to the linearization1−e−λ​τ≈λ​τ1\-e^\{\-\\lambda\\tau\}\\approx\\lambda\\tau\.

To mitigate discretization errors that accumulate over these Euler steps, we employ the corrector procedure ofHavasi et al\. \[[20](https://arxiv.org/html/2607.09039#bib.bib20)\], which leverages both the forward rateutu\_\{t\}\(insertion\) and reverse rateu~t\\tilde\{u\}\_\{t\}\(deletion\)\. The combined rate\(1\+α\)​ut\+α​u~t\(1\+\\alpha\)u\_\{t\}\+\\alpha\\tilde\{u\}\_\{t\}preserves the marginal distributionptp\_\{t\}for any corrector strengthα≥0\\alpha\\geq 0\. Note that whenα=0\\alpha=0, this reduces to the standard forward\-only Euler sampler\. This allows the introduction of additional self\-correction without altering the target distribution\. Concretely, at each time steptkt\_\{k\}with step sizehh:

1. 1\.Insertion sub\-step\.For each positioniiof the current sequenceXtkX\_\{t\_\{k\}\}, we sample insertions from a Poisson distribution with intensity\(1\+α\)​λtk,iins​\(Xtk\)​τ\(1\+\\alpha\)\\lambda\_\{t\_\{k\},i\}^\{\\mathrm\{ins\}\}\(X\_\{t\_\{k\}\}\)\\tau\. Inserted token identities are drawn fromQtk,iins\(⋅\|Xtk\)Q\_\{t\_\{k\},i\}^\{\\mathrm\{ins\}\}\(\\cdot\|X\_\{t\_\{k\}\}\)\. This produces an intermediate stateYYthat overshoots byα​τ\\alpha\\tauin time\.
2. 2\.Deletion sub\-step\.For each positioniiofYY, we sample whether to delete that position with probabilityα​λ~tk′,idel​\(Y\)​τ\\alpha\\tilde\{\\lambda\}\_\{t\_\{k\}^\{\\prime\},i\}^\{\\mathrm\{del\}\}\(Y\)\\tau, wheretk′=tk\+\(1\+α\)​τt\_\{k\}^\{\\prime\}=t\_\{k\}\+\(1\+\\alpha\)\\tauandλ~del\\tilde\{\\lambda\}^\{\\mathrm\{del\}\}is the learned reverse \(deletion\) rate\. Deleted positions are removed, yieldingXtk\+1X\_\{t\_\{k\+1\}\}\. This pulls back byα​τ\\alpha\\tau, so the net time advancement isτ\\tau\.

We further apply classifier\-free guidance \(CFG\) independently to both the rate magnitudeλ\\lambdaand the token distributionQQfor the insertion model and only to the latter for the deletion model\. Following the formulation inHavasi et al\. \[[20](https://arxiv.org/html/2607.09039#bib.bib20)\], the guided quantities are

λ~t,i​\(xt,c\)\\displaystyle\\tilde\{\\lambda\}\_\{t,i\}\(x\_\{t\},c\)=λt,i​\(xt\|c\)1\+w​λt,i​\(xt\)−w,\\displaystyle=\\lambda\_\{t,i\}\(x\_\{t\}\|c\)^\{1\+w\}\\,\\lambda\_\{t,i\}\(x\_\{t\}\)^\{\-w\},\(93\)Q~t,i​\(a\|xt,c\)\\displaystyle\\tilde\{Q\}\_\{t,i\}\(a\|x\_\{t\},c\)∝Qt,i​\(a\|xt,c\)1\+w​Qt,i​\(a\|xt\)−w,\\displaystyle\\propto Q\_\{t,i\}\(a\|x\_\{t\},c\)^\{1\+w\}\\,Q\_\{t,i\}\(a\|x\_\{t\}\)^\{\-w\},\(94\)wherew≥0w\\geq 0is the guidance weight,ccis the target length, and the conditional and unconditional outputs are produced by the same model \(with and without the length condition, respectively\)\. Eq\.[93](https://arxiv.org/html/2607.09039#A3.E93)corresponds to log\-linear extrapolation in rate space,logλ~=\(1\+w\)logλ\(⋅\|c\)−wlogλ\(⋅\)\\log\\tilde\{\\lambda\}=\(1\+w\)\\log\\lambda\(\\cdot\|c\)\-w\\log\\lambda\(\\cdot\), amplifying the conditional rate relative to the unconditional baseline\. Eq\.[94](https://arxiv.org/html/2607.09039#A3.E94)applies log\-linear extrapolation in token\-distribution space, analogous to the rate guidance\. Together, guidance modulates both*where*edits occur \(viaλ~\\tilde\{\\lambda\}\) and*what*tokens are inserted \(viaQ~\\tilde\{Q\}\)\. Figure[8](https://arxiv.org/html/2607.09039#A5.F8)demonstrates the effectiveness of this length control mechanism\.

For the results in the main text, we generate 500 sequences usingT=5,000T=5\{,\}000steps with corrector strengthα=1\\alpha=1and guidance weightw=0\.75w=0\.75for both the forward and reverse models\. For the length distribution evaluation \(Figure[2](https://arxiv.org/html/2607.09039#S5.F2)\(e\)\), we generate a larger set of 1,000 sequences usingT=10,000T=10\{,\}000steps with the forward\-only Euler sampler \(α=0\\alpha=0\), nucleus sampling withp=0\.95p=0\.95, and temperature annealing fromτ=2\.0\\tau=2\.0toτ=0\.1\\tau=0\.1\. For sequence\-based motif scaffolding, generation starts from the motif tokens and proceeds for 5,000 steps with corrector strengthα=1\\alpha=1, with no classifier\-free guidance or length conditioning applied\.

#### C\.2\.3Evaluation Pipeline

We evaluate GPFlow samples against ground\-truth sequences in UniRef50 and samples generated by the baseline models\. For all baselines, we follow the DPLM\[[48](https://arxiv.org/html/2607.09039#bib.bib48)\]evaluation protocol to generate 500 sequences in total: 50 sequences at each of the 10 fixed lengthsL∈\{100,200,…,1000\}L\\in\\\{100,200,\\ldots,1000\\\}\. UniRef50 reference sequences are obtained from 500 randomly sampled training sequences\. We fold generated sequences into 3D structures with ESMFold\[[32](https://arxiv.org/html/2607.09039#bib.bib32)\]\. The predicted structures are then assessed with the metrics detailed below, following a similar protocol in structure design\.

Foldabilityis measured by the sequence pLDDT score, as folded by ESMFold\. We report the mean CA pLDDT across all residues of each structure, stratified by sequence length in bins of width 100 residues\.Noveltyis evaluated by performing an exhaustive structural search across the entire PDB database with FoldSeek\. The maximum TM\-score across all hits defines the pdb\-TM, and structures withpdb\-TM<0\.5\\text\{pdb\-TM\}<0\.5are classified as structurally novel\.Structural diversityis reported as the ratio of distinct clusters to the total number of generated structures clustered by FoldSeek with a TM\-score threshold of 0\.5\.Secondary structure compositionsusing DSSP are also reported to provide population\-level structural similarity\.

#### C\.2\.4Baseline Information

We compare GPFlow against four protein sequence generative models\. For each baseline, we use the authors’ official model checkpoints and follow the default sampling procedures\.

EvoDiff\[[1](https://arxiv.org/html/2607.09039#bib.bib1)\]provides two discrete diffusion variants\. EvoDiff\-OADM uses a masking corruption process that replaces one residue with a mask token at each step until the sequence is fully masked\. EvoDiff\-D3PM uses a categorical transition\-matrix corruption process that introduces substitution noise, so that the terminal distribution approaches a uniform amino\-acid sequence\. Sampling starts from a uniform amino\-acid initialization and iteratively denoises\. We report EvoDiff\-OADM, the better\-performing variant in the original evaluation\. The model is trained on UniRef50 sequences 10\-1024 residues long\. Sequences longer than 1024 residues are subsampled into 1024\-residue chunks\. In our evaluation, we use the EvoDiff\-OADM\-640M checkpoint\.

DPLM\[[48](https://arxiv.org/html/2607.09039#bib.bib48)\]is a discrete diffusion language model instantiated using an absorbing\-mask corruption process\. The corruption schedule replaces residues with a mask token and treats the mask as an absorbing noise state\. In our evaluation, we use the DPLM\-650M checkpoint trained on the same dataset asAlamdari et al\. \[[1](https://arxiv.org/html/2607.09039#bib.bib1)\]\. For sampling, we perform a 500\-step discrete diffusion denoising process starting from a fully masked sequence, using temperature annealing from 2\.0 to 0\.1, with stochastic unmasking and top\-p=0\.95p=0\.95filtering\.

ESM3\[[21](https://arxiv.org/html/2607.09039#bib.bib21)\]is a generative masked language model trained over multiple discrete token tracks that represent sequence, structure, and function\. In our experiments, we evaluate only the sequence track, which is trained on a mixture of large\-scale protein sequence databases including UniRef, MGnify, JGI, and OAS, together with sequences paired with structure datasets such as PDB, AlphaFoldDB, and ESMAtlas, and further augmented with inverse\-folded sequences\[[21](https://arxiv.org/html/2607.09039#bib.bib21)\]\. We generate sequences using 20\-step iterative masked language model decoding with a cosine unmasking schedule and temperature annealing\.

SCISOR\[[4](https://arxiv.org/html/2607.09039#bib.bib4)\]is a discrete diffusion model specialized for deletions\. Its forward noising process inserts random residues into natural protein sequences, and the learned reverse process removes residues to recover natural\-like sequences\. SCISOR is trained on the same dataset asAlamdari et al\. \[[1](https://arxiv.org/html/2607.09039#bib.bib1)\]\. For unconditional generation, we initialize with a random amino acid string and run a 10\-step reverse deletion diffusion procedure at temperature 1\.0\.

#### C\.2\.5Localized Probability Paths

In Algorithm[2](https://arxiv.org/html/2607.09039#alg2), step 4, each target positioniiis independently included inztz\_\{t\}with probabilityκt\\kappa\_\{t\}\. This corresponds to the standard factorized token\-wise mixture path\[[17](https://arxiv.org/html/2607.09039#bib.bib17)\], in which the inclusion of each position is an independent Bernoulli trial\. While simple, this factorized construction produces partial sequencesS^t\\hat\{S\}\_\{t\}consisting of non\-neighboring tokens interspersed with gaps\. For long sequences, the local context around each insertion point consists largely of gap positions, limiting the signal for predicting token identities\. Adapting the construction introduced inHavasi et al\. \[[20](https://arxiv.org/html/2607.09039#bib.bib20)\], we replace the factorized probability path with a localized construction that introduces spatial correlation: when a target position becomes present, its neighbors are encouraged to also become present, yielding local neighborhoods inS^t\\hat\{S\}\_\{t\}\.

##### Construction\.

LetX1=yX\_\{1\}=ydenote the target sequence length\. We introduce an auxiliary boolean indicator𝐦t∈\{true,false\}y\\mathbf\{m\}\_\{t\}\\in\\\{\\texttt\{true\},\\texttt\{false\}\\\}^\{y\}, where𝐦tj=true\\mathbf\{m\}\_\{t\}^\{j\}=\\texttt\{true\}means target positionjjis present in the partial sequenceS^t\\hat\{S\}\_\{t\}at timett\. Rather than sampling each𝐦tj\\mathbf\{m\}\_\{t\}^\{j\}independently, we construct𝐦t\\mathbf\{m\}\_\{t\}via an auxiliary space of Boolean variables𝐌∈\{true,false\}y×y\\mathbf\{M\}\\in\\\{\\texttt\{true\},\\texttt\{false\}\\\}^\{y\\times y\}, consisting ofyyindependent CTMC processes, one per row, all initialized at𝐌0=false\\mathbf\{M\}\_\{0\}=\\texttt\{false\}\. The dynamics of rowiiare:

ut​\(𝐌i,j\|𝐌ti\)=\(λtindep​δi​j\+𝟙\[𝐌ti,j−1∨𝐌ti,j\+1\]​λprop\)​\(𝟙\[𝐌i,j\]−δ𝐌i,j​\(𝐌ti,j\)\)\.u\_\{t\}\(\\mathbf\{M\}^\{i,j\}\|\\mathbf\{M\}\_\{t\}^\{i\}\)=\\Bigl\(\\lambda\_\{t\}^\{\\mathrm\{indep\}\}\\,\\delta\_\{ij\}\+\\mathds\{1\}\_\{\[\\mathbf\{M\}\_\{t\}^\{i,j\-1\}\\vee\\mathbf\{M\}\_\{t\}^\{i,j\+1\}\]\}\\,\\lambda^\{\\mathrm\{prop\}\}\\Bigr\)\\bigl\(\\mathds\{1\}\_\{\[\\mathbf\{M\}^\{i,j\}\]\}\-\\delta\_\{\\mathbf\{M\}^\{i,j\}\}\(\\mathbf\{M\}\_\{t\}^\{i,j\}\)\\bigr\)\.\(95\)The diagonal entry𝐌i,i\\mathbf\{M\}^\{i,i\}switches totrueat the time\-dependent rateλtindep=κ˙t/\(1−κt\)\\lambda\_\{t\}^\{\\mathrm\{indep\}\}=\\dot\{\\kappa\}\_\{t\}/\(1\-\\kappa\_\{t\}\), independently of all other positions\. Once𝐌i,i\\mathbf\{M\}^\{i,i\}is active, neighboring off\-diagonal entries𝐌i,j\\mathbf\{M\}^\{i,j\}\(j≠ij\\neq i\) switch totrueat a constant rateλprop\\lambda^\{\\mathrm\{prop\}\}whenever an adjacent entry \(𝐌ti,j−1\\mathbf\{M\}\_\{t\}^\{i,j\-1\}or𝐌ti,j\+1\\mathbf\{M\}\_\{t\}^\{i,j\+1\}\) is alreadytrue, propagating thetruestate outward from positionii\. We use a cubic scheduleκt=t3\\kappa\_\{t\}=t^\{3\}and setλprop=3\.0\\lambda^\{\\mathrm\{prop\}\}=3\.0\. The per\-position indicator is recovered by taking the column\-wise disjunction:

𝐦tj=𝐌t1,j∨𝐌t2,j∨⋯∨𝐌ty,j\.\\mathbf\{m\}\_\{t\}^\{j\}=\\mathbf\{M\}\_\{t\}^\{1,j\}\\vee\\mathbf\{M\}\_\{t\}^\{2,j\}\\vee\\cdots\\vee\\mathbf\{M\}\_\{t\}^\{y,j\}\.\(96\)Whenλprop=0\\lambda^\{\\mathrm\{prop\}\}=0, each row activates only its diagonal entry, recovering the factorized probability path in which each𝐦tj\\mathbf\{m\}\_\{t\}^\{j\}is independent\.

##### Effective rate\.

When computing the training loss, the rate at which positionjjtransitions from absent to present \(i\.e\.,𝐦tj\\mathbf\{m\}\_\{t\}^\{j\}switches totrue\) is no longer the uniformκ˙t/\(1−κt\)\\dot\{\\kappa\}\_\{t\}/\(1\-\\kappa\_\{t\}\)but an effective rate that depends on the local propagation state:

λj,teff=λtindep\+∑l=1y𝟙​\{𝐌tl,j−1∨𝐌tl,j\+1\}​λprop\.\\lambda\_\{j,t\}^\{\\mathrm\{eff\}\}=\\lambda\_\{t\}^\{\\mathrm\{indep\}\}\+\\sum\_\{l=1\}^\{y\}\\mathds\{1\}\\\{\\mathbf\{M\}\_\{t\}^\{l,j\-1\}\\vee\\mathbf\{M\}\_\{t\}^\{l,j\+1\}\\\}\\,\\lambda^\{\\mathrm\{prop\}\}\.\(97\)Positions adjacent to already\-present neighbors have a higher effective rate, while isolated positions receive only the independent rate\.

##### Training loss\.

The localized path modifies both the Poisson NLLℒPP\\mathcal\{L\}\_\{\\mathrm\{PP\}\}\(Eq\.[77](https://arxiv.org/html/2607.09039#A2.E77)\) and the reconstruction lossℒrec\\mathcal\{L\}\_\{\\mathrm\{rec\}\}\(Eq\.[78](https://arxiv.org/html/2607.09039#A2.E78)\) by replacing the uniform weightκ˙t/\(1−κt\)\\dot\{\\kappa\}\_\{t\}/\(1\-\\kappa\_\{t\}\)with the position\-dependent effective rate\. The modified Poisson NLL is:

ℒGPloc=𝔼t,Y1​\[∑i=0Xtλti−∑i=0Xt\(∑k∈Ωiλk,teff\)​log⁡λti\],\\mathcal\{L\}\_\{\\mathrm\{GP\}\}^\{\\mathrm\{loc\}\}=\\mathbb\{E\}\_\{t,Y\_\{1\}\}\\\!\\left\[\\sum\_\{i=0\}^\{X\_\{t\}\}\\lambda\_\{t\}^\{i\}\-\\sum\_\{i=0\}^\{X\_\{t\}\}\\biggl\(\\sum\_\{k\\in\\Omega\_\{i\}\}\\lambda\_\{k,t\}^\{\\mathrm\{eff\}\}\\biggr\)\\log\\lambda\_\{t\}^\{i\}\\right\],\(98\)and the modified reconstruction loss is:

ℒrecloc=𝔼Y1,Yt​\[−∑i=0Xt∑k∈Ωiλk,teff​log⁡ρθi​\(sk\|Yt\)\],\\mathcal\{L\}\_\{\\mathrm\{rec\}\}^\{\\mathrm\{loc\}\}=\\mathbb\{E\}\_\{Y\_\{1\},Y\_\{t\}\}\\\!\\left\[\-\\sum\_\{i=0\}^\{X\_\{t\}\}\\sum\_\{k\\in\\Omega\_\{i\}\}\\lambda\_\{k,t\}^\{\\mathrm\{eff\}\}\\,\\log\\rho\_\{\\theta\}^\{i\}\(s\_\{k\}\|Y\_\{t\}\)\\right\],\(99\)where the sum runs over theX1−XtX\_\{1\}\-X\_\{t\}not\-yet\-present positionskk\(grouped into binsΩi\\Omega\_\{i\}\), andρθi\\rho\_\{\\theta\}^\{i\}is the sampling distribution at insertion positionii\. The effective rateλk,teff\\lambda\_\{k,t\}^\{\\mathrm\{eff\}\}upweights reconstruction loss at positions adjacent to already\-present residues, encouraging the model to extend locally consistent subsequences rather than predict isolated positions\.

### C\.3Peptide Co\-Design

#### C\.3\.1Data Specification

We use the dataset derived in PepFlow\[[29](https://arxiv.org/html/2607.09039#bib.bib29)\], which contains filtered PDB structure entries with peptide lengths ranging from 3 to 25\. The full dataset contains 8,365 protein–peptide complex structures, which are clustered at 40% peptide sequence identity, yielding 292 unique clusters\. Ten clusters across different lengths are selected as the test set, comprising 158 structures\. The binding pocket is given as additional conditioning information, following the practice in PepFlow\.

Each data pair is a complex of a target receptor and its bound peptide\. Following PepFlow, the binding pocket is provided as a fixed conditioning context, and the peptide is the target for generation\. The peptide is modeled at the all\-atom level: a backbone frame \(rotation and translation\), the residue type, and up to four sidechain torsion angles per residue\. During training, uniform sampling over the training clusters is performed\.

#### C\.3\.2Model Specification

##### Architecture

We adapt PepFlow\[[29](https://arxiv.org/html/2607.09039#bib.bib29)\], an equivariant geometric GNN with heads to regress vector fields for all\-atom peptide co\-design\. PepFlow represents all\-atom structures with per\-residue rotation and translation, amino acid logit, and up to four sidechain torsion angles\. Built upon the IPA structure, PepFlow denoises these modalities together using \(Riemannian\) flow matching to co\-design the peptide conditioned on the target protein structure\. We enable variable\-length generation with an order\-preserving process: a rate head predicts how many residues to insert at each inter\-residue gap, so the peptide length is generated rather than conditioned on the native length\. Each inserted residue is predicted across all of PepFlow’s modalities with additional heads when such an insertion happens\.

##### Training

We optimize GPFlow with AdamW \(learning rate10−410^\{\-4\}, weight decay 0\.01\) at a total batch size of 48, keeping an EMA of the weights \(decay 0\.999\) for evaluation\. The model is trained across four NVIDIA RTX PRO 6000 GPUs for approximately two days\. We adopt two techniques to improve the early rate prediction during generation\.Noisy context learning \(NCL\)\.During training, the kept peptide context fed to the decoder is corrupted asxncl=w​xt\+\(1−w\)​ϵx\_\{\\mathrm\{ncl\}\}=w\\,x\_\{t\}\+\(1\-w\)\\,\\epsilon, withw∼𝒰​\[0,1\]w\\sim\\mathcal\{U\}\[0,1\]andϵ∼𝒩​\(0,I\)\\epsilon\\sim\\mathcal\{N\}\(0,I\), so the model learns to read structural guidance without relying on perfectly accurate context\.Scheduled sampling \(SS\)\.With probability0\.50\.5, the self\-conditioning channel is fed the decoder’s own clean\-data predictionxpred=xt\+\(1−t\)​vθ​\(xt,t,z\)x\_\{\\mathrm\{pred\}\}=x\_\{t\}\+\(1\-t\)\\,v\_\{\\theta\}\(x\_\{t\},t,z\)in place of the ground truth, exposing the model to its own output\.

##### Sampling

For each target, we generate 10 peptides over 250 sampling steps\. The all\-atom structure and sequence are produced with a low\-temperature score\-SDE over the continuous modalities \(noise scale 0\.05\), while the peptide length is not supplied and emerges from the learned insertion rate\.

#### C\.3\.3Evaluation Pipeline

We use the metrics in PepFlow to evaluate the quality of generation across various modalities\. For each baseline model, 10 samples are generated for each holdout target, and the metrics are averaged across all generations\.

##### AAR

Amino Acid Recovery \(AAR\) quantifies the sequence identity between the generated and the native peptide\. AAR is computed as the overlap ratio between the generated and the native sequence\. A higher AAR indicates a closer match between the generated and native peptides in terms of amino acid composition, reflecting the accuracy of the generated sequence\.

##### RMSD

To evaluate the generation quality of the peptides, the RMSD between the ground\-truth and generated peptides is calculated from their CA coordinates using the Kabsch algorithm\. A lower RMSD indicates a closer structural alignment between the generated and native peptides\.

##### SSR

The Secondary Structure Ratio \(SSR\) assesses the similarity between the secondary structure of the generated peptide and the ground\-truth peptide\. This metric is computed by determining the ratio of identical entries in the secondary structure labels of the two peptides\. The secondary structure labels are obtained using the DSSP software\. A higher SSR indicates a closer match in the secondary structure between the generated and native peptides\.

##### BSR

Binding Site Recovery \(BSR\) gauges the interaction similarity between the generated peptide–target pair and the native pair\. BSR assesses whether the generated peptide recognizes residues in the target protein similarly to the native peptide\. Following PepFlow, a residue is considered to be in the binding site if its CB atom is within a 6 Å radius of any peptide residue\. BSR is calculated as the overlap between the binding sites derived from the generated and native peptides\. A higher BSR indicates a closer match between the generated and native peptide\-protein interactions\.

##### Affinity

The Affinity ratio represents the percentage of designed peptides exhibiting lower binding energy, indicating a higher binding affinity to the target protein compared to the native peptide\. Following PepFlow, the binding energy is computed using theInterfaceAnalyzerMoverin PyRosetta\[[10](https://arxiv.org/html/2607.09039#bib.bib10)\], after relaxing the complex and defining the interface between the peptide and target protein\. A higher Affinity percentage indicates that the designed peptides exhibit enhanced binding affinity, suggesting potential improvements in their functional capabilities\.

##### Stability

The Stability ratio is defined as the proportion of designed peptides that exhibit a lower energy score compared to the native complex\. Following PepFlow, we utilize theFastRelaxmethod in PyRosetta, where each complex first undergoes relaxation, and the total score is evaluated using the REF2015 score function\. The Stability metric is then calculated as the ratio of complexes with reduced energy scores, highlighting the designed peptides that contribute to the enhanced stability in the protein\-peptide complex\.

##### Designability

Similar to previous structure design tasks, Designability assesses whether a generated peptide structure corresponds to a sequence that can fold into a structure similar to itself\. Specifically, we use ESMFold to predict the peptide structure and subsequently employ RosettaFlexPep\[[36](https://arxiv.org/html/2607.09039#bib.bib36),[39](https://arxiv.org/html/2607.09039#bib.bib39)\]to dock the predicted structure into the target\. The resulting docked structure is then compared with the native peptide structure, where a*designable*peptide should exhibit less than 2 Å RMSD to the native structure\. Designability is then reported as the ratio of designable peptides across all test targets\.

##### Diversity

Similar to previous structure design tasks, Diversity is quantified by calculating all the pairwise TM\-scores among the generated peptides for a given target using the original TM\-align program\[[53](https://arxiv.org/html/2607.09039#bib.bib53)\]\. The diversity metric is then defined as 1 minus the average TM\-score\. A higher diversity value indicates greater structural variation among the generated peptides, showcasing the extent of structural exploration in the design process\.

Table 6:Per\-subset breakdown of GPFlow’s composite AAR/RMSD/SSR\.
##### Adaptations to Variable\-Length Design

Among the metrics above, AAR, RMSD, and SSR are computed per residue and therefore require a correspondence between the generated peptide of lengthnnand the native peptide of lengthmm\. Whenn=mn=m, this is the identity correspondencei↦ii\\mapsto i\. Since GPFlow is a variable\-length model, the generated peptide lengthnndiffers frommmon67\.9%67\.9\\%of samples, and no canonical identity correspondence exists\. We therefore align the two peptides with an order\-preserving map\. Whenn=mn=m, we use the identity correspondence; otherwise, we compute a Needleman\-Wunsch alignment with zero gap cost, which yields a monotone matching ofmin⁡\(n,m\)\\min\(n,m\)residue pairs\. We average each metric over the matched residues only and exclude the\|n−m\|\|n\-m\|unmatched residues, so the three metrics do not penalize the length difference, which is instead reported as a separate length statistic\. The AAR, RMSD, and SSR computed under this scheme are reported in Table[6](https://arxiv.org/html/2607.09039#A3.T6)\.

#### C\.3\.4Baseline Information

In addition to the base PepFlow model, for which we use the official PepFlow checkpoint for evaluation with 200 flow sampling steps, we also compare GPFlow instantiation against two additional peptide co\-design models or pipelines\. We follow PepFlow’s configurations for all three baselines for fair comparison\.

##### RFDiffusion

The same RFDiffusion\[[49](https://arxiv.org/html/2607.09039#bib.bib49)\]model used in motif scaffolding is also used here for conditional peptide design\. The official RFDiffusion implementation is used with 200 diffusion steps\. ProteinMPNN\[[15](https://arxiv.org/html/2607.09039#bib.bib15)\]is subsequently applied to redesign the amino acid types for evaluation\.

##### ProteinGenerator

ProteinGenerator\[[35](https://arxiv.org/html/2607.09039#bib.bib35)\]is built on RFDiffusion and additionally supports co\-designing protein backbone structures and sequences\. The official inference scripts are used with 200 diffusion steps\. For a fair comparison, no additional hotspot or DSSP information is fed into the model\.

## Appendix DAblation Studies

In this section, we ablate the key design choices of GPFlow on the unconditional structure generation task using the PDB dataset, reporting the same metrics as in Table[1](https://arxiv.org/html/2607.09039#S5.T1)\.

### D\.1Insertion Scheduler and Sampling Step Impact

Table[7](https://arxiv.org/html/2607.09039#A4.T7)presents an ablation study of the length\-insertion schedulerκt\\kappa\_\{t\}and the number of sampling steps\. Both the linear scheduler and the early\-insertion scheduler \(κt:=min⁡\(t/0\.6,1\)\\kappa\_\{t\}:=\\min\(t/0\.6,1\), used in our main results\) yield comparably high designability \(96\.796\.7and96\.196\.1, respectively\), whereas the late\-insertion scheduler \(κt:=t2\\kappa\_\{t\}:=t^\{2\}\) noticeably degrades designability to91\.691\.6\. This suggests that growing the structure earlier in the generative trajectory leaves more steps for the continuous vector field to refine the inserted coordinates\. For the number of sampling steps, increasing from200200to400400substantially improves designability \(89\.3→96\.789\.3\\to 96\.7\), while800800steps provide no further gain \(95\.795\.7\); we therefore adopt400400steps as a balance between generation quality and sampling cost\.

Table 7:Ablation studies of different GPFlow variants on the PDB dataset\.
### D\.2Designability\-Diversity Trade\-Off

Table 8:Ablation studies of different sampling\-noise scalesγ\\gammafor GPFlow on the PDB dataset\.Table[8](https://arxiv.org/html/2607.09039#A4.T8)reports the effect of the sampling noise scaleγ\\gammain the SDE sampler \(Eq\.[81](https://arxiv.org/html/2607.09039#A2.E81)\)\. A smallerγ\\gammamimics low\-temperature sampling, increasing designability at the cost of diversity \(e\.g\.,γ=0\.3\\gamma=0\.3attains97\.797\.7designability but only0\.270\.27diversity\), whereas a largerγ=0\.5\\gamma=0\.5reverses this trade\-off \(82\.682\.6designability,0\.450\.45diversity\)\. We adoptγ=0\.35\\gamma=0\.35as the default, which balances high designability with reasonable diversity\. This trade\-off also explains the generally lower diversity of the Proteina\-family models reported in Table[1](https://arxiv.org/html/2607.09039#S5.T1)\.

## Appendix EAdditional Results

In this section, we provide additional experimental results and visualizations of the GPFlow generations to further validate its performance and generation quality\.

### E\.1Unconditional Protein Generation

![Refer to caption](https://arxiv.org/html/2607.09039v1/figs/length_design.png)Figure 7:Designability per length range for generated structures\.For unconditional protein structure design, Figure[7](https://arxiv.org/html/2607.09039#A5.F7)reports GPFlow’s designability within each generated length bin\. Designability remains above 93% across all bins, indicating that the aggregate result is not driven solely by a narrow length range\. Therefore, we treat it as strong evidence that the superior designability of GPFlow is not due to length mismatch but rather to its capture of the joint distribution over length and continuous coordinates\.

### E\.2Sequence\-Based Length Control

In protein design settings where the desired length is knowna priori, it would be useful to generate samples of a specified length\. To this end, we evaluate the effectiveness of using classifier\-free guidance \(see Figure[8](https://arxiv.org/html/2607.09039#A5.F8)\) for controlling protein length\. Specifically, we generate 100 sequences at each target length from 100 to 1000 in increments of 100, using a guidance weight ofw=0\.75w=0\.75and 5,000 Euler steps\. As shown in Figure[8](https://arxiv.org/html/2607.09039#A5.F8), the generated lengths closely match the specified targets across the full range, demonstrating that GPFlow is capable of accurate and reliable protein length control\.

![Refer to caption](https://arxiv.org/html/2607.09039v1/figs/length_control.png)Figure 8:Evaluation of protein length control via classifier\-free guidance\. Each ridge shows the distribution of generated sequence lengths for a given target length, with the mean±\\pmstandard deviation annotated\.
### E\.3Structure\-Based Motif Scaffolding

In Figure[9](https://arxiv.org/html/2607.09039#A5.F9), we showcase the length distribution across all 16 motif tasks generated by GPFlow and Proteina\. Compared to Proteina, our method yields a broader and smoother length distribution, reflecting its ability to adaptively generate backbones of appropriate lengths conditioned on the motif target, rather than relying on predefined contig templates\. Remarkably, GPFlow can generate successful scaffolds longer than 200 residues, whereas Proteina generations are strictly limited by the contig templates, with most generations clustered around 75 residues\.

We additionally demonstrate the per\-task length distributions for representative tasks 1YCR, 3IXT, 5TPN, and 6E6R\. For each task, the raw success lengths \(red\) and the average length within each unique success \(blue\) are plotted on the same scale in Figure[10](https://arxiv.org/html/2607.09039#A5.F10)\. Because the sampling hyperparameters are identical across tasks, these results suggest that GPFlow adapts its generated lengths to the motif input, spanning a wider range for 3IXT while concentrating on a shorter\-length peak for 6E6R\. The longer unique successes observed for 3IXT and 5TPN further indicate that the model can produce successful and diverse scaffolds across a broad length range\.

![Refer to caption](https://arxiv.org/html/2607.09039v1/figs/motif_len.png)Figure 9:Aggregate protein length distribution across all 16 motif scaffolding tasks for GPFlow and Proteina\.![Refer to caption](https://arxiv.org/html/2607.09039v1/figs/motif_len_example.png)Figure 10:Protein length distributions for representative tasks for GPFlow\. Raw success lengths and the average length within each unique success are plotted on the same scale\.
### E\.4Sequence\-Based Motif Scaffolding

We adapt our 642M sequence\-generation GPFlow with the localized probability path described in Appendix[C\.2\.5](https://arxiv.org/html/2607.09039#A3.SS2.SSS5)to perform sequence\-based motif scaffolding\. During inference, generation starts from the motif tokens att=0t=0, and the model grows a scaffold through progressive insertions with motif positions frozen\. The learned deletion model serves as a corrector to refine the scaffold\. The scaffold length emerges naturally from the learned rate function\.

Table 9:Sequence\-based zero\-shot motif scaffolding success rates \(100 generations per task\)\. Superscripts denote the number of motif segments\. Best results are shown inbold\.We compare against two sequence\-based baselines: EvoDiff\-OADM\[[1](https://arxiv.org/html/2607.09039#bib.bib1)\]and DPLM\[[48](https://arxiv.org/html/2607.09039#bib.bib48)\], using their official checkpoints and inference protocols\. A design is successful if CA motif RMSD<<1Å and CA pLDDT\>\>70 as evaluated by ESMFold\[[33](https://arxiv.org/html/2607.09039#bib.bib33)\]\. We generate 100 candidates per task across the 17 motif problems adapted from the RFdiffusion benchmark\[[49](https://arxiv.org/html/2607.09039#bib.bib49)\]to the sequence setting byAlamdari et al\. \[[1](https://arxiv.org/html/2607.09039#bib.bib1)\]\.

As shown in Table[9](https://arxiv.org/html/2607.09039#A5.T9), GPFlow passes 10 out of 17 tasks, compared to 5/17 for EvoDiff and 6/17 for DPLM\. Notably, GPFlow uniquely solves four tasks \(5WN9, 2KL8, 4ZYP, 6VW1\) where both baselines achieve zero success\. On tasks where multiple methods succeed, GPFlow shows substantial gains: 87% on 3IXT versus 2% \(EvoDiff\) and 23% \(DPLM\); 59% on 7MRX versus 0% and 22%\.

### E\.5Peptide Co\-Design

We compare the lengths of GPFlow\-generated samples with those of native peptides\. In Figure[11](https://arxiv.org/html/2607.09039#A5.F11)\(left\), the generated and native length distributions agree closely\. We further compare the generated length with the native length for each target \(Figure[11](https://arxiv.org/html/2607.09039#A5.F11), right\)\. GPFlow recovers the lengths of short peptides nearly exactly; however, its conditional mean increasingly falls below the identity line as the native length increases, under\-generating long peptides by approximately1\.61\.6residues for native lengths above1212\. A possible contributing factor is the limited conditioning signal: GPFlow observes only the binding pocket and is not given the target peptide length\. Thus, for longer peptides, residues extending beyond the pocket interface may be less directly constrained, potentially making their lengths harder to recover accurately\.

![Refer to caption](https://arxiv.org/html/2607.09039v1/figs/peptide_length_combined.png)Figure 11:Peptide length\.Left: length distributions of GPFlow\-generated and native peptides\.Right: mean generated length at each native length; the dashed line marks exact recovery\.

## Appendix FBroader Impacts

This work introduces Generalized Poisson Flow \(GPFlow\), a framework for variable\-length generation evaluated on protein design\. Its connection between generalized Poisson processes and probability flows may also be relevant to other domains with variable\-dimensional data, although such applications are not evaluated here\. In computational biology, GPFlow addresses the dependence of structural feasibility on sequence length in*de novo*protein design\. Learning a distribution over backbone lengths may reduce manual length searches in motif\-scaffolding applications relevant to vaccine and enzyme design\. As with any advanced protein design technology, there is a potential for misuse, such as the design of harmful biological agents or toxins\. The generalizability of our framework implies that it could potentially be applied to biological systems with unknown or hazardous properties\. We are dedicated to ensuring the responsible use of our model for societal benefit\.

Similar Articles

DanceOPD: On-Policy Generative Field Distillation

Hugging Face Daily Papers

DanceOPD proposes an on-policy generative field distillation framework for flow-matching models that unifies text-to-image generation, local editing, and global editing via capability-specific routing and velocity-based training, improving multi-capability composition while preserving anchor generation quality.