Design-CP: Context Parallelism for Design of Protein Nanoparticles

arXiv cs.LG Papers

Summary

Design-CP introduces context-parallel inference strategies for RFdiffusion 3 that enable the all-atom design of large multimeric protein nanoparticles by distributing quadratic activations across multiple GPUs, making large-assembly protein design feasible on smaller GPU clusters.

arXiv:2607.05439v1 Announce Type: new Abstract: Many all-atom generative protein models can in principle design large multimeric complexes by jointly modelling all chains, but their quadratic token- and atom-pair representations quickly exceed single-GPU memory as the number of chains and residues modelled grows. We introduce Design-CP, two context-parallel (CP) inference strategies for RFdiffusion 3 (1D row-sharding and 2D grid sharding with ring attention) that distribute the quadratic activations across a multi-GPU mesh while preserving pretrained weights. We characterise their scaling when sampling icosahedral assemblies, showing that the maximum feasible asymmetric subunit (ASU) size grows with the expected square-root trend in GPU count and that 2D sharding achieves better wall-clock scaling. Moreover, we show how strong point-group symmetry constraints make CP usable out of the box for end-to-end, all-atom design of icosahedral nanoparticles, yielding favourable in silico structural and interface metrics. Finally, we demonstrate octahedral nanoparticle design on a small cluster of workstation-grade 16GB GPUs, illustrating how Design-CP can be a practical path towards democratising large-assembly protein design.
Original Article
View Cached Full Text

Cached at: 07/08/26, 04:43 AM

# Design-CP: Context Parallelism for Design of Protein Nanoparticles
Source: [https://arxiv.org/html/2607.05439](https://arxiv.org/html/2607.05439)
###### Abstract

Many all\-atom generative protein models can in principle design large multimeric complexes by jointly modelling all chains, but their quadratic token\- and atom\-pair representations quickly exceed single\-GPU memory as the number of chains and residues modelled grows\. We introduce*Design\-CP*, two context\-parallel \(CP\) inference strategies for RFdiffusion 3 \(1D row\-sharding and 2D grid sharding with ring attention\) that distribute the quadratic activations across a multi\-GPU mesh while preserving pretrained weights\. We characterise their scaling when sampling icosahedral assemblies, showing that the maximum feasible asymmetric subunit \(ASU\) size grows with the expected square\-root trend in GPU count and that 2D sharding achieves better wall\-clock scaling\. Moreover, we show how strong point\-group symmetry constraints make CP usable out of the box for end\-to\-end, all\-atom design of icosahedral nanoparticles, yielding favourable*in silico*structural and interface metrics\. Finally, we demonstrate octahedral nanoparticle design on a small cluster of workstation\-grade 16 GB GPUs, illustrating how Design\-CP can be a practical path towards democratising large\-assembly protein design\.

Machine Learning, BioML, Protein Design, Parallelism, Protein Nanoparticles

## 1Introduction

Deep learning is transforming computational protein design from a predominantly physics\-based endeavour into a data\-driven discipline\. Families of structure prediction models such as AlphaFold\(Jumperet al\.,[2021](https://arxiv.org/html/2607.05439#bib.bib160); Abramsonet al\.,[2024](https://arxiv.org/html/2607.05439#bib.bib12)\), RoseTTAFold\(Baeket al\.,[2021](https://arxiv.org/html/2607.05439#bib.bib8),[2023](https://arxiv.org/html/2607.05439#bib.bib147); Corleyet al\.,[2025](https://arxiv.org/html/2607.05439#bib.bib138)\)and Boltz\(Wohlwendet al\.,[2024](https://arxiv.org/html/2607.05439#bib.bib89); Passaroet al\.,[2025](https://arxiv.org/html/2607.05439#bib.bib3)\)now achieve near\-experimental accuracy on many single\-chain targets\. In parallel, a rapidly expanding family of denoising\-based generative design frameworks including RFDiffusion\(Watsonet al\.,[2023](https://arxiv.org/html/2607.05439#bib.bib159); Butcheret al\.,[2025](https://arxiv.org/html/2607.05439#bib.bib143)\), Chroma\(Ingrahamet al\.,[2023](https://arxiv.org/html/2607.05439#bib.bib9)\), Genie\(Lin and AlQuraishi,[2023](https://arxiv.org/html/2607.05439#bib.bib38); Linet al\.,[2024](https://arxiv.org/html/2607.05439#bib.bib52)\), and Proteina\(Geffneret al\.,[2025b](https://arxiv.org/html/2607.05439#bib.bib41),[a](https://arxiv.org/html/2607.05439#bib.bib140); Didiet al\.,[2026](https://arxiv.org/html/2607.05439#bib.bib1)\)enable the*de novo*creation of proteins with prescribed structural and functional properties\. These advances have already yielded a tangible impact across diverse application domains\. In therapeutic design alone, examples include de novo minibinders against therapeutically relevant targets such as bioactive peptide hormones\(Vázquez Torreset al\.,[2024](https://arxiv.org/html/2607.05439#bib.bib43)\)and bacterial toxins\(Ragotteet al\.,[2025](https://arxiv.org/html/2607.05439#bib.bib42)\)as well as the de novo design of epitope\-targeted antibodies, from diffusion\-based co\-design of CDR sequence and structure on a fixed framework\(Luoet al\.,[2022](https://arxiv.org/html/2607.05439#bib.bib121)\)to atomically accurate in\-silico design of VHHs and scFvs\(Bennettet al\.,[2026](https://arxiv.org/html/2607.05439#bib.bib120)\)\.

The success of these methods motivates scaling such generative tools to larger and more biologically complex targets, but this requires the ability to reliably design*multimeric*protein complexes\. In nature, the majority of proteins carry out their functions not as isolated monomers but as oligomeric assemblies, such as homodimers, heteromeric complexes, and higher\-order symmetric architectures\(Goodsell and Olson,[2000](https://arxiv.org/html/2607.05439#bib.bib115); Marsh and Teichmann,[2015](https://arxiv.org/html/2607.05439#bib.bib114)\)\. Designing symmetric assemblies*de novo*could unlock applications ranging from biomolecular machines inspired by rotary motors such as ATP synthase\(Courbetet al\.,[2022](https://arxiv.org/html/2607.05439#bib.bib111)\)to vaccine scaffolds inspired by viral capsids\(Butterfieldet al\.,[2017](https://arxiv.org/html/2607.05439#bib.bib110); Marcandalliet al\.,[2019](https://arxiv.org/html/2607.05439#bib.bib150); Wallset al\.,[2020](https://arxiv.org/html/2607.05439#bib.bib109)\)\. The computational design of such assemblies has so far relied on rigid\-body docking of independently\-designed oligomers\(Kinget al\.,[2012](https://arxiv.org/html/2607.05439#bib.bib37); Baleet al\.,[2016](https://arxiv.org/html/2607.05439#bib.bib36); Hsiaet al\.,[2016](https://arxiv.org/html/2607.05439#bib.bib66); Sheffleret al\.,[2023](https://arxiv.org/html/2607.05439#bib.bib39)\), a paradigm that recent ML\-era pipelines\(De Haaset al\.,[2024](https://arxiv.org/html/2607.05439#bib.bib65); Haaset al\.,[2025](https://arxiv.org/html/2607.05439#bib.bib64),[2026](https://arxiv.org/html/2607.05439#bib.bib60)\)have refined but not fundamentally replaced\. Crucially, this reliance on docking is often a practical workaround rather than a modelling choice: end\-to\-end all\-atom generators exist, but they struggle to fit whole assemblies in memory when many subunits must be modelled jointly\.

Recent all\-atom generative models such as RFDiffusion 3\(Butcheret al\.,[2025](https://arxiv.org/html/2607.05439#bib.bib143)\)\(RFD3\), which is inspired by the AlphaFold 3 architecture\(Abramsonet al\.,[2024](https://arxiv.org/html/2607.05439#bib.bib12)\)\(AF3\), can in principle generate multimeric structures by jointly modelling all chains with an atomistic level of precision and designing proteins with predefined point\-group symmetries\. However, the underlying architecture maintains pairwise representations whose memory cost scales quadratically with the number of tokensII\(and atomsLL\)\. For large protein assemblies,IIandLLgrow linearly with the number of chains modelled, causing the𝒪\\mathcal\{O\}\(I2I^\{2\}\) and𝒪\\mathcal\{O\}\(L2L^\{2\}\) pairwise feature tensors and related quadratic intermediates to exceed the memory capacity of a single GPU\. This practical bottleneck heavily limits the size of what can be designed, particularly when modelling symmetric protein assemblies\. This single\-device ceiling, however, contrasts with broader trends in computational infrastructure: while single\-device memory capacity has grown only incrementally, access to multi\-GPU clusters is now routine for both academic and industrial research groups\. Accordingly, partitioning large transformer workloads across such clusters has become standard practice in adjacent fields such as large language modelling, both for training\(Liet al\.,[2022](https://arxiv.org/html/2607.05439#bib.bib40)\)and inference\(Popeet al\.,[2022](https://arxiv.org/html/2607.05439#bib.bib87)\)\.

We implement and compare two context\-parallel \(CP\) inference strategies for RFD3, which we call*Design\-CP*\. The first, a*1D*scheme, stripes the pair representation acrossPPGPUs along a single axis; the second, a*2D*scheme following Fold\-CP\(Linet al\.,[2026](https://arxiv.org/html/2607.05439#bib.bib82)\), tiles it over aP×P\\sqrt\{P\}\\times\\sqrt\{P\}device grid with ring attention\. Both shard the dominant quadratic memory cost while preserving numerical equivalence with single\-GPU inference\. We evaluate the two schemes on large point\-group\-symmetric assemblies, including icosahedral and octahedral nanoparticles, and show that symmetry constraints sharpen practical sample quality and make end\-to\-end all\-atom sampling tractable on modest multi\-GPU setups without additional training or fine\-tuning\.

#### Contributions\.

Our main contributions are:

- •Design\-CP: context\-parallel inference for RFD3 with strong scaling\.We introduce two CP schemes for RFD3 inference \(a lightweight*1D row\-sharding*strategy and a*2D grid*strategy based on Fold\-CP\(Linet al\.,[2026](https://arxiv.org/html/2607.05439#bib.bib82)\)\), and we characterise their memory ceilings and wall\-clock scaling when sampling large symmetric icosahedral assemblies\.
- •Symmetry makes CP usable out of the box for icosahedral design\.We show that imposing strong point\-group symmetry constraints sharpens practical sample quality beyond the native training crop and makes end\-to\-end, all\-atom generation of icosahedral nanoparticles tractable without retraining or fine\-tuning\.
- •Octahedral design on small GPUs\.We demonstrate that the same approach enables*de novo*design of octahedral nanoparticles on a small cluster of workstation\-grade GPUs, showcasing a workable route to making large\-assembly protein design broadly accessible\.

## 2Related Work

Two lines of prior work are directly relevant to Design\-CP: \(i\) methods that reduce the memory cost of AF3\-class architectures, which we build on technically, and \(ii\) computational pipelines for designing large symmetric protein assemblies, where Design\-CP aims to make a methodological contribution\. For the first, we briefly cover IO\-efficient attention on a single device, while devoting most of this section to distributed parallelism for structure models\. For the second, we situate Design\-CP within the broader landscape of protein nanoparticle design methods\.

#### IO\-efficient attention\.

FlashAttention\(Daoet al\.,[2022](https://arxiv.org/html/2607.05439#bib.bib107); Dao,[2023](https://arxiv.org/html/2607.05439#bib.bib106)\)computes exact attention with memory linear in sequence length \(𝒪​\(I\)\\mathcal\{O\}\(I\)or𝒪​\(L\)\\mathcal\{O\}\(L\)in our notation\) by tiling queries, keys, and values into SRAM\-resident blocks while maintaining running softmax statistics\(Milakov and Gimelshein,[2018](https://arxiv.org/html/2607.05439#bib.bib105)\), building on the earlier observation that attention admits a linear\-memory implementation\(Rabe and Staats,[2022](https://arxiv.org/html/2607.05439#bib.bib85)\)\. For architectures without pair representations, this suffices and motivates its use in protein language models like ESM3\(Hayeset al\.,[2025](https://arxiv.org/html/2607.05439#bib.bib156)\), which represent sequence, structure, and function as discrete tokens and condense pairwise geometry into a single SE\(3\)\-invariant geometric\-attention block at the input\. Flash‑IPA\(Liuet al\.,[2025](https://arxiv.org/html/2607.05439#bib.bib102)\)and FlashBias\(Wuet al\.,[2025](https://arxiv.org/html/2607.05439#bib.bib101)\)extend this idea to pair‑biased attention, such as the one present in AlphaFold‑3 Pairformer, by re\-expressing geometric and pairwise bias terms via additional low‑rank features concatenated into the query and key projections\. For complex learned pair biases, these methods generally require training auxiliary parameters and are not a weight‑preserving drop‑in for an arbitrary pretrained model\. These approaches are orthogonal to Design\-CP: they reduce the per\-block memory footprint on a single device, while we shard the persistent quadratic activations across devices, and the two could in principle be composed\.

#### Distributed parallelism for structure models\.

Among the standard parallelism axes \(data, tensor, pipeline, expert, activation\),*context parallelism*uniquely shards activations along the sequence dimension of every layer, which makes it a natural fit for the𝒪​\(I2\)\\mathcal\{O\}\(I^\{2\}\)pair tensor of AlphaFold\-class models\. FastFold\(Chenget al\.,[2023](https://arxiv.org/html/2607.05439#bib.bib84)\)introduced Dynamic Axial Parallelism \(DAP\) for AlphaFold 2, replicating parameters on every device and sharding activations along a single sequence axis at a time; ScaleFold\(Zhuet al\.,[2024](https://arxiv.org/html/2607.05439#bib.bib108)\)adopted DAP and scaled to 2048 H100s for training\. As the Evoformer interleaves row\- and column\-wise attention over the MSA track, DAP must insert an all\-to\-all communication step whenever the active axis flips, incurring six all\-to\-all redistributions per Evoformer block at inference, together with oneAllGatherin the outer\-product\-mean and two in the triangular updates\(Chenget al\.,[2023](https://arxiv.org/html/2607.05439#bib.bib84)\)\. These collectives contribute non\-trivially to inference latency at scale and increase transient memory relative to the steady\-state \(𝒪​\(I2/P\)\\mathcal\{O\}\(I^\{2\}/P\)\) sharded pair representation\.

Fold\-CP\(Linet al\.,[2026](https://arxiv.org/html/2607.05439#bib.bib82)\)generalises axis sharding to a full two\-dimensional context\-parallel strategy for AF3\-class models \(implementing it for Boltz\-2\(Passaroet al\.,[2025](https://arxiv.org/html/2607.05439#bib.bib3)\)\), extending Ring Attention\(Liuet al\.,[2023](https://arxiv.org/html/2607.05439#bib.bib129)\)into a Cannon\-style 2D ring tailored to dense triangular updates, while window\-batched atom attention is handled by a complementary shardwise kernel that keeps each window’s attention local to its rank\. The pair tensor is tiled over aP×P\\sqrt\{P\}\\times\\sqrt\{P\}device grid, so that rank\(r,c\)\(r,c\)holds the block indexed by residue ranges\[r⋅I/P,\(r\+1\)⋅I/P\]×\[c⋅I/P,\(c\+1\)⋅I/P\]\[r\\cdot I/\\sqrt\{P\},\(r\{\+\}1\)\\cdot I/\\sqrt\{P\}\]\\times\[c\\cdot I/\\sqrt\{P\},\(c\{\+\}1\)\\cdot I/\\sqrt\{P\}\]\. Triangle attention, triangle multiplication, attention\-with\-pair\-bias, pair\-weighted\-averaging, and outer\-product\-mean are each reformulated as ring algorithms in which𝐊\\mathbf\{K\}and𝐕\\mathbf\{V\}shards \(and, where required, triangular biases and masks\) circulate between neighbouring devices while a numerically\-stable tiled softmax merges partial outputs without ever materialising the fullI×II\\times Iattention on any rank\. The resulting steady\-state pair memory per device is \(𝒪​\(I2/P\)\\mathcal\{O\}\(I^\{2\}/P\)\), matching 1D axis\-sharded DAP\. However, under 1D sharding, triangular updates typically require collectives over all \(PP\) ranks \(e\.g\., to obtain the necessary key/value or bias shards along the active axis\), which can inflate the transient working\-set memory during attention/multiplication beyond the steady shard\. In Fold‑CP’s 2D tiling, collectives are restricted to a single row or column subgroup of size \(P\\sqrt\{P\}\)\. This reduces communication volume and confines the transient working\-set memory growth in triangular updates \(e\.g\., gathered/circulated K/V shards, triangle biases, masks, and other pair\-like intermediates\) to \(P\\sqrt\{P\}\)\-sized groups, rather than requiring global \(PP\)\-way collectives\.

To the best of our knowledge, context parallelism has so far been developed and evaluated exclusively for structure*prediction*\. We take this as motivation to apply the same techniques to generative*design*, and adopt RFdiffusion 3\(Butcheret al\.,[2025](https://arxiv.org/html/2607.05439#bib.bib143)\)as our target\. Our 2D scheme is a direct port of Fold\-CP\(Linet al\.,[2026](https://arxiv.org/html/2607.05439#bib.bib82)\)\. At the same time, we observe that RFD3 has no MSA processing or triangular operations, which makes a much lighter 1D row\-sharded scheme practical as well\. We implement both and compare them on symmetric\-design tasks\.

#### Computational design of protein nanoparticles\.

Kinget al\.\([2012](https://arxiv.org/html/2607.05439#bib.bib37)\)introduced the modern*dock\-and\-design*recipe in Rosetta: pre\-existing oligomeric building blocks with compatible point\-group symmetry are docked as rigid bodies along the rotational axes of the target architecture, sampling only the radial displacementrrand axial rotationω\\omega, and a new low\-energy interface is then sequence\-designed between them\. Extending the pipeline to nanoparticles of multiple components\(Kinget al\.,[2014](https://arxiv.org/html/2607.05439#bib.bib35)\)enabled scaling to megadalton\-scale icosahedral assemblies such as the 120\-subunit I53\-50\(Baleet al\.,[2016](https://arxiv.org/html/2607.05439#bib.bib36)\)and the hyperstable 60\-subunit I3\-01\(Hsiaet al\.,[2016](https://arxiv.org/html/2607.05439#bib.bib66)\)\. Recent work has progressively replaced individual modules of this pipeline in favour of ML\-based methods: ProteinMPNN substitutes for Rosetta in interface design\(De Haaset al\.,[2024](https://arxiv.org/html/2607.05439#bib.bib65)\); AlphaFold2 predictions of thermophilic homologs supply building blocks in place of experimental structures\(Haaset al\.,[2025](https://arxiv.org/html/2607.05439#bib.bib64)\); and RFdiffusion\-generated*de novo*oligomers now serve as the building\-block library\(Haaset al\.,[2026](https://arxiv.org/html/2607.05439#bib.bib60)\)\. Crucially, all of these methods still rely on a rigid\-body docking step over pre\-computed, independently generated oligomers\. A natural next step is to try to model the entire assembly with a single generative network, but the per\-device memory footprint of representing a megadalton\-scale complex at full\-atom resolution currently makes joint generative design infeasible on a single GPU\.

Design\-CP removes this single\-GPU memory barrier, enabling RFdiffusion 3 to jointly denoise all atoms of large multimeric proteins in a single trajectory: the first end\-to\-end, all\-atom generative design of symmetric protein assemblies that models the full set of inter\-ASU interactions without an intermediate docking step\.

![Refer to caption](https://arxiv.org/html/2607.05439v1/Figure_1.png)Figure 1:Symmetric design with Design\-CP\.a, Schematic depiction of the sharding techniques implemented in Design\-CP\. Every square represents a sub\-tensor of the self\-attention matrix, and its colour represents its assigned device\. 1D sharding partitions the queries across different GPUs, where they are used to calculate cross\-attention against all keys\. 2D sharding partitions both queries and keys and uses ring attention to compute attention scores on the fly\.b,Qualitative comparison of large\-design samples\.Designs of a 10800 amino acid protein using different symmetries\. Designs that exceed the native crop limits can show visible degradation when sampled without additional structure constraints; imposing a strong symmetry prior mitigates this effect by reducing the effective design space\.

## 3Methods

### 3\.1RFDiffusion 3 preliminaries and notation

RFDiffusion 3 \(RFD3\) jointly modelsIItokens \(each token representing, for example, a single residue or the heavy atom of a small molecule\) together withLLatoms\. Tokens carry a single\-track tensor𝐒∈ℝI×cs\\mathbf\{S\}\\in\\mathbb\{R\}^\{I\\times c\_\{s\}\}and a pair representation𝐙∈ℝI×I×cz\\mathbf\{Z\}\\in\\mathbb\{R\}^\{I\\times I\\times c\_\{z\}\}; atoms carry a single\-track𝐀∈ℝL×catom\\mathbf\{A\}\\in\\mathbb\{R\}^\{L\\times c\_\{\\text\{atom\}\}\}and an analogous pair representation𝐏∈ℝL×L×catompair\\mathbf\{P\}\\in\\mathbb\{R\}^\{L\\times L\\times c\_\{\\text\{atompair\}\}\}\. Within every token\-level attention block, a learned projection of𝐙\\mathbf\{Z\}produces a per\-head pair bias𝐁∈ℝI×I×H\\mathbf\{B\}\\in\\mathbb\{R\}^\{I\\times I\\times H\}that is added to the attention logits before the softmax,softmax⁡\(QK⊤/d\+𝐁\)\\operatorname\{softmax\}\(\\textbf\{QK\}^\{\\top\}/\\sqrt\{d\}\+\\mathbf\{B\}\)\. Atom\-level attention is*sparse*in computation: each query atom attends to a budget ofkkneighbours assembled from a small set of atoms close in sequence together with the spatially closest atoms, with usuallyk≪Lk\\\!\\ll\\\!L\. The relevant slice of𝐏\\mathbf\{P\}is gathered at thosekkindices and then projected to the per\-head bias, so the attention computation itself is𝒪​\(L​k\)\\mathcal\{O\}\(L\\,k\)\. The storage cost of𝐏\\mathbf\{P\}, however, remains quadratic, and the dense\[L,L,catompair\]\[L,L,c\_\{\\text\{atompair\}\}\]tensor is the dominant atom\-level memory consumer at the scales we target\.111The pre\-existing RFD3 inference codebase also exposes an optional low\-memory mode based on a chunked pairwise embedder that constructs𝐏\\mathbf\{P\}on the fly at thekkkNN indices, avoiding the dense materialisation; this is orthogonal to Design\-CP and we describe it in Appendix[A\.1](https://arxiv.org/html/2607.05439#A1.SS1)\.In this work, we successfully partitioned the𝒪​\(I2\)\\mathcal\{O\}\(I^\{2\}\)pair track𝐙\\mathbf\{Z\}and the𝒪​\(L2\)\\mathcal\{O\}\(L^\{2\}\)pair track𝐏\\mathbf\{P\}acrossNNGPUs while never communicating full pair tensors\. A fuller description of the five\-stage RFD3 architecture and of the𝐏\\mathbf\{P\}–kNN interaction is deferred to Appendix[A\.1](https://arxiv.org/html/2607.05439#A1.SS1)\.

When designing with a pre\-defined point symmetry, the diffusion model is always run on the entire complex at every denoising step\. For a configurable fraction of the trajectory \(default 90%\), RFD3 then*resymmetrises*its prediction by extracting the coordinates of a single asymmetric subunit \(ASU\) and generating the remaining copies by applying the point\-group operations\. The resulting resymmetrised complex is the state that is fed into the next iteration of the sampling loop\. Additional details on the symmetrisation procedure are deferred to Appendix[A\.2](https://arxiv.org/html/2607.05439#A1.SS2)\.

### 3\.21D row\-sharding

Our first scheme, implemented in PyTorch with NCCL and without custom kernels or DTensor machinery, partitions the query dimension of everyI×II\\times IandL×LL\\times Loperation acrossPPGPUs \([Figure1](https://arxiv.org/html/2607.05439#S2.F1)a\)\. GPUppmaterialises only its row stripe of the pair track:

𝐙\(p\)∈ℝIp×I×cz,Ip=⌊I/P⌋\+𝕀​\[p<ImodP\],\\mathbf\{Z\}^\{\(p\)\}\\in\\mathbb\{R\}^\{I\_\{p\}\\times I\\times c\_\{z\}\},\\qquad I\_\{p\}=\\lfloor I/P\\rfloor\+\\mathbb\{I\}\[p<I\\bmod P\],\(1\)where𝕀​\[⋅\]∈\{0,1\}\\mathbb\{I\}\[\\,\\cdot\\,\]\\in\\\{0,1\\\}denotes the indicator function of its predicate\. WhenIIis not divisible byPP, the floor⌊I/P⌋\\lfloor I/P\\rfloorleaves a remainder ofImodPI\\bmod Pelements that must still be assigned\. We absorb this remainder by giving the firstImodPI\\bmod Pranks one additional row each: rankppreceives the extra row precisely whenp<ImodPp<I\\bmod P, which is what the indicator encodes\. By construction∑pIp=I\\sum\_\{p\}I\_\{p\}=I, no element is dropped, and chunk sizes differ by at most one across GPUs, so the load imbalance per attention block is bounded by a single row regardless ofPP\. The pair tracks𝐙\(p\)\\mathbf\{Z\}^\{\(p\)\}and the self\-conditioning distogram𝐃self\(p\)\\mathbf\{D\}^\{\(p\)\}\_\{\\text\{self\}\}are in this way*never*gathered to their full\[I,I\]\[I,I\]shape\.

Every self\-attention block is reformulated as a cross\-attention: GPUppprojects queries from its stripe𝐐\(p\)=fQ​\(𝐒\(p\)\)∈ℝIp×H×d\\mathbf\{Q\}^\{\(p\)\}=f\_\{Q\}\(\\mathbf\{S\}^\{\(p\)\}\)\\in\\mathbb\{R\}^\{I\_\{p\}\\times H\\times d\}, while keys and values are projected from the*full*single\-track𝐒\\mathbf\{S\}, which is replicated on every GPU\. The pair bias is drawn from the local stripe,𝐁\(p\)=fB​\(𝐙\(p\)\)∈ℝIp×I×H\\mathbf\{B\}^\{\(p\)\}=f\_\{B\}\(\\mathbf\{Z\}^\{\(p\)\}\)\\in\\mathbb\{R\}^\{I\_\{p\}\\times I\\times H\}, and

Attn\(p\)=softmax⁡\(𝐐\(p\)​𝐊⊤/d\+𝐁\(p\)\)​𝐕\\operatorname\{Attn\}^\{\(p\)\}=\\operatorname\{softmax\}\\\!\\left\(\\mathbf\{Q\}^\{\(p\)\}\\mathbf\{K\}^\{\\top\}/\\sqrt\{d\}\+\\mathbf\{B\}^\{\(p\)\}\\right\)\\mathbf\{V\}\(2\)A singleAllGatherover the 1D track reconstructs𝐒=⨁p𝐒\(p\)∈ℝI×cs\\mathbf\{S\}=\\bigoplus\_\{p\}\\mathbf\{S\}^\{\(p\)\}\\in\\mathbb\{R\}^\{I\\times c\_\{s\}\}after each block, so that all GPUs hold identical K and V for the next block\. The same partition applies at the atom level\. The dense atom pair track is striped as𝐏\(p\)∈ℝLp×L×catompair\\mathbf\{P\}^\{\(p\)\}\\in\\mathbb\{R\}^\{L\_\{p\}\\times L\\times c\_\{\\text\{atompair\}\}\}, reducing its per\-GPU storage from𝒪​\(L2\)\\mathcal\{O\}\(L^\{2\}\)to𝒪​\(L2/P\)\\mathcal\{O\}\(L^\{2\}/P\)\. The sparse kNN sequence\-local structure\-local attention is then applied per\-shard: each GPU gathers, from its stripe of𝐏\(p\)\\mathbf\{P\}^\{\(p\)\}, thekkneighbour entries for each of itsLpL\_\{p\}query atoms, yielding a local bias𝐏sparse\(p\)∈ℝLp×k×catompair\\mathbf\{P\}^\{\(p\)\}\_\{\\text\{sparse\}\}\\in\\mathbb\{R\}^\{L\_\{p\}\\times k\\times c\_\{\\text\{atompair\}\}\}at attention\-computation cost𝒪​\(L​k/P\)\\mathcal\{O\}\(L\\,k/P\)\. At diffusion\-step boundaries, rank 0 broadcasts the noised coordinates, sampled Gaussian noise, and denoised prediction so that all ranks share a single stochastic trajectory\. Importantly, per\-GPU pair\-track memory is𝒪​\(I2/P\)\\mathcal\{O\}\(I^\{2\}/P\)for𝐙\\mathbf\{Z\}and𝒪​\(L2/P\)\\mathcal\{O\}\(L^\{2\}/P\)for𝐏\\mathbf\{P\}\. A more detailed description of this process is provided in the Appendix[B](https://arxiv.org/html/2607.05439#A2)\.

### 3\.32D grid context parallelism

Our second scheme adopts the Fold\-CP framework\(Linet al\.,[2026](https://arxiv.org/html/2607.05439#bib.bib82)\)and specialises it to RFD3 \([Figure1](https://arxiv.org/html/2607.05439#S2.F1)a\)\. We arrangePPGPUs on aP×P\\sqrt\{P\}\\times\\sqrt\{P\}grid and tile the pair tensor into local quadrants𝐙\(r,c\)∈ℝIr×Ic×cz\\mathbf\{Z\}^\{\(r,c\)\}\\in\\mathbb\{R\}^\{I\_\{r\}\\times I\_\{c\}\\times c\_\{z\}\}at grid position\(r,c\)\(r,c\), withIr=I/PI\_\{r\}=I/\\sqrt\{P\}\. Queries are sharded along the row axis and replicated along the column axis; keys and values are sharded along the column axis and circulated byP\\sqrt\{P\}ring shifts, with an initial\(r,c\)↔\(c,r\)\(r,c\)\\leftrightarrow\(c,r\)transpose to align them with the query rows\. At each ring step, a GPU computes attention between its resident𝐐\\mathbf\{Q\}rows and the currently visited K/V shard, using the local bias quadrant, and merges the partial output into a running total via an online softmax\. We refer the reader to Fold\-CP’s §2–3 and Figure 2 for the full ring\-attention derivation, the Cannon\-style shift patterns for triangular updates, and the per\-module complexity table\. All tensors are represented as PyTorch DTensors; parameters are replicated across the grid and verified via a runtime check that every trainable tensor is a DTensor\.

Asymptotically, per\-device pair\-track memory is𝒪​\(I2/P\)\\mathcal\{O\}\(I^\{2\}/P\)\(the same as the 1D scheme\) while K/V are ring\-rotated rather than replicated, so each device only ever holds an𝒪​\(I/P\)\\mathcal\{O\}\(I/\\sqrt\{P\}\)slab of K/V at a time\. TheP\\sqrt\{P\}ring shifts per attention block are overlapped with local computation\.

Relative to Fold\-CP, the RFD3 adaptation mainly concerns the sampling loop and the atom\-level path: the ring primitive is invoked inside every recycling iteration of every denoising step, with rank\-0broadcasts at step boundaries that mirror the 1D scheme, and the sparse atom attention requires a distributed kNN that avoids materialising any\[L,L\]\[L,L\]distance tensor\. Concrete descriptions of these adaptations \(the distributed kNN, the boundary communicators, and the DTensor parameter distribution\) are deferred to Appendix[C](https://arxiv.org/html/2607.05439#A3)\. Importantly, this 2D grid scheme applies only when the available GPU countPPis a perfect square, so that devices can be arranged as aP×P\\sqrt\{P\}\\times\\sqrt\{P\}mesh\.

## 4Results

### 4\.1Design quality and symmetry

![Refer to caption](https://arxiv.org/html/2607.05439v1/Figure_2.png)Figure 2:Designing icosahedral nanoparticles with Design\-CP\.a, Depiction of the ASU \(red\) and its eight nearest neighbours in the icosahedral assembly \(cyan\), shown schematically \(left\) and annotated on a naturally occurring icosahedral protein nanoparticle: Lumazine Synthase \(right, PDB:[1NQX](https://www.rcsb.org/structure/1NQX)\), which has been used as a scaffold for vaccine development\.b–e, Per\-chain comparison of Design\-CP\-generated icosahedral designs \(blue,n=40n=40,6060chains×210\\times 210residues\) against original single\-GPU RFD3 monomers of length210210\(green,n=40n=40\):baverage chain breaks,caverage backbone clashes,dnon\-loop fraction,emax CA deviation; the dotted line gives the corresponding value for Lumazine Synthase \(1NQX\)\.f, Visualisation of a selected icosahedral design \(left\) with a zoomed view of the interaction between two ASUs and representative inter\-subunit Cα\\alpha–Cα\\alphadistances \(right\)\.g–l, Symmetry\-aware interface metrics for the same4040Design\-CP\-generated icosahedral designs, with the 1NQX reference overlaid:gASU clashes,hmean contacts per interface,iminimum inter\-chain distance,lnumber of interfaces with contacts\. Full metric definitions in Appendix[D](https://arxiv.org/html/2607.05439#A4)\.Context parallelism can be applied as an*inference\-only*modification: it changes how intermediate activations are partitioned and communicated, but does not alter the learned parameters of RFD3\. This makes it immediately applicable to pretrained checkpoints, but also means that when CP is used to exceed RFD3’s native crop limits, the model is sampled outside the regime it was trained on \(384 tokens and 5000 atoms, according to §1\.6 of the supplementary material ofButcheret al\.\([2025](https://arxiv.org/html/2607.05439#bib.bib143)\)\)\. In practice, we find that this distribution shift can degrade sample quality when the target displays limited symmetry\.

[Figure1](https://arxiv.org/html/2607.05439#S2.F1)b qualitatively illustrates this effect: when the number of atoms/tokens modelled is much higher than RFD3 saw during training \(in this example, designing a single monomer with 10800 amino acids and 2D sharding\), we observe visibly unnatural designs\. By contrast, when introducing constraints through symmetry, the resulting assemblies appear visibly more protein\-like, even while keeping the system size constant\. For icosahedral targets, this observation is supported quantitatively in[Section4\.2](https://arxiv.org/html/2607.05439#S4.SS2)\. We also assess that this apparent increase in sample quality is not an effect due to chain length by designing and visually checking a series of asymmetric designs \(Appendix[AppendixG](https://arxiv.org/html/2607.05439#A7)\) having the same number of chains \(60\) and same amino acids per chain \(180\) as the icosahedron shown in[Figure1](https://arxiv.org/html/2607.05439#S2.F1)b\.

We attribute this effect to the symmetry sampling mechanism described in[Section3\.1](https://arxiv.org/html/2607.05439#S3.SS1): at each symmetrised denoising step the network performs a full forward pass over all atoms of the assembly, but only the asymmetric subunit \(ASU\) of its prediction is retained, and the remaining subunits are overwritten by deterministic group\-operation copies of that ASU\. The sampler therefore varies the coordinates of a single ASU rather than those of the full assembly, which shrinks the design space and couples distant regions of the assembly by forcing every subunit to share the same ASU backbone\.

This motivates the usability of CP for the design of highly symmetric protein assemblies such as octahedral and icosahedral capsids and cages\. CP allows the model to consider \(and remain self\-consistent with respect to\) the full set of residue/atom interactions in the assembly, while the symmetry prior ensures that the number of*free*coordinates the sampler must generate is that of a single ASU\. We next focus our studies on icosahedral and octahedral protein nanoparticle design\.

### 4\.2Designing icosahedral nanoparticles

A standard*de novo*design pipeline often couples a structure generator \(RFDiffusion/RFD3\), a sequence designer \(e\.g\., ProteinMPNN\), and an external structure predictor for refolding\-based validation\. For very large multimeric assemblies, however, prediction\-based oracles become both more expensive and less reliable: AlphaFold\-Multimer accuracy, for example, has been found to degrade with chain count\(Bryantet al\.,[2022](https://arxiv.org/html/2607.05439#bib.bib81)\)\. We therefore assess design quality primarily through*in silico*structural sanity checks computed directly from the generated all\-atom coordinates, and we organise the assessment into two questions: \(i\) does sampling at this scale via Design\-CP preserve the per\-chain backbone quality of the original RFD3 model, and \(ii\) does the symmetrised assembly realise icosahedral interfaces consistent with a functional natural baseline\.

We generatedn=40n=40icosahedral nanoparticles \(210210residues per chain,6060chains,12,60012\{,\}600residues per assembly\) with the 2D sharding scheme on a2×22\\times 2grid of HG200 GPUs \(9595GB each\) at batch size one\. As a paired control for question \(i\), we additionally sampledn=40n=40single\-chain210210\-residue monomers with the standard single\-GPU RFD3\. Symmetry\-aware metrics for question \(ii\) are computed between the ASU and its eight nearest icosahedral neighbours \([Figure2](https://arxiv.org/html/2607.05439#S4.F2)a\)\. The corresponding Lumazine Synthase \(PDB:[1NQX](https://www.rcsb.org/structure/1NQX)\) values are overlaid as a dotted reference line in every panel\. This particular icosahedral nanoparticle was selected as an example of a naturally occurring nanoparticle that has previously been employed as a vaccine scaffold\(Ladenstein and Morgunova,[2020](https://arxiv.org/html/2607.05439#bib.bib57); Josephet al\.,[2025](https://arxiv.org/html/2607.05439#bib.bib58)\)\. Full metric definitions and additional supporting plots are reported in Appendix[D](https://arxiv.org/html/2607.05439#A4)and[E](https://arxiv.org/html/2607.05439#A5)\.

Table 1:Scaling of Design\-CP under icosahedral symmetry\.For each GPU countPP, we report the maximum ASU length before out\-of\-memory and the wall\-clock inference time*at that maximum ASU length*, for the 1D and 2D sharding schemes\.*Capacity ratio*is defined asLmax2​D/Lmax1​DL\_\{\\max\}^\{\\mathrm\{2D\}\}/L\_\{\\max\}^\{\\mathrm\{1D\}\}and*speedup*ast1​D/t2​Dt\_\{\\mathrm\{1D\}\}/t\_\{\\mathrm\{2D\}\}, so that values greater than unity indicate that the 2D scheme outperforms the 1D scheme\. All measurements use HG200 GPUs \(95 GB each\); a dash denotes a configuration that cannot be measured\.Max ASU length before OOM\(residues\)Wall\-clock time at max ASU\(s\)PP1D2DCapacity ratio1D2DSpeedup158581\.00×1\.00\\times102310251\.00×1\.00\\times2102––2340––3120––2943––41411370\.97×0\.97\\times432122301\.94×1\.94\\times5160––3845––6173––3662––7186––4802––8196––4689––92012011\.00×1\.00\\times724732532\.23×2\.23\\times162192371\.08×1\.08\\times703037851\.86×1\.86\\times#### Per\-chain backbone quality broadly matches the original RFD3\.

On the backbone\-sanity indicators \([Figure2](https://arxiv.org/html/2607.05439#S4.F2)b–d\), the two populations have broadly similar distributions: the average number of chain breaks per chain and the average count of backbone heavy\-atom clashes per chain both have medians at zero\. The ASUs extracted by the nanoparticles designed with Design\-CP show, however, more outliers than the designed monomers, especially for the number of backbone clashes\. The fraction of residues assigned to a helix orβ\\beta\-strand secondary\-structure element is closely matched between them and to the 1NQX reference \(medians of≈0\.6\\approx 0\.6for icosahedral nanoparticles’ ASUs vs\.≈0\.65\\approx 0\.65for monomers, against≈0\.68\\approx 0\.68for 1NQX\)\. Another notable gap is in the maximum per\-chain deviation of consecutive Cα\\alpha–Cα\\alphabond lengths from the canonical3\.83\.8Å value \([Figure2](https://arxiv.org/html/2607.05439#S4.F2)e\): monomers have a distribution concentrated around≈0\.15\\approx 0\.15Å, whereas icosahedral nanoparticles’ ASUs chains show a broader bulk around≈0\.30\\approx 0\.30Å\. We partially attribute the difference in some of the analysed metrics between Design\-CP ASUs and single\-GPU RFD3 monomers to the geometric strain of having to remain self\-consistent with symmetry mates and their inter\-subunit interfaces inside a12,60012\{,\}600\-residue joint context, rather than to a degradation introduced by the context\-parallel sampler itself\. The icosahedral design problem is strictly harder per chain than designing monomers of the same length\. The full eight\-panel comparison and a more detailed discussion are in Appendix[E](https://arxiv.org/html/2607.05439#A5)\.

#### Symmetric interfaces track a functional natural assembly\.

Symmetry\-aware metrics \([Figure2](https://arxiv.org/html/2607.05439#S4.F2)g–l\) generally follow the 1NQX baseline, but not every sample produces a sterically clean assembly\. The number of inter\-subunit clashes within the ASU’s neighbourhood \([Figure2](https://arxiv.org/html/2607.05439#S4.F2)g\) is non\-zero for a substantial fraction of designs, despite the mean being around zero\. Concordantly, the minimum Cα\\alpha–Cα\\alphadistance between distinct chains \([Figure2](https://arxiv.org/html/2607.05439#S4.F2)i\) is centred near the 1NQX reference at≈4\\approx 4Å, but the lower tail of the distribution descends toward and below the3\.53\.5Å steric\-overlap threshold defined in Appendix[D](https://arxiv.org/html/2607.05439#A4), indicating that a fraction of designs realise inter\-chain contacts inside the steric\-overlap regime\. The average number of Cα\\alpha–Cα\\alphacontacts per inter\-subunit interface centres around≈125\\approx 125\([Figure2](https://arxiv.org/html/2607.05439#S4.F2)h\), slightly above 1NQX rather than collapsing toward zero, and the average number of contacts per interface resembles that of 1NQX \([Figure2](https://arxiv.org/html/2607.05439#S4.F2)l\)\. Combined with the local visual evidence from one selected sample in[Figure2](https://arxiv.org/html/2607.05439#S4.F2)f, where ASU–ASU contacts settle into well\-formed packing geometries with Cα\\alpha–Cα\\alphadistances in the expected≈4\\approx 4–1010Å band, this indicates that the symmetrised denoising trajectory can recover physically reasonable inter\-subunit interfaces rather than independently designing non\-interacting chains\.

Together, these results show that Design\-CP scales RFD3 sampling to large icosahedral assemblies without retraining or fine\-tuning, broadly preserving per\-chain backbone quality on par with vanilla RFD3 monomers and producing symmetry\-consistent interfaces comparable to a functional natural icosahedral particle\.

![Refer to caption](https://arxiv.org/html/2607.05439v1/Octahedral_figure_main_text.png)Figure 3:Designing octahedral nanoparticles on workstation\-grade GPUs\.a, A representative Design\-CP octahedral assembly with176176residues per chain \(2424chains,4,2244\{,\}224residues total\) generated on1616NVIDIA RTX A4000 GPUs \(1616GB per device\)\.b–e, Backbone\-sanity metrics overn=12n=12such designs:bchain breaks,cbackbone clashes,dnon\-loop fraction,emax CA deviation\.f–g, Symmetry\-interface metrics:fminimum inter\-chain distance,gmean contacts per interface\. Metric definitions in Appendix[D](https://arxiv.org/html/2607.05439#A4); the remaining secondary\-structure and contact metrics are reported in Appendix[F](https://arxiv.org/html/2607.05439#A6)\.

### 4\.3Scaling: max ASU size and inference time vs\. number of GPUs

To validate the extent of applicability of our proposed CP techniques, we performed a scaling analysis on the task of designing proteins with icosahedral point\-group symmetry\. Specifically, we target the icosahedral symmetry and sweep the amino\-acid length of the asymmetric subunit \(ASU\) until inference runs out of memory \(OOM\) for a fixed number of GPUsPP\.[Table1](https://arxiv.org/html/2607.05439#S4.T1)shows these values and the inference time for that specific run right before reaching OOM error\. Unless explicitly mentioned, all experiments in this subsection use HG200 GPUs with 95 GB memory and a batch size of one\.

#### Maximum ASU size\.

With either sharding scheme, the dominant quadratic pair tracks are distributed across the device mesh, so the memory ceiling is set primarily by the*aggregate*available memory rather than by a single\-device limit\. We believe this is the reason why we observe this value of max ASU length before OOM to be similar, independently of the sharding scheme used\. Empirically, the maximum feasible ASU length increases withPPwith the expected square\-root trend predicted in[Sections3\.2](https://arxiv.org/html/2607.05439#S3.SS2)and[3\.3](https://arxiv.org/html/2607.05439#S3.SS3): asPPgrows linearly while all the GPUs are maximally used, each device stores a quadratically smaller proportion of the growing pair tensors\.

#### Inference time\.

While the two schemes reach comparable memory ceilings at a givenPP, their wall\-clock scaling differs\. The 1D scheme performs a singleAllGatherof the single\-track activations after each attention block \([Section3\.2](https://arxiv.org/html/2607.05439#S3.SS2)\), which introduces a latency cost that grows with cluster size and impacts runtime\. The 2D scheme replaces global gathers withP\\sqrt\{P\}\-step ring communication over smaller shards and overlaps these shifts with local computation \([Section3\.3](https://arxiv.org/html/2607.05439#S3.SS3)\), yielding better time scaling in practice\. This effect is reflected in[Table1](https://arxiv.org/html/2607.05439#S4.T1), where 2D CP achieves lower per\-step inference time at the samePP, often finishing inference in around half the time with respect to the 1D inference\.

### 4\.4Designing octahedral nanoparticles on small GPUs

Octahedral protein cages are a promising but underexplored therapeutic scaffold class: their rigid, multivalent geometry is suited to higher\-order receptor engagement, and their interior volume can encapsulate biologic cargoes\. Despite this potential, only a few*de novo*octahedral cages have been characterised in therapy\-relevant cell\-based functional assays\(Divineet al\.,[2021](https://arxiv.org/html/2607.05439#bib.bib50); Yanget al\.,[2024](https://arxiv.org/html/2607.05439#bib.bib49)\), and as noted in[Section1](https://arxiv.org/html/2607.05439#S1), the memory cost of jointly sampling such large symmetric assemblies has been a practical barrier to broadening this design space\.

To test whether Design\-CP lifts this barrier on commodity hardware, we repeated the scaling protocol of[Section4\.3](https://arxiv.org/html/2607.05439#S4.SS3)for octahedral targets, this time using a small cluster of workstation\-grade NVIDIA RTX A4000 GPUs \(1616GB per device,∼6×\{\\sim\}6\\timesless per\-device memory than the HG200 GPUs used so far\) and we report the results in[AppendixF](https://arxiv.org/html/2607.05439#A6)\. In practice, the 2D sharding scheme allowed us to sample full octahedral assemblies up to178178residues per chain \(with2424chains, this means modelling4,2724\{,\}272residues total\) on1616A4000 GPUs \([Figure3](https://arxiv.org/html/2607.05439#S4.F3)a\), a per\-chain size comparable to existing*de novo*octahedral nanoparticles\(Kinget al\.,[2012](https://arxiv.org/html/2607.05439#bib.bib37)\); we then evaluatedn=12n=12such designs against the same families of*in silico*metrics used for the icosahedral targets in[Section4\.2](https://arxiv.org/html/2607.05439#S4.SS2)\. Additional metrics and a discussion of the scaling sweep on this hardware are deferred to Appendix[F](https://arxiv.org/html/2607.05439#A6)\.

#### Backbone sanity tracks the icosahedral baseline\.

The backbone quality indicators \([Figure3](https://arxiv.org/html/2607.05439#S4.F3)b–e\) are healthy and consistent with the icosahedral designs population: the per\-chain count of chain breaks concentrates at zero with a single outlier at22, the count of backbone heavy\-atom clashes is essentially zero across all designs \(one outlier near7070\), the fraction of residues assigned to a helix or strand element sits at a median of≈0\.65\\approx 0\.65, and the maximum per\-chain deviation of consecutive Cα\\alpha–Cα\\alphabond lengths from the canonical3\.83\.8Å value clusters tightly around≈0\.3\\approx 0\.3Å, well below the0\.750\.75Å chain\-break threshold\. This last distribution is in fact noticeably tighter than for the icosahedral chains, and might suggest that the lower oligomeric state \(2424vs\.6060subunits\) imposes less geometric strain per chain\.

#### Symmetric interfaces are well\-formed\.

The two main symmetry\-interface metrics shown in[Figure3](https://arxiv.org/html/2607.05439#S4.F3)f–g are likewise healthy\. The minimum Cα\\alpha–Cα\\alphadistance between distinct chains sits in the≈4\\approx 4–1010Å band, consistent with proper inter\-subunit packing without steric overlap, and the average number of Cα\\alpha–Cα\\alphacontacts per inter\-subunit interface centres around≈30\\approx 30with a positive tail beyond6060\. An example of octahedral design is provided in[Figure3](https://arxiv.org/html/2607.05439#S4.F3)a, where the chains organise into a closed, octahedrally symmetric cage\. This confirms that the symmetrised trajectory realises the intended point\-group geometry with favourable in\-silico metrics on commodity GPUs\.

Overall, these results indicate that large\-assembly protein design can be made feasible on modest, widely available GPU hardware, lowering the barrier to entry for groups without access to large\-memory accelerators\.

## 5Conclusion

We introduced*Design\-CP*, two context\-parallel inference strategies for RFDiffusion 3 that shard the quadratic token\- and atom\-pair representations across multiple GPUs while preserving the model’s architecture and weights\. We characterised the memory ceilings and strong\-scaling behaviour of a lightweight 1D row\-sharding scheme and a 2D grid scheme \(which requires a perfect\-square GPU countPPto form aP×P\\sqrt\{P\}\\times\\sqrt\{P\}mesh\), and found that 2D sharding achieves better wall\-clock scaling in practice by overlapping ring communication with computation\.

Beyond RFD3, the underlying approach is model\-agnostic: any protein design model whose inference is dominated by self\-attention and other𝒪​\(n2\)\\mathcal\{O\}\(n^\{2\}\)pairwise activations can, in principle, adopt the same sharding patterns\. This suggests that context parallelism is broadly applicable across modern design pipelines, which largely build on Transformer\-style attention mechanisms\.

We further showed that strong point\-group symmetry constraints can make CP usable*out of the box*for end\-to\-end all\-atom design of large protein nanoparticles without retraining or fine\-tuning\. Prior work on context\-parallelism for structure prediction \(e\.g\., Fold\-CP\(Linet al\.,[2026](https://arxiv.org/html/2607.05439#bib.bib82)\)\) indicates that scaling inference to very large complexes can be limited by out\-of\-distribution generalisation once inputs exceed the model’s native training regime\. Our results refine this picture: for highly symmetric assemblies, symmetrised sampling collapses the effective degrees of freedom to a single ASU while still modelling the correct inter\-subunit geometry, enabling CP to scale to full icosahedral and octahedral cages while preserving promising*in silico*backbone\-sanity and symmetry\-interface metrics \([Sections4\.2](https://arxiv.org/html/2607.05439#S4.SS2)and[4\.4](https://arxiv.org/html/2607.05439#S4.SS4)\)\.

This changes the design paradigm: whereas existing nanoparticle pipelines often rely on rigid\-body docking of pre\-computed oligomeric building blocks, Design\-CP enables joint denoising of all atoms of all asymmetric subunits and their inter\-subunit interactions within a single trajectory\. We also demonstrated that the same approach enables*de novo*design of octahedral nanoparticles on a small cluster of workstation\-grade GPUs, indicating a practical path towards democratising large\-assembly protein design\.

A key limitation is that pushing beyond the native training crop can still introduce distribution shift and degrade outputs for highly asymmetric targets\. Future work should therefore explore training or fine\-tuning design models with longer contexts and/or explicit context\-parallelism in the training loop, as well as experimental validation of the designed assemblies to assess how in\-silico quality metrics translate to expression, stability, and correct self\-assembly*in vitro*\.

## Acknowledgements

L\.T\. acknowledges support from Ellision Institute of Technology, Oxford Ltd\. The authors are grateful to Prof\. David Baker and Prof\. Frank DiMaio for valuable discussions, and to the Institute of Protein Design at the University of Washington for access to computational resources\. The authors acknowledge the use of resources provided by the Isambard\-AI National AI Research Resource \(AIRR\)\. Isambard\-AI is operated by the University of Bristol and is funded by the UK Government’s Department for Science, Innovation and Technology \(DSIT\) via UK Research and Innovation; and the Science and Technology Facilities Council \[ST/AIRR/I\-A\-I/1023\]\.

## Impact Statement

This work studies inference\-time scaling for all\-atom generative protein design models\. By distributing the dominant quadratic activations across a multi\-GPU mesh, Design\-CP makes end\-to\-end design of large symmetric protein assemblies feasible on hardware ranging from data\-centre clusters to workstation\-grade GPUs\. We see the primary positive impact as democratising access to large\-assembly design: lowering the hardware barrier broadens participation beyond well\-resourced groups and could accelerate progress in therapeutic delivery systems, vaccine scaffolds, and other biomedical applications of protein nanoparticles\. The contributions are at the inference and parallelisation layer and do not introduce new model weights, training data, or predictive capabilities\. The same broadening of access does, however, raise dual\-use concerns: tools that make symmetric\-assembly design easier could, in principle, be misused for harmful protein engineering\. We view such risks as best mitigated through standard biosecurity governance, responsible disclosure norms, and careful application review at the point of use\.

## References

- J\. Abramson, J\. Adler, J\. Dunger, R\. Evans, T\. Green, A\. Pritzel, O\. Ronneberger, L\. Willmore, A\. J\. Ballard, J\. Bambrick, S\. W\. Bodenstein, D\. A\. Evans, C\. Hung, M\. O’Neill, D\. Reiman, K\. Tunyasuvunakool, Z\. Wu, A\. Žemgulytė, E\. Arvaniti, C\. Beattie, O\. Bertolli, A\. Bridgland, A\. Cherepanov, M\. Congreve, A\. I\. Cowen\-Rivers, A\. Cowie, M\. Figurnov, F\. B\. Fuchs, H\. Gladman, R\. Jain, Y\. A\. Khan, C\. M\. R\. Low, K\. Perlin, A\. Potapenko, P\. Savy, S\. Singh, A\. Stecula, A\. Thillaisundaram, C\. Tong, S\. Yakneen, E\. D\. Zhong, M\. Zielinski, A\. Žídek, V\. Bapst, P\. Kohli, M\. Jaderberg, D\. Hassabis, and J\. M\. Jumper \(2024\)Accurate structure prediction of biomolecular interactions with AlphaFold 3\.Nature630\(8016\),pp\. 493–500\(en\)\.External Links:ISSN 1476\-4687,[Link](https://www.nature.com/articles/s41586-024-07487-w),[Document](https://dx.doi.org/10.1038/s41586-024-07487-w)Cited by:[§1](https://arxiv.org/html/2607.05439#S1.p1.1),[§1](https://arxiv.org/html/2607.05439#S1.p3.8)\.
- M\. Baek, I\. Anishchenko, I\. R\. Humphreys, Q\. Cong, D\. Baker, and F\. DiMaio \(2023\)Efficient and accurate prediction of protein structure using RoseTTAFold2\.bioRxiv\(en\)\.Note:Pages: 2023\.05\.24\.542179 Section: New ResultsExternal Links:[Link](https://www.biorxiv.org/content/10.1101/2023.05.24.542179v1),[Document](https://dx.doi.org/10.1101/2023.05.24.542179)Cited by:[§1](https://arxiv.org/html/2607.05439#S1.p1.1)\.
- M\. Baek, F\. DiMaio, I\. Anishchenko, J\. Dauparas, S\. Ovchinnikov, G\. R\. Lee, J\. Wang, Q\. Cong, L\. N\. Kinch, R\. D\. Schaeffer, C\. Millán, H\. Park, C\. Adams, C\. R\. Glassman, A\. DeGiovanni, J\. H\. Pereira, A\. V\. Rodrigues, A\. A\. van Dijk, A\. C\. Ebrecht, D\. J\. Opperman, T\. Sagmeister, C\. Buhlheller, T\. Pavkov\-Keller, M\. K\. Rathinaswamy, U\. Dalwadi, C\. K\. Yip, J\. E\. Burke, K\. C\. Garcia, N\. V\. Grishin, P\. D\. Adams, R\. J\. Read, and D\. Baker \(2021\)Accurate prediction of protein structures and interactions using a three\-track neural network\.Science373\(6557\),pp\. 871–876\.External Links:[Link](https://www.science.org/doi/10.1126/science.abj8754),[Document](https://dx.doi.org/10.1126/science.abj8754)Cited by:[§1](https://arxiv.org/html/2607.05439#S1.p1.1)\.
- J\. B\. Bale, S\. Gonen, Y\. Liu, W\. Sheffler, D\. Ellis, C\. Thomas, D\. Cascio, T\. O\. Yeates, T\. Gonen, N\. P\. King, and D\. Baker \(2016\)Accurate design of megadalton\-scale two\-component icosahedral protein complexes\.Science353\(6297\),pp\. 389–394\.External Links:[Link](https://www.science.org/doi/10.1126/science.aaf8818),[Document](https://dx.doi.org/10.1126/science.aaf8818)Cited by:[§1](https://arxiv.org/html/2607.05439#S1.p2.1),[§2](https://arxiv.org/html/2607.05439#S2.SS0.SSS0.Px3.p1.2)\.
- N\. R\. Bennett, J\. L\. Watson, R\. J\. Ragotte, A\. J\. Borst, D\. L\. See, C\. Weidle, R\. Biswas, Y\. Yu, E\. L\. Shrock, R\. Ault, P\. J\. Y\. Leung, B\. Huang, I\. Goreshnik, J\. Tam, K\. D\. Carr, B\. Singer, C\. Criswell, B\. I\. M\. Wicky, D\. Vafeados, M\. Garcia Sanchez, H\. M\. Kim, S\. Vázquez Torres, S\. Chan, S\. M\. Sun, T\. T\. Spear, Y\. Sun, K\. O’Reilly, J\. M\. Maris, N\. G\. Sgourakis, R\. A\. Melnyk, C\. C\. Liu, and D\. Baker \(2026\)Atomically accurate de novo design of antibodies with RFdiffusion\.Nature649\(8095\),pp\. 183–193\(en\)\.External Links:ISSN 1476\-4687,[Link](https://www.nature.com/articles/s41586-025-09721-5),[Document](https://dx.doi.org/10.1038/s41586-025-09721-5)Cited by:[§1](https://arxiv.org/html/2607.05439#S1.p1.1)\.
- P\. Bryant, G\. Pozzati, W\. Zhu, A\. Shenoy, P\. Kundrotas, and A\. Elofsson \(2022\)Predicting the structure of large protein complexes using AlphaFold and Monte Carlo tree search\.Nature Communications13\(1\),pp\. 6028\(en\)\.External Links:ISSN 2041\-1723,[Link](https://www.nature.com/articles/s41467-022-33729-4),[Document](https://dx.doi.org/10.1038/s41467-022-33729-4)Cited by:[§4\.2](https://arxiv.org/html/2607.05439#S4.SS2.p1.1)\.
- J\. Butcher, R\. Krishna, R\. Mitra, R\. I\. Brent, Y\. Li, N\. Corley, P\. Kim, J\. Funk, S\. Mathis, S\. Salike, A\. Muraishi, H\. Eisenach, T\. R\. Thompson, J\. Chen, Y\. Politanska, E\. Sehgal, B\. Coventry, O\. Zhang, B\. Qiang, K\. Didi, M\. Kazman, F\. DiMaio, and D\. Baker \(2025\)De novo Design of All\-atom Biomolecular Interactions with RFdiffusion3\.bioRxiv\(en\)\.Note:ISSN: 2692\-8205 Pages: 2025\.09\.18\.676967 Section: New ResultsExternal Links:[Link](https://www.biorxiv.org/content/10.1101/2025.09.18.676967v1),[Document](https://dx.doi.org/10.1101/2025.09.18.676967)Cited by:[§A\.1](https://arxiv.org/html/2607.05439#A1.SS1.SSS0.Px1.p1.10),[Appendix E](https://arxiv.org/html/2607.05439#A5.SS0.SSS0.Px4.p1.6),[§1](https://arxiv.org/html/2607.05439#S1.p1.1),[§1](https://arxiv.org/html/2607.05439#S1.p3.8),[§2](https://arxiv.org/html/2607.05439#S2.SS0.SSS0.Px2.p3.1),[§4\.1](https://arxiv.org/html/2607.05439#S4.SS1.p1.1)\.
- G\. L\. Butterfield, M\. J\. Lajoie, H\. H\. Gustafson, D\. L\. Sellers, U\. Nattermann, D\. Ellis, J\. B\. Bale, S\. Ke, G\. H\. Lenz, A\. Yehdego, R\. Ravichandran, S\. H\. Pun, N\. P\. King, and D\. Baker \(2017\)Evolution of a designed protein assembly encapsulating its own RNA genome\.Nature552\(7685\),pp\. 415–420\(en\)\.External Links:ISSN 0028\-0836, 1476\-4687,[Link](https://www.nature.com/articles/nature25157),[Document](https://dx.doi.org/10.1038/nature25157)Cited by:[§1](https://arxiv.org/html/2607.05439#S1.p2.1)\.
- S\. Cheng, X\. Zhao, G\. Lu, J\. Fang, Z\. Yu, T\. Zheng, R\. Wu, X\. Zhang, J\. Peng, and Y\. You \(2023\)FastFold: Reducing AlphaFold Training Time from 11 Days to 67 Hours\.arXiv\.Note:arXiv:2203\.00854 \[cs\]External Links:[Link](http://arxiv.org/abs/2203.00854),[Document](https://dx.doi.org/10.48550/arXiv.2203.00854)Cited by:[§2](https://arxiv.org/html/2607.05439#S2.SS0.SSS0.Px2.p1.2)\.
- N\. Corley, S\. Mathis, R\. Krishna, M\. S\. Bauer, T\. R\. Thompson, W\. Ahern, M\. W\. Kazman, R\. I\. Brent, K\. Didi, A\. Kubaney, L\. McHugh, A\. Nagle, A\. Favor, M\. Kshirsagar, P\. Sturmfels, Y\. Li, J\. Butcher, B\. Qiang, L\. L\. Schaaf, R\. Mitra, K\. Campbell, O\. Zhang, R\. Weissman, I\. R\. Humphreys, Q\. Cong, J\. Funk, S\. Sonthalia, P\. Liò, D\. Baker, and F\. DiMaio \(2025\)Accelerating Biomolecular Modeling with AtomWorks and RF3\.Biochemistry\(en\)\.External Links:[Link](http://biorxiv.org/lookup/doi/10.1101/2025.08.14.670328),[Document](https://dx.doi.org/10.1101/2025.08.14.670328)Cited by:[§1](https://arxiv.org/html/2607.05439#S1.p1.1)\.
- A\. Courbet, J\. Hansen, Y\. Hsia, N\. Bethel, Y\.\-J\. Park, C\. Xu, A\. Moyer, S\. E\. Boyken, G\. Ueda, U\. Nattermann, D\. Nagarajan, D\.\-A\. Silva, W\. Sheffler, J\. Quispe, A\. Nord, N\. King, P\. Bradley, D\. Veesler, J\. Kollman, and D\. Baker \(2022\)Computational design of mechanically coupled axle\-rotor protein assemblies\.Science376\(6591\),pp\. 383–390\(en\)\.External Links:ISSN 0036\-8075, 1095\-9203,[Link](https://www.science.org/doi/10.1126/science.abm1183),[Document](https://dx.doi.org/10.1126/science.abm1183)Cited by:[§1](https://arxiv.org/html/2607.05439#S1.p2.1)\.
- T\. Dao, D\. Y\. Fu, S\. Ermon, A\. Rudra, and C\. Ré \(2022\)FlashAttention: Fast and Memory\-Efficient Exact Attention with IO\-Awareness\.arXiv\.Note:arXiv:2205\.14135 \[cs\]External Links:[Link](http://arxiv.org/abs/2205.14135),[Document](https://dx.doi.org/10.48550/arXiv.2205.14135)Cited by:[§2](https://arxiv.org/html/2607.05439#S2.SS0.SSS0.Px1.p1.2)\.
- T\. Dao \(2023\)FlashAttention\-2: Faster Attention with Better Parallelism and Work Partitioning\.arXiv\.Note:arXiv:2307\.08691 \[cs\]External Links:[Link](http://arxiv.org/abs/2307.08691),[Document](https://dx.doi.org/10.48550/arXiv.2307.08691)Cited by:[§2](https://arxiv.org/html/2607.05439#S2.SS0.SSS0.Px1.p1.2)\.
- R\. J\. De Haas, N\. Brunette, A\. Goodson, J\. Dauparas, S\. Y\. Yi, E\. C\. Yang, Q\. Dowling, H\. Nguyen, A\. Kang, A\. K\. Bera, B\. Sankaran, R\. De Vries, D\. Baker, and N\. P\. King \(2024\)Rapid and automated design of two\-component protein nanomaterials using ProteinMPNN\.Proceedings of the National Academy of Sciences121\(13\),pp\. e2314646121\(en\)\.External Links:ISSN 0027\-8424, 1091\-6490,[Link](https://pnas.org/doi/10.1073/pnas.2314646121),[Document](https://dx.doi.org/10.1073/pnas.2314646121)Cited by:[§1](https://arxiv.org/html/2607.05439#S1.p2.1),[§2](https://arxiv.org/html/2607.05439#S2.SS0.SSS0.Px3.p1.2)\.
- K\. Didi, Z\. Zhang, G\. Zhou, D\. Reidenbach, Z\. Cao, S\. Cha, T\. Geffner, C\. Dallago, J\. Tang, M\. M\. Bronstein, M\. Steinegger, E\. Kucukbenli, A\. Vahdat, and K\. Kreis \(2026\)Scaling Atomistic Protein Binder Design with Generative Pretraining and Test\-Time Compute\.arXiv\.Note:arXiv:2603\.27950 \[cs\.LG\]External Links:[Link](http://arxiv.org/abs/2603.27950),[Document](https://dx.doi.org/10.48550/arXiv.2603.27950)Cited by:[§1](https://arxiv.org/html/2607.05439#S1.p1.1)\.
- R\. Divine, H\. V\. Dang, G\. Ueda, J\. A\. Fallas, I\. Vulovic, W\. Sheffler, S\. Saini, Y\. T\. Zhao, I\. X\. Raj, P\. A\. Morawski, M\. F\. Jennewein, L\. J\. Homad, Y\. Wan, M\. R\. Tooley, F\. Seeger, A\. Etemadi, M\. L\. Fahning, J\. Lazarovits, A\. Roederer, A\. C\. Walls, L\. Stewart, M\. Mazloomi, N\. P\. King, D\. J\. Campbell, A\. T\. McGuire, L\. Stamatatos, H\. Ruohola\-Baker, J\. Mathieu, D\. Veesler, and D\. Baker \(2021\)Designed proteins assemble antibodies into modular nanocages\.Science372\(6537\),pp\. eabd9994\.External Links:[Link](https://www.science.org/doi/10.1126/science.abd9994),[Document](https://dx.doi.org/10.1126/science.abd9994)Cited by:[§4\.4](https://arxiv.org/html/2607.05439#S4.SS4.p1.1)\.
- T\. Geffner, K\. Didi, Z\. Cao, D\. Reidenbach, Z\. Zhang, C\. Dallago, E\. Kucukbenli, K\. Kreis, and A\. Vahdat \(2025a\)La\-Proteina: Atomistic Protein Generation via Partially Latent Flow Matching\.External Links:[Link](https://arxiv.org/abs/2507.09466),[Document](https://dx.doi.org/10.48550/ARXIV.2507.09466)Cited by:[§1](https://arxiv.org/html/2607.05439#S1.p1.1)\.
- T\. Geffner, K\. Didi, Z\. Zhang, D\. Reidenbach, Z\. Cao, J\. Yim, M\. Geiger, C\. Dallago, E\. Kucukbenli, A\. Vahdat, and K\. Kreis \(2025b\)Proteina: Scaling Flow\-based Protein Structure Generative Models\.arXiv\.Note:arXiv:2503\.00710 \[cs\]External Links:[Link](http://arxiv.org/abs/2503.00710),[Document](https://dx.doi.org/10.48550/arXiv.2503.00710)Cited by:[§1](https://arxiv.org/html/2607.05439#S1.p1.1)\.
- D\. S\. Goodsell and A\. J\. Olson \(2000\)Structural Symmetry and Protein Function\.Annual Review of Biophysics29\(Volume 29, 2000\),pp\. 105–153\(en\)\.External Links:ISSN 1936\-122X, 1936\-1238,[Link](https://www.annualreviews.org/content/journals/10.1146/annurev.biophys.29.1.105),[Document](https://dx.doi.org/10.1146/annurev.biophys.29.1.105)Cited by:[§1](https://arxiv.org/html/2607.05439#S1.p2.1)\.
- C\. M\. Haas, N\. Jasti, A\. Dosey, J\. D\. Allen, R\. Gillespie, J\. McGowan, E\. M\. Leaf, M\. Crispin, C\. A\. DeForest, M\. Kanekiyo, and N\. P\. King \(2025\)From sequence to scaffold: Computational design of protein nanoparticle vaccines from AlphaFold2\-predicted building blocks\.Proceedings of the National Academy of Sciences122\(45\),pp\. e2409566122\(en\)\.External Links:ISSN 0027\-8424, 1091\-6490,[Link](https://pnas.org/doi/10.1073/pnas.2409566122),[Document](https://dx.doi.org/10.1073/pnas.2409566122)Cited by:[§1](https://arxiv.org/html/2607.05439#S1.p2.1),[§2](https://arxiv.org/html/2607.05439#S2.SS0.SSS0.Px3.p1.2)\.
- C\. M\. Haas, S\. Rankovic, H\. K\. Lewis, K\. D\. Carr, C\. Weidle, S\. S\. Gerdes, L\. R\. Nuss, F\. Ruiz, S\. Moiz, M\. Fiorelli, E\. Grey, J\. McGowan, N\. Kumar, A\. Creanga, A\. Kang, H\. Nguyen, Y\. Wang, B\. Sankaran, A\. Dosey, R\. Ravichandran, A\. K\. Bera, E\. M\. Leaf, C\. A\. DeForest, M\. Kanekiyo, A\. J\. Borst, and N\. P\. King \(2026\)De novo design of protein nanoparticles with integrated functional motifs\.bioRxiv,pp\. 2025\.12\.19\.695620\.External Links:ISSN 2692\-8205,[Link](https://pmc.ncbi.nlm.nih.gov/articles/PMC12776285/),[Document](https://dx.doi.org/10.64898/2025.12.19.695620)Cited by:[§1](https://arxiv.org/html/2607.05439#S1.p2.1),[§2](https://arxiv.org/html/2607.05439#S2.SS0.SSS0.Px3.p1.2)\.
- T\. Hayes, R\. Rao, H\. Akin, N\. J\. Sofroniew, D\. Oktay, Z\. Lin, R\. Verkuil, V\. Q\. Tran, J\. Deaton, M\. Wiggert, R\. Badkundri, I\. Shafkat, J\. Gong, A\. Derry, R\. S\. Molina, N\. Thomas, Y\. A\. Khan, C\. Mishra, C\. Kim, L\. J\. Bartie, M\. Nemeth, P\. D\. Hsu, T\. Sercu, S\. Candido, and A\. Rives \(2025\)Simulating 500 million years of evolution with a language model\.Science\(EN\)\.External Links:[Link](https://www.science.org/doi/10.1126/science.ads0018),[Document](https://dx.doi.org/10.1126/science.ads0018)Cited by:[§2](https://arxiv.org/html/2607.05439#S2.SS0.SSS0.Px1.p1.2)\.
- Y\. Hsia, J\. B\. Bale, S\. Gonen, D\. Shi, W\. Sheffler, K\. K\. Fong, U\. Nattermann, C\. Xu, P\. Huang, R\. Ravichandran, S\. Yi, T\. N\. Davis, T\. Gonen, N\. P\. King, and D\. Baker \(2016\)Design of a hyperstable 60\-subunit protein icosahedron\.Nature535\(7610\),pp\. 136–139\(en\)\.External Links:ISSN 0028\-0836, 1476\-4687,[Link](https://www.nature.com/articles/nature18010),[Document](https://dx.doi.org/10.1038/nature18010)Cited by:[§1](https://arxiv.org/html/2607.05439#S1.p2.1),[§2](https://arxiv.org/html/2607.05439#S2.SS0.SSS0.Px3.p1.2)\.
- J\. B\. Ingraham, M\. Baranov, Z\. Costello, K\. W\. Barber, W\. Wang, A\. Ismail, V\. Frappier, D\. M\. Lord, C\. Ng\-Thow\-Hing, E\. R\. Van Vlack, S\. Tie, V\. Xue, S\. C\. Cowles, A\. Leung, J\. V\. Rodrigues, C\. L\. Morales\-Perez, A\. M\. Ayoub, R\. Green, K\. Puentes, F\. Oplinger, N\. V\. Panwar, F\. Obermeyer, A\. R\. Root, A\. L\. Beam, F\. J\. Poelwijk, and G\. Grigoryan \(2023\)Illuminating protein space with a programmable generative model\.Nature623\(7989\),pp\. 1070–1078\(en\)\.External Links:ISSN 1476\-4687,[Link](https://www.nature.com/articles/s41586-023-06728-8),[Document](https://dx.doi.org/10.1038/s41586-023-06728-8)Cited by:[§1](https://arxiv.org/html/2607.05439#S1.p1.1)\.
- J\. Joseph, K\. Modenkattil Sethumadhavan, P\. Ahlawat, M\. Prakash, G\. Kandpal, G\. Raj, H\. Srivastava, P\. Charulekha, A\. K Dev, A\. Radhakrishnan, V\. Singh, R\. Yadav, P\. Chandramohanan, R\. Varghese, Z\. A\. Rizvi, A\. Awasthi, and V\. S\. Raj \(2025\)Lumazine Synthase Nanoparticles as a Versatile Platform for Multivalent Antigen Presentation and Cross\-Protective Coronavirus Vaccines\.ACS nano19\(31\),pp\. 28295–28314\(eng\)\.External Links:ISSN 1936\-086X,[Document](https://dx.doi.org/10.1021/acsnano.5c06081)Cited by:[§4\.2](https://arxiv.org/html/2607.05439#S4.SS2.p2.8)\.
- J\. Jumper, R\. Evans, A\. Pritzel, T\. Green, M\. Figurnov, O\. Ronneberger, K\. Tunyasuvunakool, R\. Bates, A\. Žídek, A\. Potapenko, A\. Bridgland, C\. Meyer, S\. A\. A\. Kohl, A\. J\. Ballard, A\. Cowie, B\. Romera\-Paredes, S\. Nikolov, R\. Jain, J\. Adler, T\. Back, S\. Petersen, D\. Reiman, E\. Clancy, M\. Zielinski, M\. Steinegger, M\. Pacholska, T\. Berghammer, S\. Bodenstein, D\. Silver, O\. Vinyals, A\. W\. Senior, K\. Kavukcuoglu, P\. Kohli, and D\. Hassabis \(2021\)Highly accurate protein structure prediction with AlphaFold\.Nature596\(7873\),pp\. 583–589\(en\)\.External Links:ISSN 1476\-4687,[Link](https://www.nature.com/articles/s41586-021-03819-2),[Document](https://dx.doi.org/10.1038/s41586-021-03819-2)Cited by:[§1](https://arxiv.org/html/2607.05439#S1.p1.1)\.
- T\. Karras, M\. Aittala, T\. Aila, and S\. Laine \(2022\)Elucidating the Design Space of Diffusion\-Based Generative Models\.arXiv\.Note:arXiv:2206\.00364 \[cs\]External Links:[Link](http://arxiv.org/abs/2206.00364),[Document](https://dx.doi.org/10.48550/arXiv.2206.00364)Cited by:[§A\.1](https://arxiv.org/html/2607.05439#A1.SS1.SSS0.Px7.p1.2)\.
- N\. P\. King, J\. B\. Bale, W\. Sheffler, D\. E\. McNamara, S\. Gonen, T\. Gonen, T\. O\. Yeates, and D\. Baker \(2014\)Accurate design of co\-assembling multi\-component protein nanomaterials\.Nature510\(7503\),pp\. 103–108\(en\)\.External Links:ISSN 1476\-4687,[Link](https://www.nature.com/articles/nature13404),[Document](https://dx.doi.org/10.1038/nature13404)Cited by:[§2](https://arxiv.org/html/2607.05439#S2.SS0.SSS0.Px3.p1.2)\.
- N\. P\. King, W\. Sheffler, M\. R\. Sawaya, B\. S\. Vollmar, J\. P\. Sumida, I\. André, T\. Gonen, T\. O\. Yeates, and D\. Baker \(2012\)Computational Design of Self\-Assembling Protein Nanomaterials with Atomic Level Accuracy\.Science336\(6085\),pp\. 1171–1174\.External Links:[Link](https://www.science.org/doi/10.1126/science.1219364),[Document](https://dx.doi.org/10.1126/science.1219364)Cited by:[§1](https://arxiv.org/html/2607.05439#S1.p2.1),[§2](https://arxiv.org/html/2607.05439#S2.SS0.SSS0.Px3.p1.2),[§4\.4](https://arxiv.org/html/2607.05439#S4.SS4.p2.7)\.
- G\. Labesse, N\. Colloc’h, J\. Pothier, and J\.\-P\. Mornon \(1997\)P\-SEA: a new efficient assignment of secondary structure from CA trace of proteins\.Bioinformatics13\(3\),pp\. 291–295\.External Links:ISSN 1367\-4803,[Link](https://doi.org/10.1093/bioinformatics/13.3.291),[Document](https://dx.doi.org/10.1093/bioinformatics/13.3.291)Cited by:[3rd item](https://arxiv.org/html/2607.05439#A4.I1.i3.p1.1),[Appendix E](https://arxiv.org/html/2607.05439#A5.SS0.SSS0.Px3.p1.3)\.
- R\. Ladenstein and E\. Morgunova \(2020\)Second career of a biosynthetic enzyme: Lumazine synthase as a virus\-like nanoparticle in vaccine development\.Biotechnology Reports27,pp\. e00494\.External Links:ISSN 2215\-017X,[Link](https://www.sciencedirect.com/science/article/pii/S2215017X20303593),[Document](https://dx.doi.org/10.1016/j.btre.2020.e00494)Cited by:[§4\.2](https://arxiv.org/html/2607.05439#S4.SS2.p2.8)\.
- S\. Li, F\. Xue, C\. Baranwal, Y\. Li, and Y\. You \(2022\)Sequence Parallelism: Long Sequence Training from System Perspective\.arXiv\.Note:arXiv:2105\.13120 \[cs\]External Links:[Link](http://arxiv.org/abs/2105.13120),[Document](https://dx.doi.org/10.48550/arXiv.2105.13120)Cited by:[§1](https://arxiv.org/html/2607.05439#S1.p3.8)\.
- D\. Lin, S\. Chu, V\. Iyer, Y\. Lee, J\. S\. John, K\. Boyd, B\. Roland, X\. Ren, G\. Zhou, Z\. Cao, P\. Binder, Y\. Zhautouskaya, J\. Zakrzewski, M\. Stadler, K\. Gion, Y\. Peng, X\. Chen, T\. Zhang, P\. Junk, M\. Dimon, P\. Gniewek, F\. Ortega, M\. Polen, I\. Grubisic, A\. Bashir, G\. Holt, D\. Kovtun, M\. Grass, L\. Naef, R\. Wang, J\. Peng, A\. Costa, S\. Paliwal, E\. Calleja, T\. Rvachov, N\. Tadimeti, R\. Tal, and E\. Kucukbenli \(2026\)Fold\-CP: A Context Parallelism Framework for Biomolecular Modeling\.arXiv\.Note:arXiv:2603\.14806 \[q\-bio\]External Links:[Link](http://arxiv.org/abs/2603.14806),[Document](https://dx.doi.org/10.48550/arXiv.2603.14806)Cited by:[Appendix C](https://arxiv.org/html/2607.05439#A3.p1.1),[1st item](https://arxiv.org/html/2607.05439#S1.I1.i1.p1.1),[§1](https://arxiv.org/html/2607.05439#S1.p4.2),[§2](https://arxiv.org/html/2607.05439#S2.SS0.SSS0.Px2.p2.11),[§2](https://arxiv.org/html/2607.05439#S2.SS0.SSS0.Px2.p3.1),[§3\.3](https://arxiv.org/html/2607.05439#S3.SS3.p1.8),[§5](https://arxiv.org/html/2607.05439#S5.p3.1)\.
- Y\. Lin and M\. AlQuraishi \(2023\)Generating Novel, Designable, and Diverse Protein Structures by Equivariantly Diffusing Oriented Residue Clouds\.arXiv\.Note:arXiv:2301\.12485 \[q\-bio\]External Links:[Link](http://arxiv.org/abs/2301.12485),[Document](https://dx.doi.org/10.48550/arXiv.2301.12485)Cited by:[§1](https://arxiv.org/html/2607.05439#S1.p1.1)\.
- Y\. Lin, M\. Lee, Z\. Zhang, and M\. AlQuraishi \(2024\)Out of Many, One: Designing and Scaffolding Proteins at the Scale of the Structural Universe with Genie 2\.arXiv\.Note:arXiv:2405\.15489 \[q\-bio\]External Links:[Link](http://arxiv.org/abs/2405.15489),[Document](https://dx.doi.org/10.48550/arXiv.2405.15489)Cited by:[§1](https://arxiv.org/html/2607.05439#S1.p1.1)\.
- A\. Liu, A\. Elaldi, N\. T\. Franklin, N\. Russell, G\. S\. Atwal, Y\. A\. Ban, and O\. Viessmann \(2025\)Flash Invariant Point Attention\.arXiv\.Note:arXiv:2505\.11580 \[cs\]External Links:[Link](http://arxiv.org/abs/2505.11580),[Document](https://dx.doi.org/10.48550/arXiv.2505.11580)Cited by:[§2](https://arxiv.org/html/2607.05439#S2.SS0.SSS0.Px1.p1.2)\.
- H\. Liu, M\. Zaharia, and P\. Abbeel \(2023\)Ring Attention with Blockwise Transformers for Near\-Infinite Context\.arXiv\.Note:arXiv:2310\.01889 \[cs\]External Links:[Link](http://arxiv.org/abs/2310.01889),[Document](https://dx.doi.org/10.48550/arXiv.2310.01889)Cited by:[§2](https://arxiv.org/html/2607.05439#S2.SS0.SSS0.Px2.p2.11)\.
- S\. Luo, Y\. Su, X\. Peng, S\. Wang, J\. Peng, and J\. Ma \(2022\)Antigen\-Specific Antibody Design and Optimization with Diffusion\-Based Generative Models for Protein Structures\.Bioinformatics\(en\)\.External Links:[Link](http://biorxiv.org/lookup/doi/10.1101/2022.07.10.499510),[Document](https://dx.doi.org/10.1101/2022.07.10.499510)Cited by:[§1](https://arxiv.org/html/2607.05439#S1.p1.1)\.
- J\. Marcandalli, B\. Fiala, S\. Ols, M\. Perotti, W\. de van der Schueren, J\. Snijder, E\. Hodge, M\. Benhaim, R\. Ravichandran, L\. Carter, W\. Sheffler, L\. Brunner, M\. Lawrenz, P\. Dubois, A\. Lanzavecchia, F\. Sallusto, K\. K\. Lee, D\. Veesler, C\. E\. Correnti, L\. J\. Stewart, D\. Baker, K\. Loré, L\. Perez, and N\. P\. King \(2019\)Induction of Potent Neutralizing Antibody Responses by a Designed Protein Nanoparticle Vaccine for Respiratory Syncytial Virus\.Cell176\(6\),pp\. 1420–1431\.e17\.External Links:ISSN 0092\-8674,[Link](https://pmc.ncbi.nlm.nih.gov/articles/PMC6424820/),[Document](https://dx.doi.org/10.1016/j.cell.2019.01.046)Cited by:[§1](https://arxiv.org/html/2607.05439#S1.p2.1)\.
- J\. A\. Marsh and S\. A\. Teichmann \(2015\)Structure, dynamics, assembly, and evolution of protein complexes\.Annual Review of Biochemistry84,pp\. 551–575\(eng\)\.External Links:ISSN 1545\-4509,[Document](https://dx.doi.org/10.1146/annurev-biochem-060614-034142)Cited by:[§1](https://arxiv.org/html/2607.05439#S1.p2.1)\.
- M\. Milakov and N\. Gimelshein \(2018\)Online normalizer calculation for softmax\.arXiv\.Note:arXiv:1805\.02867 \[cs\]External Links:[Link](http://arxiv.org/abs/1805.02867),[Document](https://dx.doi.org/10.48550/arXiv.1805.02867)Cited by:[§2](https://arxiv.org/html/2607.05439#S2.SS0.SSS0.Px1.p1.2)\.
- S\. Passaro, G\. Corso, and J\. Wohlwend \(2025\)Boltz\-2: Towards Accurate and Efficient Binding Affinity Prediction\.\(en\)\.Cited by:[§1](https://arxiv.org/html/2607.05439#S1.p1.1),[§2](https://arxiv.org/html/2607.05439#S2.SS0.SSS0.Px2.p2.11)\.
- R\. Pope, S\. Douglas, A\. Chowdhery, J\. Devlin, J\. Bradbury, A\. Levskaya, J\. Heek, K\. Xiao, S\. Agrawal, and J\. Dean \(2022\)Efficiently Scaling Transformer Inference\.arXiv\.Note:arXiv:2211\.05102 \[cs\]External Links:[Link](http://arxiv.org/abs/2211.05102),[Document](https://dx.doi.org/10.48550/arXiv.2211.05102)Cited by:[§1](https://arxiv.org/html/2607.05439#S1.p3.8)\.
- M\. N\. Rabe and C\. Staats \(2022\)Self\-attention Does Not Need $O\(n^2\)$ Memory\.arXiv\.Note:arXiv:2112\.05682 \[cs\]External Links:[Link](http://arxiv.org/abs/2112.05682),[Document](https://dx.doi.org/10.48550/arXiv.2112.05682)Cited by:[§2](https://arxiv.org/html/2607.05439#S2.SS0.SSS0.Px1.p1.2)\.
- R\. J\. Ragotte, H\. Liang, J\. Tam, S\. Miletic, J\. M\. Berman, R\. Palou, C\. Weidle, Z\. Li, M\. Glögl, G\. L\. Beilhartz, K\. D\. Carr, A\. J\. Borst, B\. Coventry, X\. Wang, J\. L\. Rubinstein, M\. Tyers, D\. Schramek, R\. A\. Melnyk, and D\. Baker \(2025\)De novo design of potent inhibitors of clostridial family toxins\.Proceedings of the National Academy of Sciences122\(39\),pp\. e2509329122\.External Links:[Link](https://www.pnas.org/doi/10.1073/pnas.2509329122),[Document](https://dx.doi.org/10.1073/pnas.2509329122)Cited by:[§1](https://arxiv.org/html/2607.05439#S1.p1.1)\.
- N\. Shazeer \(2020\)GLU Variants Improve Transformer\.arXiv\.Note:arXiv:2002\.05202 \[cs\]External Links:[Link](http://arxiv.org/abs/2002.05202),[Document](https://dx.doi.org/10.48550/arXiv.2002.05202)Cited by:[§B\.4](https://arxiv.org/html/2607.05439#A2.SS4.SSS0.Px1.p1.6)\.
- W\. Sheffler, E\. C\. Yang, Q\. Dowling, Y\. Hsia, C\. N\. Fries, J\. Stanislaw, M\. D\. Langowski, M\. Brandys, Z\. Li, R\. Skotheim, A\. J\. Borst, A\. Khmelinskaia, N\. P\. King, and D\. Baker \(2023\)Fast and versatile sequence\-independent protein docking for nanomaterials design using RPXDock\.PLOS Computational Biology19\(5\),pp\. e1010680\(en\)\.External Links:ISSN 1553\-7358,[Link](https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1010680),[Document](https://dx.doi.org/10.1371/journal.pcbi.1010680)Cited by:[§1](https://arxiv.org/html/2607.05439#S1.p2.1)\.
- S\. Vázquez Torres, P\. J\. Y\. Leung, P\. Venkatesh, I\. D\. Lutz, F\. Hink, H\. Huynh, J\. Becker, A\. H\. Yeh, D\. Juergens, N\. R\. Bennett, A\. N\. Hoofnagle, E\. Huang, M\. J\. MacCoss, M\. Expòsit, G\. R\. Lee, A\. K\. Bera, A\. Kang, J\. De La Cruz, P\. M\. Levine, X\. Li, M\. Lamb, S\. R\. Gerben, A\. Murray, P\. Heine, E\. N\. Korkmaz, J\. Nivala, L\. Stewart, J\. L\. Watson, J\. M\. Rogers, and D\. Baker \(2024\)De novo design of high\-affinity binders of bioactive helical peptides\.Nature626\(7998\),pp\. 435–442\(en\)\.External Links:ISSN 1476\-4687,[Link](https://www.nature.com/articles/s41586-023-06953-1),[Document](https://dx.doi.org/10.1038/s41586-023-06953-1)Cited by:[§1](https://arxiv.org/html/2607.05439#S1.p1.1)\.
- A\. C\. Walls, B\. Fiala, A\. Schäfer, S\. Wrenn, M\. N\. Pham, M\. Murphy, L\. V\. Tse, L\. Shehata, M\. A\. O’Connor, C\. Chen, M\. J\. Navarro, M\. C\. Miranda, D\. Pettie, R\. Ravichandran, J\. C\. Kraft, C\. Ogohara, A\. Palser, S\. Chalk, E\. Lee, K\. Guerriero, E\. Kepl, C\. M\. Chow, C\. Sydeman, E\. A\. Hodge, B\. Brown, J\. T\. Fuller, K\. H\. Dinnon, L\. E\. Gralinski, S\. R\. Leist, K\. L\. Gully, T\. B\. Lewis, M\. Guttman, H\. Y\. Chu, K\. K\. Lee, D\. H\. Fuller, R\. S\. Baric, P\. Kellam, L\. Carter, M\. Pepper, T\. P\. Sheahan, D\. Veesler, and N\. P\. King \(2020\)Elicitation of Potent Neutralizing Antibody Responses by Designed Protein Nanoparticle Vaccines for SARS\-CoV\-2\.Cell183\(5\),pp\. 1367–1382\.e17\(en\)\.External Links:ISSN 00928674,[Link](https://linkinghub.elsevier.com/retrieve/pii/S0092867420314501),[Document](https://dx.doi.org/10.1016/j.cell.2020.10.043)Cited by:[§1](https://arxiv.org/html/2607.05439#S1.p2.1)\.
- J\. L\. Watson, D\. Juergens, N\. R\. Bennett, B\. L\. Trippe, J\. Yim, H\. E\. Eisenach, W\. Ahern, A\. J\. Borst, R\. J\. Ragotte, L\. F\. Milles, B\. I\. M\. Wicky, N\. Hanikel, S\. J\. Pellock, A\. Courbet, W\. Sheffler, J\. Wang, P\. Venkatesh, I\. Sappington, S\. V\. Torres, A\. Lauko, V\. De Bortoli, E\. Mathieu, S\. Ovchinnikov, R\. Barzilay, T\. S\. Jaakkola, F\. DiMaio, M\. Baek, and D\. Baker \(2023\)De novo design of protein structure and function with RFdiffusion\.Nature620\(7976\),pp\. 1089–1100\(en\)\.External Links:ISSN 1476\-4687,[Link](https://www.nature.com/articles/s41586-023-06415-8),[Document](https://dx.doi.org/10.1038/s41586-023-06415-8)Cited by:[§1](https://arxiv.org/html/2607.05439#S1.p1.1)\.
- J\. Wohlwend, G\. Corso, S\. Passaro, M\. Reveiz, K\. Leidal, W\. Swiderski, T\. Portnoi, I\. Chinn, J\. Silterra, T\. Jaakkola, and R\. Barzilay \(2024\)Boltz\-1 Democratizing Biomolecular Interaction Modeling\.bioRxiv\(en\)\.Note:Pages: 2024\.11\.19\.624167 Section: New ResultsExternal Links:[Link](https://www.biorxiv.org/content/10.1101/2024.11.19.624167v1),[Document](https://dx.doi.org/10.1101/2024.11.19.624167)Cited by:[§1](https://arxiv.org/html/2607.05439#S1.p1.1)\.
- H\. Wu, M\. Guo, Y\. Ma, Y\. Sun, J\. Wang, W\. Matusik, and M\. Long \(2025\)FlashBias: Fast Computation of Attention with Bias\.arXiv\.Note:arXiv:2505\.12044 \[cs\]External Links:[Link](http://arxiv.org/abs/2505.12044),[Document](https://dx.doi.org/10.48550/arXiv.2505.12044)Cited by:[§2](https://arxiv.org/html/2607.05439#S2.SS0.SSS0.Px1.p1.2)\.
- E\. C\. Yang, R\. Divine, M\. C\. Miranda, A\. J\. Borst, W\. Sheffler, J\. Z\. Zhang, J\. Decarreau, A\. Saragovi, M\. Abedi, N\. Goldbach, M\. Ahlrichs, C\. Dobbins, A\. Hand, S\. Cheng, M\. Lamb, P\. M\. Levine, S\. Chan, R\. Skotheim, J\. Fallas, G\. Ueda, J\. Lubner, M\. Somiya, A\. Khmelinskaia, N\. P\. King, and D\. Baker \(2024\)Computational design of non\-porous pH\-responsive antibody nanoparticles\.Nature Structural & Molecular Biology31\(9\),pp\. 1404–1412\(en\)\.External Links:ISSN 1545\-9985,[Link](https://www.nature.com/articles/s41594-024-01288-5),[Document](https://dx.doi.org/10.1038/s41594-024-01288-5)Cited by:[§4\.4](https://arxiv.org/html/2607.05439#S4.SS4.p1.1)\.
- F\. Zhu, A\. Nowaczynski, R\. Li, J\. Xin, Y\. Song, M\. Marcinkiewicz, S\. B\. Eryilmaz, J\. Yang, and M\. Andersch \(2024\)ScaleFold: Reducing AlphaFold Initial Training Time to 10 Hours\.arXiv\.Note:arXiv:2404\.11068 \[cs\]External Links:[Link](http://arxiv.org/abs/2404.11068),[Document](https://dx.doi.org/10.48550/arXiv.2404.11068)Cited by:[§2](https://arxiv.org/html/2607.05439#S2.SS0.SSS0.Px2.p1.2)\.

## Appendix ARFD3 background

This appendix expands on §[3\.1](https://arxiv.org/html/2607.05439#S3.SS1)by recording the architectural choices, hyper\-parameters, and inference\-loop mechanics of stock RFDiffusion 3 that Design\-CP wraps without modification\. The intent is to make the rest of the paper readable in isolation: every constant referenced in §[3](https://arxiv.org/html/2607.05439#S3)–§[4\.4](https://arxiv.org/html/2607.05439#S4.SS4)is fixed here, with values taken from the configuration files of the open\-source codebase\.

### A\.1RFDiffusion 3 architecture and denoising loop

#### Token and atom representations\.

As in the main text, RFD3\(Butcheret al\.,[2025](https://arxiv.org/html/2607.05439#bib.bib143)\)jointly maintains a token\-level single track𝐒∈ℝI×cs\\mathbf\{S\}\\in\\mathbb\{R\}^\{I\\times c\_\{s\}\}and pair track𝐙∈ℝI×I×cz\\mathbf\{Z\}\\in\\mathbb\{R\}^\{I\\times I\\times c\_\{z\}\}, together with an atom\-level single track𝐀∈ℝL×catom\\mathbf\{A\}\\in\\mathbb\{R\}^\{L\\times c\_\{\\text\{atom\}\}\}and dense pair track𝐏∈ℝL×L×catompair\\mathbf\{P\}\\in\\mathbb\{R\}^\{L\\times L\\times c\_\{\\text\{atompair\}\}\}\. The default channel widths in the open\-source configuration arecs=384c\_\{s\}=384,cz=128c\_\{z\}=128,catom=128c\_\{\\text\{atom\}\}=128, andcatompair=16c\_\{\\text\{atompair\}\}=16\. An additional internal channelctoken=768c\_\{\\text\{token\}\}=768is used inside the diffusion path for the upcast/downcast cross\-attention modules, and a Fourier time embedding of dimensionct,embed=256c\_\{t,\\text\{embed\}\}=256encodes the current noise level\.

#### Pairformer blocks\.

The five\-stage architecture announced in §[3\.1](https://arxiv.org/html/2607.05439#S3.SS1)is built from two recurring micro\-architectures\. The Pairformer block used inside the token initialiser and the diffusion token encoder is, in this codebase, an*AttentionPairBias*\(1616heads, optional QK\-norm\) followed by a Transition module; both stock configurations explicitly disable the AlphaFold\-3\-style triangle multiplication and triangle attention paths \(use\_triangle\_attn=false,use\_triangle\_mult=false\)\. The atom\-attention block used in the atom encoder, the atom decoder, and inside the token initialiser’s atom embedder is an*AttentionPairBiasDiffusion*attention layer with44heads followed by a transition; the open\-source default applies0\.100\.10dropout inside the diffusion\-side atom blocks\. All blocks share a common*conditional*layer\-norm path that injects the time embedding and, where applicable, the conditioning state\.

#### Five\-stages architecutre\.

With those primitives, the stock open\-source configuration realises the following pipeline:

- •Token initialiser\(one call per trajectory\)\. A token\-level single\-track is built from the discrete features \(residue type, motif tokens, predicted\-LDDT, “non\-loopy” flag\), the pair track𝐙\\mathbf\{Z\}is initialised from the outer sum of two independent linear projections of𝐒\\mathbf\{S\}plus a relative\-position\-encoding bias, and an atom embedder pre\-aggregates𝐀\\mathbf\{A\}into a starting per\-token feature\. Two*non\-triangular*Pairformer blocks \(each1616\-head AttentionPairBias \+ Transition\) then refine\(𝐒,𝐙\)\(\\mathbf\{S\},\\mathbf\{Z\}\), and the dense atom pair tensor𝐏∈ℝL×L×catompair\\mathbf\{P\}\\in\\mathbb\{R\}^\{L\\times L\\times c\_\{\\text\{atompair\}\}\}is materialised at this stage\.
- •Atom encoder\(one call per trajectory\)\. Three blocks of sparse atom attention \(each a44\-head AttentionPairBiasDiffusion \+ Transition\) refine the atom\-level latent𝐀\\mathbf\{A\}using the kNN budget described below\.
- •Diffusion token encoder\(one call per recycling iteration\)\. The pair track is augmented along the channel dimension with a noised\-coordinate distogram𝐃dist\\mathbf\{D\}\_\{\\text\{dist\}\}and the self\-conditioning distogram𝐃self∈ℝB×I×I×nbins\\mathbf\{D\}\_\{\\text\{self\}\}\\in\\mathbb\{R\}^\{B\\times I\\times I\\times n\_\{\\text\{bins\}\}\}\(withnbins=65n\_\{\\text\{bins\}\}=65\), then passed through two further non\-triangular Pairformer blocks\. The single track𝐒\\mathbf\{S\}receives an AdaLN injection of the Fourier time embedding before each block\.
- •Diffusion transformer\(one call per recycling iteration\)\. Eighteen sequential AttentionPairBiasDiffusion \+ ConditionedTransitionBlock blocks \(1616heads each,0\.100\.10dropout\) update the token\-level latent𝐀I∈ℝB×I×ctoken\\mathbf\{A\}\_\{I\}\\in\\mathbb\{R\}^\{B\\times I\\times c\_\{\\text\{token\}\}\}using the local pair bias derived from𝐙\\mathbf\{Z\}\. This is the stage that dominates per\-step compute and that motivates the per\-blockAllGatherin the 1D parallel scheme\.
- •Atom decoder\(one call per recycling iteration\)\. Three blocks each apply a token\-to\-atom upcast \(cross\-attention from atoms onto the current token state, with annsplit=3n\_\{\\text\{split\}\}=3chunking of the upcast linear\), a44\-head atom\-level self\-attention block with0\.100\.10dropout, and a token\-level downcast that scatter\-means atom features back to tokens\. The final block emits a per\-atom coordinate update which is unwound through the EDM update of §[A\.1](https://arxiv.org/html/2607.05439#A1.SS1.SSS0.Px7)\.

Within each denoising step, stages 1–2 fire once, while stages 3–5 are repeated fornrecycle=2n\_\{\\text\{recycle\}\}=2recycling iterations: the first iteration uses a zero\-initialised𝐃self\\mathbf\{D\}\_\{\\text\{self\}\}, and each subsequent iteration feeds back the distogram computed by bucketising the previous iteration’s predicted Cα\\alphacoordinates\.

#### Token\- and atom\-level features at the boundary\.

The token initialiser reads a per\-token feature dictionary that, in the open\-source configuration, sums tocs,inputs=37c\_\{s,\\text\{inputs\}\}=37channels before projection: a3232\-dim residue\-type embedding, a33\-dim motif\-token\-type one\-hot, a scalar reference plDDT, and a scalar non\-loopy flag\. The atom embedder consumes a much wider402402\-dim per\-atom feature:256256\-dim character\-level atom\-name encodings, a128128\-dim element embedding,33\-dim reference coordinates, and a battery of scalar flags \(formal charge, occupancy mask, motif\-atom\-with\-fixed\-coord, motif\-atom\-unindexed, has\-zero\-occupancy\) plus the conditioning channels \(per\-atom RASA, hydrogen\-bond donor/acceptor activity, atom\-level hotspot\)\. The relative\-position\-encoding bias added to𝐙init\\mathbf\{Z\}\_\{\\text\{init\}\}is built from one\-hot encodings of residue offset \(range±32\\pm 32\), token offset \(same range\), chain\-hop separation \(range±2\\pm 2\), and a same\-entity boolean\.

#### Atom\-level kNN attention budget\.

The𝒪​\(L​k\)\\mathcal\{O\}\(L\\,k\)atom\-level attention announced in §[3\.1](https://arxiv.org/html/2607.05439#S3.SS1)draws itskkneighbours from a structured budget: a fixed numbernseqn\_\{\\text\{seq\}\}of per\-residue sequence\-local neighbours \(defaultnseq=2n\_\{\\text\{seq\}\}=2, namely the query atom plus its immediate flanking residues’ atoms\), and the spatially closest atoms beyond those, until the per\-query budgetk=nattn\-keysk=n\_\{\\text\{attn\-keys\}\}\(defaultk=128k=128\) is filled\. The kNN indices are computed once per denoising step from the current noised coordinates𝐗\(t\)\\mathbf\{X\}^\{\(t\)\}via a single𝐜𝐝𝐢𝐬𝐭\\mathbf\{cdist\}call and reused across all recycling iterations of that step\. When the input is a multi\-chain assembly with more than three chains, the budget is split into an intra\-chain quota ofk−max⁡\(32,k/4\)k\-\\max\(32,k/4\)neighbours and an inter\-chain quota of at least3232neighbours, ensuring that every query atom always retains at least3232context atoms outside its own chain; this is the mechanism by which the dense\[L,L\]\[L,L\]distogram gets replaced by a sparse\[L,k\]\[L,k\]slice without losing inter\-chain interactions\.

#### Chunked pairwise embedder\.

The dense\[L,L,catompair\]\[L,L,c\_\{\\text\{atompair\}\}\]allocation that the standard token initialiser performs becomes prohibitive on the assemblies we target\. The optional low\-memory mode that the main text mentions in passing replaces this dense allocation by a coordinate\-dependent on\-the\-fly construction: every linear projection of𝐏\\mathbf\{P\}that depends only on per\-atom features \(reference coordinates, element, charge, residue\-bond graph\) is precomputed and*cached*once at tokenisation, while the coordinate\-dependent components – the inverse\-distance and same\-residue masks that vary with𝐗\(t\)\\mathbf\{X\}^\{\(t\)\}– are recomputed at the kNN indices at the start of every atom\-attention block\. The block therefore consumes a\[L,k,catompair\]\[L,k,c\_\{\\text\{atompair\}\}\]slice rather than a\[L,L,catompair\]\[L,L,c\_\{\\text\{atompair\}\}\]tensor, at the cost of recomputing the coordinate\-dependent embeddingsnblockn\_\{\\text\{block\}\}times per step\. This path is gated by the environment variableRFD3\_LOW\_MEMORY\_MODE; both Design\-CP schemes auto\-enable it wheneverRFD3\_ATTENTION\_PARALLELis set, because the row/quadrant striping reduces𝒪​\(L2\)\\mathcal\{O\}\(L^\{2\}\)to𝒪​\(L2/P\)\\mathcal\{O\}\(L^\{2\}/P\)but only the chunked path pushes the per\-rank atom\-pair memory all the way down to𝒪​\(L​k/P\)\\mathcal\{O\}\(Lk/P\)\.

#### EDM denoising loop and Karras parameters\.

Structure generation follows the EDM framework\(Karraset al\.,[2022](https://arxiv.org/html/2607.05439#bib.bib142)\)with the Algorithm\-18 schedule from the AlphaFold\-3 supplement\. The default trajectory isT=200T=200denoising steps witht∈\[0,1\]t\\in\[0,1\]linearly spaced, and the per\-step noise level is

t^=σdata​\(smax1/p\+t​\(smin1/p−smax1/p\)\)p,\\hat\{t\}\\;=\\;\\sigma\_\{\\text\{data\}\}\\bigl\(s\_\{\\text\{max\}\}^\{1/p\}\+t\\,\(s\_\{\\text\{min\}\}^\{1/p\}\-s\_\{\\text\{max\}\}^\{1/p\}\)\\bigr\)^\{p\},\(3\)with stock defaultsσdata=16\\sigma\_\{\\text\{data\}\}=16,smin=4×10−4s\_\{\\text\{min\}\}=4\\\!\\times\\\!10^\{\-4\},smax=160s\_\{\\text\{max\}\}=160, andp=7p=7\. At each step the sampler optionally injects a Karras\-style stochastic noise augmentation of magnitudeγ0=0\.6\\gamma\_\{0\}=0\.6whent^\>γmin=1\.0\\hat\{t\}\>\\gamma\_\{\\text\{min\}\}=1\.0\(otherwiseγ=0\\gamma=0\), perturbs the current state byϵ∼𝒩​\(0,σnoise2​I\)\\boldsymbol\{\\epsilon\}\\sim\\mathcal\{N\}\(0,\\sigma\_\{\\text\{noise\}\}^\{2\}I\)withσnoise=1\.003\\sigma\_\{\\text\{noise\}\}=1\.003, calls the diffusion module to obtain the clean prediction𝐗^0\\hat\{\\mathbf\{X\}\}\_\{0\}, and updates

𝐗\(t\+1\)=𝐗noisy\(t\)\+s​\(ct−t^\)​𝐗noisy\(t\)−𝐗^0t^,\\mathbf\{X\}^\{\(t\+1\)\}\\;=\\;\\mathbf\{X\}^\{\(t\)\}\_\{\\text\{noisy\}\}\+s\\,\(c\_\{t\}\-\\hat\{t\}\)\\,\\frac\{\\mathbf\{X\}^\{\(t\)\}\_\{\\text\{noisy\}\}\-\\hat\{\\mathbf\{X\}\}\_\{0\}\}\{\\hat\{t\}\},\(4\)with step scales=1\.5s=1\.5\. The recycling loop sits inside this update: between successive denoising steps the model holds the predicted distogram𝐃self\\mathbf\{D\}\_\{\\text\{self\}\}obtained by bucketising the predicted Cα\\alphacoordinates \(uniform6565\-bin distance grid over\[2,22\]​Å\[2,22\]\\,\\textup\{\\AA \}\) and consumes it as the self\-conditioning input of the next iteration\. The bucketising routine and the self\-conditioning channel are unchanged from stock RFD3 and are reused as\-is by both Design\-CP schemes\.

### A\.2Symmetry in RFD3 design

#### Symmetric inference loop\.

The stock RFD3 inference path supports cyclic \(CnC\_\{n\}\), dihedral \(DnD\_\{n\}\), and arbitrary user\-supplied \(input\_defined\) point groups: the first two are produced analytically from the ordernnon the fly, while the third is loaded from a user\-provided frames file and validated against the ASU at run\-time\. Each point group is materialised as a list ofGGrotation matrices\{Rg\}g=1G∈SO​\(3\)\\\{R\_\{g\}\\\}\_\{g=1\}^\{G\}\\in\\mathrm\{SO\}\(3\)paired with zero translations; forCnC\_\{n\},G=nG=nrotations about a common axis are evenly spaced by2​π/n2\\pi/n, and forDnD\_\{n\},2​n2nrotations combine the cyclic axis withnnorthogonalC2C\_\{2\}axes\. ASymmetryConfigdataclass groups the chosen identifier together with two additional handles –is\_unsym\_motif, a comma\-separated list of contig or ligand identifiers \(DNA strands, small\-molecule cofactors\) that should not be replicated by the symmetry operation, andis\_symmetric\_motif, a Boolean controlling whether the supplied input is already symmetric or must itself be replicated – and is the single entry point used by both Design\-CP schemes\.

#### ASU\-based input construction\.

Given an ASU atom array, the inference engine first appends the symmetry metadata \(a per\-atom subunit index and the current frame stack\), then walks through the frame list and producesGGcopies of the ASU by applying eachRgR\_\{g\}to the ASU coordinates; unsymmetrised motifs \(DNA, ligands, any contig matched byis\_unsym\_motif\) are excised before the per\-frame replication and re\-appended at the end of the resulting atom array, which keeps them out of the symmetric coordinate pool while preserving their indexing inside the model\. Whenis\_symmetric\_motif=True, the frames inferred analytically from the symmetry identifier are first reconciled with the empirical frames recovered from the input by aligning corresponding chains via SVD; this provides a sanity check that the user\-supplied symmetric input actually obeys the requested point group, and produces the per\-frame translations that the analytic axis\-aligned frames lack\.

#### Per\-step symmetrisation procedure\.

At each denoising step within the symmetrised portion of the trajectory \(controlled by asym\_step\_fracparameter, default0\.90\.9, covering the first90%90\\%of steps\): \(i\) the*full*\(unsymmetrised\) noised coordinates𝐗\(t\)\\mathbf\{X\}^\{\(t\)\}are passed to the network, which predicts clean coordinates𝐗^0\\hat\{\\mathbf\{X\}\}\_\{0\}for all chains; \(ii\) the predicted coordinates are centred by subtracting the centroid of the non\-fixed atoms \(atoms tagged by the motif/ligand\-exclusion masks above are excluded from the centroid\), the ASU slice of𝐗^0\\hat\{\\mathbf\{X\}\}\_\{0\}is extracted, and the non\-ASU chains are overwritten byRg​𝐗^0,ASUR\_\{g\}\\,\\hat\{\\mathbf\{X\}\}\_\{0,\\text\{ASU\}\}forg=2,…,Gg=2,\\dots,G; \(iii\) this symmetrised prediction is used in the EDM update of Eq\. \([4](https://arxiv.org/html/2607.05439#A1.E4)\)\. The noise injected at each step is*not*symmetrised, so the noised coordinates seen by the network are not exactly symmetric; only the predicted output is forced to be exactly symmetric at each symmetrised step\. For the final10%10\\%of steps, no symmetry is enforced, allowing the model to relax any residual strain at the inter\-subunit interfaces before output\.

#### Composition with context parallelism\.

The procedure above is applied independently on every rank: because rank0broadcasts the noised coordinates, the sampled Gaussian noise, and the network’s clean prediction at every diffusion step \(Appendix[B\.3](https://arxiv.org/html/2607.05439#A2.SS3)for the 1D scheme, Appendix[C\.5](https://arxiv.org/html/2607.05439#A3.SS5)for the 2D scheme\), the inputs to step \(ii\) above are bit\-identical across ranks, and the deterministic ASU extraction and frame application reproduce the same symmetrised prediction on every device\. No additional communication is needed for symmetric design beyond the ones that the parallel schemes already perform\. This is the operational sense in which the symmetrisation procedure is*orthogonal*to context parallelism, and it is the reason a single set of pretrained weights designs both the unsymmetrised baselines of §[4\.1](https://arxiv.org/html/2607.05439#S4.SS1)and the icosahedral and octahedral assemblies of §[4\.2](https://arxiv.org/html/2607.05439#S4.SS2)–§[4\.4](https://arxiv.org/html/2607.05439#S4.SS4)\.

## Appendix BDesign\-CP 1D row\-sharding implementation details

This appendix collects the engineering details that make the 1D scheme of §[3\.2](https://arxiv.org/html/2607.05439#S3.SS2)work in practice: the chunk\-distribution algorithm, the per\-collective communication volume, the determinism guard required at the atom→\\totoken boundary, the per\-component memory optimisations layered on top of row striping, and a few smaller bookkeeping items \(class structure, environment variables, checkpoint compatibility\)\. The 2D\-specific machinery that adapts the Fold\-CP framework to RFD3 is collected separately in Appendix[C](https://arxiv.org/html/2607.05439#A3)\.

### B\.1Chunk distribution algorithm

Tokens \(or atoms\) are distributed acrossPPGPUs using floor division with the remainder assigned to the early ranks:

Ip=\{⌊I/P⌋\+1if​p<ImodP,⌊I/P⌋otherwise,startp=p⋅⌊I/P⌋\+min⁡\(p,ImodP\)\.I\_\{p\}=\\begin\{cases\}\\lfloor I/P\\rfloor\+1&\\text\{if \}p<I\\bmod P,\\\\ \\lfloor I/P\\rfloor&\\text\{otherwise\},\\end\{cases\}\\quad\\text\{start\}\_\{p\}=p\\cdot\\lfloor I/P\\rfloor\+\\min\(p,\\,I\\bmod P\)\.\(5\)Chunk sizes therefore differ by at most one across GPUs, bounding the load imbalance per attention block to a single row regardless ofPP\. TheAllGatheroperations handle uneven chunks by padding each GPU’s local tensor to the per\-rank maximum before the collective and trimming the concatenated result back to the exact total sizeII\. The same algorithm is reused at the atom level withLLreplacingII\.

### B\.2Communication volume analysis

[Table2](https://arxiv.org/html/2607.05439#A2.T2)details the collective communication operations executed during each recycling iteration of the 1D scheme\. All sizes assumebfloat16precision \(2 bytes per element\)\. The total per\-recycle volume per rank is dominated by the1818AllGather\(𝐀\\mathbf\{A\}\) calls from the diffusion transformer; each rank’s message size is set byI⋅ctokenI\\cdot c\_\{\\text\{token\}\}and is independent ofPP\. The singleBroadcast\(𝐀I\\mathbf\{A\}\_\{I\}\) afterprocess\_ais a correctness requirement explained in §[B\.3](https://arxiv.org/html/2607.05439#A2.SS3)\. Per\-rank message size is independent ofPP, but the latency component of NCCLAllGathergrows withPP, so the per\-block communication time still increases with the device count even when each rank’s payload is held fixed; this latency\-bound term is the operative constant behind the 1D vs\. 2D strong\-scaling gap reported in §[4\.3](https://arxiv.org/html/2607.05439#S4.SS3)\.

Table 2:Collective communication in 1D parallel inference, counted per recycling iteration\. Multiply bynrecyclen\_\{\\text\{recycle\}\}\(default22\) to obtain the per\-denoising\-step cost\.†\\daggerThe token initialiser fires only once per trajectory, not once per step\.‡\\ddaggerSampling\-loop broadcasts are issued once per denoising step, independently ofnrecyclen\_\{\\text\{recycle\}\}; the optional rotation\-augmentation broadcast is what takes the𝐗^/ϵ/𝐗noisy\\hat\{\\mathbf\{X\}\}/\\boldsymbol\{\\epsilon\}/\\mathbf\{X\}\_\{\\text\{noisy\}\}count from22to33, and the optional sequence\-logits broadcast appears only when sequence design is enabled\.

### B\.3Non\-determinism guard:process\_abroadcast

The functionprocess\_aaggregates atom\-level features to the token level usingtorch\.Tensor\.index\_reduce\(\.\.\., "mean"\)\.index\_reduceperforms atomic floating\-point accumulations whose ordering depends on per\-device scheduling, and is therefore non\-deterministic across GPU devices\. In 1D parallel mode, this causes each rank to compute slightly different token\-level features𝐀I\\mathbf\{A\}\_\{I\}– the empirical magnitude of these discrepancies is on the order of∼0\.06\{\\sim\}0\.06inbfloat16\. Because𝐀I\\mathbf\{A\}\_\{I\}subsequently serves as keys and values for the cross\-attention transformer, the small discrepancies are amplified by the downstream linear projections to∼0\.5\{\\sim\}0\.5, which corrupts the cross\-attention invariant that all ranks must share identical𝐊\\mathbf\{K\}and𝐕\\mathbf\{V\}for the per\-block formulation in Eq\. \([2](https://arxiv.org/html/2607.05439#S3.E2)\) to produce identical outputs across ranks\. The fix is a singleBroadcastof𝐀I\\mathbf\{A\}\_\{I\}from rank0to all ranks immediately afterprocess\_areturns, before any downstream operation that depends on𝐀I\\mathbf\{A\}\_\{I\}being identical across ranks\. The corresponding boundary in the 2D scheme is handled differently and described in Appendix[C\.4](https://arxiv.org/html/2607.05439#A3.SS4)\.

### B\.4Per\-component memory optimisations

[Table3](https://arxiv.org/html/2607.05439#A2.T3)summarises the per\-GPU memory reductions achieved by each optimisation in the 1D parallel inference pipeline\. Three of these \(manual SwiGLU decomposition, pre\-allocation in place of concatenation, and relative\-position\-encoding sub\-chunking\) are non\-trivial enough to warrant prose explanations; the remaining three follow directly from row striping and from the chunked pairwise embedder that is auto\-enabled together withRFD3\_ATTENTION\_PARALLEL\(cf\. Appendix[B\.5](https://arxiv.org/html/2607.05439#A2.SS5)\)\.

Table 3:Memory optimisations in 1D parallel inference\.II: token count,LL: atom count,PP: number of GPUs,kk: sparse\-attention neighbour budget\. The asymptotic per\-GPU pair\-track memory after𝐏\\mathbf\{P\}sparsification is𝒪​\(L⋅k/P\)\\mathcal\{O\}\(L\\cdot k/P\)\.
#### Manual SwiGLU decomposition\.

The pairwise transition layers use a SwiGLU feed\-forward network\(Shazeer,[2020](https://arxiv.org/html/2607.05439#bib.bib34)\):

Transition⁡\(𝐙\)=W3​\(SiLU⁡\(W1​LN⁡\(𝐙\)\)⊙W2​LN⁡\(𝐙\)\),\\operatorname\{Transition\}\(\\mathbf\{Z\}\)=W\_\{3\}\\bigl\(\\operatorname\{SiLU\}\(W\_\{1\}\\operatorname\{LN\}\(\\mathbf\{Z\}\)\)\\odot W\_\{2\}\\operatorname\{LN\}\(\\mathbf\{Z\}\)\\bigr\),\(6\)whereW1,W2∈ℝcz×4​czW\_\{1\},W\_\{2\}\\in\\mathbb\{R\}^\{c\_\{z\}\\times 4c\_\{z\}\}andW3∈ℝ4​cz×czW\_\{3\}\\in\\mathbb\{R\}^\{4c\_\{z\}\\times c\_\{z\}\}\. In the standard residual computation𝐙←𝐙\+Transition⁡\(𝐙\)\\mathbf\{Z\}\\leftarrow\\mathbf\{Z\}\+\\operatorname\{Transition\}\(\\mathbf\{Z\}\), PyTorch retains all intermediate tensors simultaneously, including the4×4\{\\times\}\-expanded linear outputs\. We manually decompose the computation with explicit deallocation \([Algorithm1](https://arxiv.org/html/2607.05439#alg1)\)\. This is mathematically identical but ensures that at most two4​cz4c\_\{z\}\-expanded tensors coexist at any point, substantially reducing peak memory\. The optimisation is applied only during inference \(gated ontorch\.is\_grad\_enabled\(\) == False\); training uses the standard path to preserve compatibility with activation checkpointing\.

Algorithm 1Manual SwiGLU decomposition \(memory\-friendly inference implementation\)\.Input:residual activations

𝐙\\mathbf\{Z\}; weights

W1,W2,W3W\_\{1\},W\_\{2\},W\_\{3\}
Output:updated activations

𝐙\\mathbf\{Z\}
𝐍←LayerNorm⁡\(𝐙\)\\mathbf\{N\}\\leftarrow\\operatorname\{LayerNorm\}\(\\mathbf\{Z\}\)

𝐀←W1​𝐍\\mathbf\{A\}\\leftarrow W\_\{1\}\\mathbf\{N\}

𝐁←W2​𝐍\\mathbf\{B\}\\leftarrow W\_\{2\}\\mathbf\{N\}

Free:

𝐍\\mathbf\{N\}
𝐀←SiLU⁡\(𝐀\)\\mathbf\{A\}\\leftarrow\\operatorname\{SiLU\}\(\\mathbf\{A\}\)

𝐀←𝐀⊙𝐁\\mathbf\{A\}\\leftarrow\\mathbf\{A\}\\odot\\mathbf\{B\}

Free:

𝐁\\mathbf\{B\}
𝐀←W3​𝐀\\mathbf\{A\}\\leftarrow W\_\{3\}\\mathbf\{A\}

𝐙←𝐙\+𝐀\\mathbf\{Z\}\\leftarrow\\mathbf\{Z\}\+\\mathbf\{A\}

Free:

𝐀\\mathbf\{A\}

#### Pre\-allocation vs concatenation\.

The diffusion token encoder constructs an augmented pairwise representation by concatenating three components along the channel dimension:

𝐙aug=\[𝐙init​‖𝐃dist‖​𝐃self\]∈ℝB×Ip×I×\(cz\+cz\+nbins\)\.\\mathbf\{Z\}\_\{\\text\{aug\}\}=\\bigl\[\\mathbf\{Z\}\_\{\\text\{init\}\}\\;\\\|\\;\\mathbf\{D\}\_\{\\text\{dist\}\}\\;\\\|\\;\\mathbf\{D\}\_\{\\text\{self\}\}\\bigr\]\\in\\mathbb\{R\}^\{B\\times I\_\{p\}\\times I\\times\(c\_\{z\}\+c\_\{z\}\+n\_\{\\text\{bins\}\}\)\}\.\(7\)A naivetorch\.catrequires allocating the output tensor*in addition to*all three inputs, causing a transient memory spike that more than triples the peak at this stage\. The 1D parallel implementation pre\-allocates a single tensor of the final shape usingtorch\.emptyand writes each component into its designated channel slice in\-place, deleting the source tensor immediately after each copy\. This avoids the concatenation overhead and reduces the peak memory of the augmentation step from3×3\\timesto1×1\\timesthe size of𝐙aug\\mathbf\{Z\}\_\{\\text\{aug\}\}\.

#### Relative position encoding sub\-chunking\.

The relative\-position\-encoding \(RPE\) bias requires materialising one\-hot tensors of shape\[Ip,I,nbins\]\[I\_\{p\},I,n\_\{\\text\{bins\}\}\], which can be large even after row striping\. The 1D parallel implementation processes the RPE in44sub\-chunks along the query dimension: each sub\-chunk computes its slice of the residue, token, and chain one\-hot encodings, applies the linear projection that turns them into a singleczc\_\{z\}\-channel bias, and deletes the intermediates before proceeding to the next sub\-chunk\. This reduces the peak memory of the RPE computation by4×4\\times\.

### B\.5Class structure and environment variables

The 1D scheme is realised through two parallel classes –ParallelTokenInitializer\(which subclassesTokenInitializer\) andParallelDiffusionModule\(which subclassesRFD3DiffusionModule\) – that share identical learned parameters with their serial counterparts but overrideforward\(\)to operate on row\-striped representations\. A third class,ParallelDiffusionTokenEncoder, is created via a Python\_\_class\_\_swap insideParallelDiffusionModule\.\_\_init\_\_: the standardDiffusionTokenEncoderinstance is reassigned to the parallel subclass without re\-loading parameters, which is safe because the parallel variant introduces no new parameters and only replaces the forward pass\. As a consequence, checkpoints trained on a single GPU are loaded directly without any conversion or key remapping\.

A factory inRFD3\.\_\_init\_\_reads the environment at construction time and instantiates the parallel classes whenever the relevant flag is set\. The 1D scheme is controlled primarily by three environment variables\.RFD3\_ATTENTION\_PARALLELenables parallel mode and selects the world size; the launcher script sets it to the detected GPU count wheninference\_parallel=Trueis passed on the command line, after auto\-relaunching the script undertorchrunviamaybe\_relaunch\_with\_torchrun\.RFD3\_LOW\_MEMORY\_MODEenables the chunked pairwise embedder that avoids materialising the dense\[L,L,catompair\]\[L,L,c\_\{\\text\{atompair\}\}\]atom pair tensor; it is auto\-enabled wheneverRFD3\_ATTENTION\_PARALLELis set, because the row\-striped path requires sparse𝐏\\mathbf\{P\}to fit the largest assemblies that motivate context parallelism\.RFD3\_NCCL\_TIMEOUT\_SECis an optional per\-collective NCCL watchdog timeout, defaulting to18001800s – raised from PyTorch’s600600s default to tolerate the serial post\-processing that rank0performs \(atom\-array cleanup, file I/O\) while other ranks have already enqueued the next collective\. A fourth optional flag,RFD3\_EXTRA\_CHUNKING, controls a finer sub\-chunking inside the chunked pairwise embedder for very largeLLbut is rarely needed in practice\.

The setup sequence is: \(i\) the launcher detectsinference\_parallel=Trueinsys\.argvand re\-execs the script undertorchrunwith\-\-nproc\_per\_nodeequal to the visible GPU count; \(ii\) child processes callsetup\_distributed\(\), which initialises an NCCL process group with the elevated timeout; \(iii\) the engine setsRFD3\_LOW\_MEMORY\_MODE=1andRFD3\_ATTENTION\_PARALLEL=world\_size; \(iv\)RFD3\.\_\_init\_\_reads those variables and constructs the parallel module tree; \(v\) the engine broadcasts model parameters from rank0to all other ranks before inference begins, since the Lightning FabricSingleDeviceStrategywe use does not implicitly replicate weights\.

## Appendix CDesign\-CP 2D grid implementation details

This appendix expands the RFD3\-specific aspects of the 2D scheme of §[3\.3](https://arxiv.org/html/2607.05439#S3.SS3)\. The Fold\-CP framework\(Linet al\.,[2026](https://arxiv.org/html/2607.05439#bib.bib82)\)supplies the ring\-attention core – transposition of K/V shards at the start of a ring loop,P\\sqrt\{P\}ring shifts with online\-softmax merging, and the boundary communicators – which we reuse unchanged\. The contributions documented below concern \(i\) how that core is plugged into RFD3’s denoising loop, \(ii\) how the atom\-level sparse attention is made compatible with 2D sharding, \(iii\) how the dense pre\-pipeline data structures of RFD3 are deferred so that the row/column quadrants can be reconstructed on rank, and \(iv\) how parameters and gradients are laid out on the device mesh\. We will release the accompanying code in a future update to the official RFDiffusion 3 repository \(https://github\.com/RosettaCommons/foundry/tree/production\)

### C\.1Device mesh and DTensor parameter distribution

Devices are arranged on aP×P\\sqrt\{P\}\\times\\sqrt\{P\}context\-parallel mesh, optionally combined with a third data\-parallel axis when training; we refer to the two CP axes ascp0\\text\{cp\}\_\{0\}\(the row axis\) andcp1\\text\{cp\}\_\{1\}\(the column axis\)\. The mesh and the corresponding process subgroups are produced by aDistributedManagerobject that the engine instantiates before constructing the model; the perfect\-square requirement onPPis enforced inside the manager, and the model code only ever consumes the precomputeddevice\_mesh\_subgroupsandlayout\_subgroupsdictionaries\.

At model construction, everyLinear,LayerNorm, and RMSNorm\-shaped module in the serial RFD3 tree is replaced by a parameter\-replicated DTensor wrapper \(LinearParamsReplicated,LayerNormParamsReplicated\) so that its weights live as PyTorch DTensors with aReplicate\(\)placement on every CP axis\. A runtime check,validate\_all\_params\_are\_dtensors, traverses the entire module tree before inference begins and asserts that no trainable tensor has silently escaped the wrapping; this catches any custom layer that constructs parameters outside the standardnn\.Linear/nn\.LayerNormpaths\.

The 2D scheme does not use a\_\_class\_\_swap\. Instead, a top\-level helpercreate\_rfd3\_distributedinstantiates a small set of CP wrapper modules \(CPTokenInitializer,CPTokenTransformer,CPDiffusionTokenEncoder,CPDiffusionModule\) and replaces theforwardattribute of the corresponding serial submodule with the wrapper’sforwardmethod\. This composition pattern matches Fold\-CP’s convention and avoids touching the construction\-time logic of the serial classes, so checkpoints trained on a single GPU are loaded unchanged\.

### C\.2Q/K/V layout and the ring loop

At each attention block, queries are sharded alongcp0\\text\{cp\}\_\{0\}and replicated alongcp1\\text\{cp\}\_\{1\}, while keys and values are sharded alongcp1\\text\{cp\}\_\{1\}and ring\-rotated\. The ring loop is opened by a singleTransposeCommthat swaps K/V between grid positions\(r,c\)\(r,c\)and\(c,r\)\(c,r\), aligning the K/V shards with the resident Q rows\. The loop then performsP\\sqrt\{P\}ring shifts: at each step a GPU computes attention between its resident Q rows and the currently visited K/V shard, using the local pair\-bias quadrant; partial outputs are merged into a running total via the online\-softmax kerneltiled\_softmax\_attention\_updatereused from Fold\-CP\. Each ring step issues anAttentionPairBiasComm, which bundles the K/V/𝐁\\mathbf\{B\}shift and overlaps the communication with local computation, and a small number ofOne2OneCommpoint\-to\-point exchanges that move auxiliary tensors \(token indices, chain identifiers, distogram features\) along the same column path as the K/V tiles\.

Asymptotically, per\-device pair\-track memory is𝒪​\(I2/P\)\\mathcal\{O\}\(I^\{2\}/P\)– the same as the 1D scheme – but K/V are ring\-rotated rather than replicated, so each device only ever holds an𝒪​\(I/P\)\\mathcal\{O\}\(I/\\sqrt\{P\}\)slab of K/V at a time; the pair tensor itself is therefore stored as a local quadrant\[Ir,Ic,cz\]\[I\_\{r\},I\_\{c\},c\_\{z\}\]withIr=Ic=I/PI\_\{r\}=I\_\{c\}=I/\\sqrt\{P\}, and is never gathered to its full\[I,I\]\[I,I\]shape on any rank\.

### C\.3Distributed kNN for sparse atom attention

RFD3’s atom\-level attention selects thekknearest neighbours of every query atom from the current predicted coordinates\. A naive implementation computes an\[L,L\]\[L,L\]distance matrix and takes the top\-kkper row; at the scales where 2D CP is useful, that distance matrix is exactly the object we cannot afford\. Our distributed kNN proceeds in three stages\. First, the small 1D feature tensors – token identifiers and chain assignments, each of shape\[L\]\[L\]– are allgathered alongcp0\\text\{cp\}\_\{0\}so every rank has global indexing for downstream masking; these are𝒪​\(L\)\\mathcal\{O\}\(L\)in size and cheap to replicate\. Second, each rank computes Euclidean distances between its local atom rows and the atoms currently resident on its column partner, producing a per\-quadrant distance block of shape\[Lr,Lc\]\[L\_\{r\},L\_\{c\}\]withLr=Lc=L/PL\_\{r\}=L\_\{c\}=L/\\sqrt\{P\}; this is implemented as a chunkedcdist\(default chunk size10241024rows\) so that even the per\-quadrant block is never materialised in full\. Third, aring\_topkprimitive rotates partial top\-kkcandidates alongcp1\\text\{cp\}\_\{1\}; at each ring step, the local top\-kkcandidates are merged with incoming candidates from the column neighbour to produce a running global top\-kkper query atom\. AfterP\\sqrt\{P\}ring steps, every row\-resident rank holds the global top\-kkindices for its own query atoms, and no further broadcast is needed because the downstream sparse ring attention consumes the indices in place\. The end\-to\-end computation never materialises an\[L,L\]\[L,L\]tensor on any device\.

The indices then feed a sparse ring attention \(sparse\_ring\_attention\_forward\)\. At each ring step, the global indices are filtered to the current column block, K/V/𝐁\\mathbf\{B\}entries are gathered at those filtered indices, and the resulting\[D,H,Lr,k\]\[D,H,L\_\{r\},k\]logits are merged into the running softmax via the same online\-softmax kernel used by the dense token\-level ring loop\.

### C\.4Atom\-to\-token pooling and the determinism boundary

The 1D scheme requires a rank\-0Broadcastof𝐀I\\mathbf\{A\}\_\{I\}immediately afterprocess\_a\(Appendix[B\.3](https://arxiv.org/html/2607.05439#A2.SS3)\), becausetorch\.Tensor\.index\_reduce\("mean"\)is non\-deterministic across devices\. The 2D scheme handles the same boundary in a structurally different way: the atom→\\totoken pooling is implemented as aDistributedScatterReducecollective alongcp0\\text\{cp\}\_\{0\}, with the per\-element scatter indices coming from the \(already replicated\) globaltok\_idx\. Because the reduction is performed by a single deterministic collective rather than by per\-device atomic accumulators, the resulting𝐀I\\mathbf\{A\}\_\{I\}is bit\-identical across ranks by construction and no separate broadcast is needed\. The same primitive also implements the irregular Cα\\alpha/motif\-token pooling consumed by the distogram path of the diffusion token encoder\.

### C\.5Sampling\-loop integration

The ring primitive is invoked inside every recycling iteration of every denoising step, identically to how the 1D scheme invokes its per\-block cross\-attention\. At diffusion\-step boundaries, rank0broadcasts the current noised coordinates, the freshly sampled Gaussian noise, and the denoised prediction; symmetrisation, when active, is then applied independently on every rank using the broadcast inputs, so that the trajectory stays deterministic without any extra mesh\-aware bookkeeping\. Because the 1D and 2D schemes share the same step\-boundary broadcast pattern, they see the same stochastic trajectory for a given seed, which simplifies head\-to\-head comparisons of correctness and timing\.

### C\.62D\-specific data\-pipeline transforms

Two pre\-pipeline transforms are introduced for 2D CP to avoid materialising dense\[I,I\]\[I,I\]matrices in the data loader\. Both store a small 1D representation in the feature dictionary and defer the reconstruction of the local\[Ir,Ic\]\[I\_\{r\},I\_\{c\}\]quadrant to the model’s input layer, where the quadrant is built directly as a DTensor on the resident grid position\.CPAwareUnindexFlaggedTokensreplaces the standard\[I,I\]\[I,I\]unindexing pair mask by a pair of\[I\]\[I\]\-shaped tensors – a per\-tokenis\_unindexedboolean and agroup\_idsinteger assignment – from which the quadrant of the mask is recomputed on\-rank\.AddAF3TokenBondFeaturesreplaces the dense\[I,I\]\[I,I\]token\-bond matrix by a sparse COO representation when the structure exceeds a configurable atom\-count threshold \(default5×1045\\times 10^\{4\}\), again reconstructing the quadrant on\-rank insideCPTokenInitializer\. Together, these transforms keep the data loader’s memory footprint bounded by𝒪​\(I\)\\mathcal\{O\}\(I\)even for the largest assemblies we target, where the dense matrices would already exceed several hundred megabytes per sample\.

### C\.7Per\-component memory optimisations carried over to 2D

Several of the optimisations of Appendix[B\.4](https://arxiv.org/html/2607.05439#A2.SS4)carry over to the 2D scheme; others are subsumed by the inherently quadrant\-based layout\. Specifically:

- •𝐙\\mathbf\{Z\}striping and𝐃self\\mathbf\{D\}\_\{\\text\{self\}\}chunking are not separate optimisations on 2D: every shape that the 1D scheme writes as\[I/P,I,⋅\]\[I/P,I,\\cdot\]exists on 2D as a\[Ir,Ic,⋅\]\[I\_\{r\},I\_\{c\},\\cdot\]DTensor by construction\.
- •𝐏\\mathbf\{P\}sparsification reuses the same chunked pairwise embedder as 1D, with the static MLP projections cached once at tokenisation; the only difference is that on 2D, the per\-rank slice of𝐏\\mathbf\{P\}is a quadrant rather than a row stripe\.
- •𝐙aug\\mathbf\{Z\}\_\{\\text\{aug\}\}pre\-allocation in place ofcatis again applied to the diffusion\-token\-encoder concatenation, since the channel\-dimension cat is the same regardless of whether the leading two dimensions are sharded as\[I/P,I,⋅\]\[I/P,I,\\cdot\]or\[Ir,Ic,⋅\]\[I\_\{r\},I\_\{c\},\\cdot\]\.
- •Manual SwiGLU decomposition and the explicit4×4\{\\times\}RPE sub\-chunking from the 1D scheme are not currently applied on 2D, because the local per\-rank𝐙\\mathbf\{Z\}quadrant is small enough that the unmodified Fold\-CP transition and RPE primitives stay below the per\-GPU memory budget at the assembly sizes we target\. They could be ported across schemes if a future 2D configuration moved the bottleneck back to the transition or RPE blocks\.

### C\.8Class structure and environment variables for the 2D scheme

The 2D entry point iscreate\_rfd3\_distributed\(rfd3, manager\), which is invoked by the engine after model construction and after theDistributedManagerhas produced its mesh and subgroup layouts\. The helper instantiates the four CP wrapper modules listed in Appendix[C\.1](https://arxiv.org/html/2607.05439#A3.SS1), attaches theirforwardmethods to the corresponding serial submodules, redistributes the parameters viadistribute\_params, and runs the post\-construction DTensor validation check\.

The 2D scheme reusesRFD3\_ATTENTION\_PARALLELas its top\-level on/off switch and inheritsRFD3\_LOW\_MEMORY\_MODE\(which controls the chunked pairwise embedder shared with 1D\) andRFD3\_NCCL\_TIMEOUT\_SEC\(the elevated NCCL watchdog timeout\)\. The atom\-level inter\-chain attention budget is configured through three additional optional variables –RFD3\_N\_ATTN\_KEYS,RFD3\_INTER\_CHAIN\_WEIGHT, andRFD3\_ATOM\_USE\_ICA\(with companionRFD3\_ATOM\_INTER\_CHAIN\_WEIGHT\) – and two debug\-oriented variables –RFD3\_DEBUG\_STATSandRFD3\_MEM\_PROFILE– toggle a checkpoint\-statistics logger and a per\-step memory profile, respectively\. These last five variables are not strictly part of the CP machinery but are exposed by the same code path because they affect the per\-block cost of the ring attention and are therefore relevant for reproducing the timings reported in §[4\.3](https://arxiv.org/html/2607.05439#S4.SS3)\.

## Appendix DDesignability metrics

This appendix provides the formal definitions of the*in silico*metrics used in §[4\.2](https://arxiv.org/html/2607.05439#S4.SS2)and §[4\.4](https://arxiv.org/html/2607.05439#S4.SS4)\. All quantities are computed directly from the generated all\-atom coordinates of the symmetrised assembly; no external structure\-prediction oracle is required\. We split the metrics into two families: backbone sanity \(per\-design checks on the diffused chain itself\) and symmetric\-interface sanity \(checks that probe the inter\-subunit geometry of the assembly\)\.

#### Backbone sanity\.

These metrics flag designs whose monomeric chain is geometrically broken, irrespective of any symmetry consideration\.

- •Chain breaks \(n\_chainbreaks\)\.For every consecutive pair of Cα\\alphaatoms along the chain, we compute the bond length and its deviation from the canonical3\.83\.8Å; pairs with deviation aboveτcb=0\.75\\tau\_\{\\text\{cb\}\}=0\.75Å are flagged as chain breaks\. Pairs that span an intentional chain transition \(differentchain\_iid\) are masked out so they do not contribute\. We additionally reportmax\_ca\_deviation, the worst per\-design deviation, as a continuous summary\. A geometrically clean monomer should haven\_chainbreaks= 0\.
- •Inter\-residue clashes \(n\_clashing\.interresidue\_clashes\_w\_backboneand\.\.\.\_w\_sidechain\)\.We pairwise compare all heavy atoms of the protein and count pairs that \(i\) belong to residues at least two apart along the sequence and \(ii\) are closer thanτclash=1\.5\\tau\_\{\\text\{clash\}\}=1\.5Å\. The backbone\-only variant restricts the second factor to atoms in\{N,C​α,C\}\\\{\\text\{N\},\\text\{C\}\\alpha,\\text\{C\}\\\}and is the more conservative indicator of physical implausibility because the backbone has no rotameric flexibility to relieve the clash\. The sidechain variant is much noisier and is reported only as a lower bound on clash density\.
- •Non\-loopy fraction \(non\_loop\_fraction\)\.We run the P\-SEA secondary\-structure annotator\(Labesseet al\.,[1997](https://arxiv.org/html/2607.05439#bib.bib79)\)\(as implemented in Biotite’sannotate\_sse\) on the diffused chain and report the fraction of residues assigned to helix or strand, i\.e\. the complement of the coil fraction\.

#### Symmetric\-interface sanity\.

These metrics use the full symmetrised complex and exclude any atoms withsym\_transform\_id<0<0\(fixed/unsymmetrised motifs\)\. All distances are Cα\\alpha–Cα\\alpha\. We use three thresholds that come from the implementation inrfd3\.inference\.symmetry\.metrics: a hard inter\-subunit clash distanceτclashsym=3\.5\\tau\_\{\\text\{clash\}\}^\{\\text\{sym\}\}=3\.5Å, an interface\-contact band\[τcontactlo,τcontacthi\]=\[4\.0,10\.0\]\[\\tau\_\{\\text\{contact\}\}^\{\\text\{lo\}\},\\tau\_\{\\text\{contact\}\}^\{\\text\{hi\}\}\]=\[4\.0,10\.0\]Å, and a proximity cutoffτprox=15\.0\\tau\_\{\\text\{prox\}\}=15\.0Å above which two subunits are considered non\-interacting\.

- •Minimum inter\-chain distance \(complex\.min\_inter\_chain\_distance\)\.The smallest Cα\\alpha–Cα\\alphadistance between atoms belonging to two distinct subunits anywhere in the complex\. Values below∼3\.5\\sim 3\.5Å indicate steric overlap; values much above∼6\\sim 6Å indicate that the asymmetric units never actually meet, which for a closed nanoparticle is a failure mode\.
- •ASU clashes \(asu\.n\_clashes\)\.The number of inter\-subunit Cα\\alpha–Cα\\alphapairs closer thanτclashsym\\tau\_\{\\text\{clash\}\}^\{\\text\{sym\}\}, normalised by the number of subunits\. The normalisation makes the metric directly comparable across symmetries with different oligomeric states\.
- •Mean contacts per interface \(complex\.mean\_contacts\_per\_interface\)\.For each pair of subunits whose Cα\\alpha–Cα\\alphaatoms come withinτprox\\tau\_\{\\text\{prox\}\}of each other we count the number of Cα\\alpha–Cα\\alphapairs falling inside the contact band\[τcontactlo,τcontacthi\]\[\\tau\_\{\\text\{contact\}\}^\{\\text\{lo\}\},\\tau\_\{\\text\{contact\}\}^\{\\text\{hi\}\}\]\. We then average this count over all proximal interfaces\. Higher values indicate richer, better\-formed interfaces\.
- •Total interface contacts \(complex\.n\_interface\_contacts\)\.The unnormalised sum of the above over all interfaces\.
- •Number of interfaces with contacts \(complex\.n\_interfaces\_with\_contacts\)andminimum contacts per interface \(complex\.min\_contacts\_per\_interface\)\.These flag assemblies in which one of the interfaces is essentially absent \(low minimum\) even when the mean is healthy\.

#### Additional secondary\-filter metrics\.

We additionally compute the radius of gyration of the diffused chain \(Biotite’sgyration\_radius\); the helix, sheet, and loop fractions and the number of secondary\-structure elements from the same P\-SEA annotation; per\-residue amino\-acid composition \(in particular alanine and glycine content, where biased composition often signals a pathological design\); and the smallest Cα\\alpha–Cα\\alphadistance*within*the ASU \(asu\.min\_intra\_distance\), which catches intra\-subunit overlap that is invisible to the inter\-subunit clash count\.

#### Protocol\.

Unless otherwise stated, the figures in §[4\.2](https://arxiv.org/html/2607.05439#S4.SS2)aggregate overn=40n=40independent designs targeting an icosahedral assembly with210210residues per chain, generated with the 2D sharding scheme on a2×22\\times 2device grid of HG200 GPUs \(9595GB each\) and using a batch size of one\.

## Appendix EPer\-chain comparison: Design\-CP vs vanilla RFD3 \(icosahedral\)

The main\-text comparison in[Figure2](https://arxiv.org/html/2607.05439#S4.F2)b–e covers four panels \(chain breaks, backbone clashes, non\-loop fraction, max CA deviation\)\.[Figure4](https://arxiv.org/html/2607.05439#A5.F4)shows the full eight\-panel version, adding helix fraction, sheet fraction, glycine content, and average number of secondary\-structure elements per chain\.

#### Caveat on problem difficulty\.

The two design problems are not of equal difficulty and the comparison should be read with this asymmetry in mind\. Each icosahedral chain is denoised inside a12,60012\{,\}600\-residue joint context where its trajectory must remain self\-consistent with the simultaneously denoised trajectories of5959symmetry mates and with all the inter\-chain interfaces they form, while the monomer problem is a single210210\-residue chain comfortably inside RFD3’s native384384\-token training crop\. The icosahedral problem is therefore strictly harder per chain\.

#### Helix and sheet fraction\.

One of the most prominent qualitative gap in[Figure4](https://arxiv.org/html/2607.05439#A5.F4)appears in secondary\-structure composition\. Vanilla RFD3 monomers reproduce the helix/sheet balance similar to the one of our reference structure 1NQX: the helix fraction is concentrated around a median of≈0\.47\\approx 0\.47with a tight bulk against the 1NQX reference of0\.470\.47, and the sheet fraction sits at a median of≈0\.18\\approx 0\.18with a comparable spread, against0\.200\.20for 1NQX\. Design\-CP icosahedral designs, in contrast, are stronglyβ\\beta\-biased: the helix fraction collapses near zero on the bulk of chains with a few upper outliers and a handful of smaller ones, and the sheet fraction shifts upwards to a median of≈0\.5\\approx 0\.5with a noticeably broader distribution\.

#### Number of secondary\-structure elements\.

Consistent with the helix/sheet shift, the average number of secondary\-structure elements per chain rises from≈13\\approx 13in the monomers to≈16\\approx 16in the icosahedral designs, with the icosahedral distribution carrying a heavier upper tail\. Both populations include outliers, but only the icosahedral set reaches the upper\-twenties range\. At fixed chain length, a higher element count mechanically implies shorter average element length, which is again consistent with the icosahedral chains preferring multiple shortβ\\beta\-strands over a small number of long helices\. We treat this as supportive evidence for the helix/sheet shift rather than an independent observation, and note that the metric is sensitive to the P\-SEA assignment thresholds\(Labesseet al\.,[1997](https://arxiv.org/html/2607.05439#bib.bib79)\)\.

#### Amino\-acid composition\.

On amino\-acid composition the two populations are closer to each other than on secondary structure\. Alanine content is similar across both sets, with broadly overlapping distributions and medians on the same order, and is, as is typical for RFD3\(Butcheret al\.,[2025](https://arxiv.org/html/2607.05439#bib.bib143)\), slightly inflated relative to 1NQX in both cases; we do not read a meaningful difference between Design\-CP ASUs and vanilla monomers on this axis\. Glycine content, in contrast, is appreciably broader in the icosahedral designs \(≈0\.10\\approx 0\.10–0\.250\.25, with sporadic high outliers\) than in the monomers \(≈0\.05\\approx 0\.05–0\.100\.10, very tightly concentrated\), although the medians remain on the same order and within the typical range for natural proteins\. The widened glycine spread is qualitatively consistent at the population level with the broader max\-Cα\\alpha–Cα\\alphadeviation distribution of[Figure2](https://arxiv.org/html/2607.05439#S4.F2)e – chains under more inter\-subunit geometric strain might rely more on backbone\-flexible residues – but we flag this as a tentative association rather than a causal claim, since we have not tested whether the same chains drive both effects\.

![Refer to caption](https://arxiv.org/html/2607.05439v1/RFD3-vs-DesignCP.png)Figure 4:Per\-chain comparison of Design\-CP icosahedral designs against vanilla RFD3 monomers \(full eight panels\)\.Distributions of standard backbone\-sanity and composition metrics, computed chain\-by\-chain onn=40n=40Design\-CP icosahedral assemblies \(blue,6060chains×\\times210210residues per chain\) andn=40n=40vanilla single\-GPU RFD3 monomers of length210210\(green\)\. The dotted line in each panel reports the corresponding value for Lumazine Synthase \(PDB: 1NQX\)\. The first four panels reproduce[Figure2](https://arxiv.org/html/2607.05439#S4.F2)b–e and are commented on in the main text; the additional four panels \(helix fraction, sheet fraction, glycine content, average number of secondary\-structure elements\) reveal the secondary\-structure shift discussed in this appendix\.

## Appendix FOctahedral design metrics on commodity GPUs

This appendix reports the full distributions of*in silico*metrics for then=12n=12octahedral assemblies generated on1616NVIDIA RTX A4000 GPUs \(1616GB each\) with a batch size of one\. The headline metrics \(chain breaks, backbone clashes, non\-loop fraction, max CA deviation, minimum inter\-chain distance, mean contacts per interface\) are already shown in[Figure3](https://arxiv.org/html/2607.05439#S4.F3)b–g and discussed in §[4\.4](https://arxiv.org/html/2607.05439#S4.SS4)\.[Figure5](https://arxiv.org/html/2607.05439#A6.F5)shows the full set of backbone\-sanity and symmetry\-interface metrics; we comment below only on the additional panels that are not in the main text\.

#### Additional backbone\-sanity panels \([Figure5](https://arxiv.org/html/2607.05439#A6.F5)a\)\.

The radius of gyration is tightly concentrated in the≈56\.8\\approx 56\.8–57\.457\.4Å range, indicating a consistent per\-chain envelope across the population\. Helix and sheet fractions are both broad, with no obvious dominant secondary\-structure type within the population: helix fraction spans≈0\\approx 0–0\.90\.9with a median around≈0\.55\\approx 0\.55, and sheet fraction spans≈0\\approx 0–0\.850\.85with a median around≈0\.45\\approx 0\.45\. This is in qualitative contrast with the icosahedral designs of Appendix[E](https://arxiv.org/html/2607.05439#A5), where helix fraction collapses near zero\. We interpret this as plausibly reflecting the lower oligomeric state \(2424vs\.6060subunits\), which gives the symmetrisation step less reason to favour extendedβ\\beta\-sheets over helical packing, but caution that the sample size \(n=12n=12\) is small\. Alanine content has a median of≈0\.25\\approx 0\.25and is comparatively broad \(≈0\.05\\approx 0\.05–0\.60\.6\), while glycine content is tightly clustered around≈0\.05\\approx 0\.05–0\.100\.10\. These figures are in line with what observed in[AppendixE](https://arxiv.org/html/2607.05439#A5)\.

#### Additional symmetry\-interface panels \([Figure5](https://arxiv.org/html/2607.05439#A6.F5)b\)\.

The hard inter\-subunit clash counts at both the ASU and complex levels \(asu\.n\_clashesandcomplexclashes\) are at zero across the entire population, indicating that no design realises sterically forbidden inter\-subunit contacts\. The smallest intra\-ASU Cα\\alpha–Cα\\alphadistance is centred at≈3\.6\\approx 3\.6Å \(matching the canonical Cα\\alpha–Cα\\alphaspacing\) with two low outliers at≈1\.2\\approx 1\.2and≈2\.4\\approx 2\.4Å that flag designs with intra\-ASU geometric strain\. The number of interfaces showing contacts has a median of≈55\\approx 55out of the2424chains’ interface budget, and the total number of interface contacts spans≈200\\approx 200–3,7003\{,\}700with a median of≈1,700\\approx 1\{,\}700\. The minimum number of contacts per interface has a median near11, with two outliers at≈25\\approx 25and≈30\\approx 30; this reflects that some designs do contain weak interfaces whose contact count drags the per\-design minimum near zero, a useful filter signal for downstream design campaigns\.

Across both metric families, the additional panels are consistent with the headline result of §[4\.4](https://arxiv.org/html/2607.05439#S4.SS4): octahedral designs sampled on commodity GPUs satisfy the same hard\-failure criteria as the icosahedral baseline, with the main qualitative difference being a more permissive secondary\-structure distribution\.

![Refer to caption](https://arxiv.org/html/2607.05439v1/Octahedral_designability_metrics.png)
![Refer to caption](https://arxiv.org/html/2607.05439v1/Octahedral_symmetry_metrics.png)

Figure 5:Full distributions of*in silico*metrics for octahedral designs \(n=12n=12\)\.Headline metrics from[Figure3](https://arxiv.org/html/2607.05439#S4.F3)b–g are reproduced here together with the additional panels discussed in this appendix\.\(a\)Backbone\-sanity metrics: chain breaks, backbone clashes, non\-loop fraction, max CA deviation, helix fraction, sheet fraction, radius of gyration, alanine content, glycine content\.\(b\)Symmetry\-interface metrics: ASU clashes, ASU minimum intra\-distance, complex clashes, minimum inter\-chain distance, number of interfaces with contacts, total interface contacts, mean contacts per interface, minimum contacts per interface\.

## Appendix GEffect of removing the symmetry constraint

The main text argues \(§[4\.1](https://arxiv.org/html/2607.05439#S4.SS1)\) that strong point\-group symmetry constraints are what make Design\-CP usable on system sizes well beyond RFD3’s native384384\-token training crop\.[Figure1](https://arxiv.org/html/2607.05439#S2.F1)b already illustrated this for the extreme case of a single10,80010\{,\}800\-residue monomer\.[Figure6](https://arxiv.org/html/2607.05439#A7.F6)shows the complementary illustration for the multi\-chain regime relevant to nanoparticle design: six Design\-CP samples generated with exactly the same total system size as the icosahedral assemblies of §[4\.2](https://arxiv.org/html/2607.05439#S4.SS2)\(6060chains of210210residues each,12,60012\{,\}600residues in total\) but with*no*point\-group symmetry constraint imposed at sampling time\. Token and atom counts, the architecture, the weights, and the parallelisation strategy are kept identical to the symmetric runs; only the ASU restriction described in Appendix[A\.2](https://arxiv.org/html/2607.05439#A1.SS2)is removed\.

The resulting structures are visibly degenerate: large slabs ofβ\\beta\-strands packed in irregular orientations, no recognisable globular folds, and no consistent inter\-chain organisation\. None of these samples resemble naturally occurring multimeric proteins, and they bear no resemblance to the well\-formed icosahedral assemblies in[Figure2](https://arxiv.org/html/2607.05439#S4.F2)f despite using exactly the same number of tokens, atoms, and chains\. We take this as direct qualitative evidence that the dominant factor behind the sample\-quality results in §[4\.2](https://arxiv.org/html/2607.05439#S4.SS2)is the symmetry prior rather than the lenght of the modelled chains: at this scale, RFD3 \+ Design\-CP cannot recover plausible protein\-like geometry from joint denoising alone\.

![Refer to caption](https://arxiv.org/html/2607.05439v1/Monstromers.png)Figure 6:Effect of removing the symmetry constraint at the same system size\.Six Design\-CP samples generated with6060chains of210210residues each \(12,60012\{,\}600residues total\) under*no*point\-group symmetry constraint, with each colour denoting a distinct chain\. Compare with the well\-formed icosahedral nanoparticles of[Figure2](https://arxiv.org/html/2607.05439#S4.F2)f\.

Similar Articles

An Efficient Heterogeneous Co-Design for Fine-Tuning on a Single GPU

Papers with Code Trending

SlideFormer introduces a heterogeneous co-design for full-parameter LLM fine-tuning on a single GPU, leveraging GPU/CPU/RAM/NVMe with a layer-sliding engine and optimized Triton kernels, enabling fine-tuning of 123B+ models on a single RTX 4090 with significant throughput improvements.

Co-folding model guided by structural proteomics

arXiv cs.LG

Introduces AIMS-Fold, an inference-time guided-diffusion framework that integrates cross-linking mass spectrometry (XL-MS) and hydrogen-deuterium exchange (HDX-MS) data to improve protein co-folding predictions for induced proximity drug targets.