Atelier: 通过超网络学习CryoEM体积的局部自监督特征
摘要
Atelier引入了一种自监督超网络框架,该框架为CryoEM密度图生成隐式神经表示,使得局部特征提取成为可能,并提升下游注释任务的性能。
arXiv:2609.30569v1 Announce Type: new
Abstract: CryoEM map interpretation requires features that are spatially localized, consistent across samples, and informative across spatial scales. Most deep learning methods for map annotation extract features from fixed voxel grids. However, implicit neural representations (INRs) are able to model volumetric data as scale-agnostic, coordinate-conditioned functions. INRs are therefore attractive for cryoEM, but fitting a separate INR for each map is too expensive for large-scale feature extraction and produces representations that are not aligned across samples. We introduce Atelier, a self-supervised framework that amortizes INR fitting for reconstructed cryoEM maps. Pretrained on 5,439 Electron Microscopy Data Bank maps, Atelier is a transformer-based hypernetwork that generates high-fidelity reconstructions across a wide range of protein structures, including large multi-subunit assemblies. Beyond reconstruction, the INR generated by the pretrained transformer exposes a continuous, local feature field through its intermediate activations at any spatial query point, a property that voxel grid and patch-tokenizer architectures do not naturally provide. Used as auxiliary channels to a 3D nested U-Net annotation head trained from scratch, these coordinate-conditioned features improve performance on eight voxel-level property prediction tasks over a volume-only baseline. Our results demonstrate that amortized implicit neural representations are an effective primitive for geometry-aware analysis of cryoEM data.
查看缓存全文
缓存时间: 2026/09/28 09:39
# Atelier: Learning Local Self-Supervised Features for CryoEM Volumes via Hypernetworks
Source: [https://arxiv.org/html/2609.30569](https://arxiv.org/html/2609.30569)
###### Abstract
CryoEM map interpretation requires features that are spatially localized, consistent across samples, and informative across spatial scales\. Most deep learning methods for map annotation extract features from fixed voxel grids\. However, implicit neural representations \(INRs\) are able to model volumetric data as scale\-agnostic, coordinate\-conditioned functions\. INRs are therefore attractive for cryoEM, but fitting a separate INR for each map is too expensive for large\-scale feature extraction and produces representations that are not aligned across samples\. We introduce Atelier, a self\-supervised framework that amortizes INR fitting for reconstructed cryoEM maps\. Pretrained on 5,439 Electron Microscopy Data Bank maps, Atelier is a transformer\-based hypernetwork that generates high\-fidelity reconstructions across a wide range of protein structures, including large multi\-subunit assemblies\. Beyond reconstruction, the INR generated by the pretrained transformer exposes a continuous, local feature field through its intermediate activations at any spatial query point, a property that voxel grid and patch\-tokenizer architectures do not naturally provide\. Used as auxiliary channels to a 3D nested U\-Net annotation head trained from scratch, these coordinate\-conditioned features improve performance on eight voxel\-level property prediction tasks over a volume\-only baseline\. Our results demonstrate that amortized implicit neural representations are an effective primitive for geometry\-aware analysis of cryoEM data\.
## 1Introduction
Cryo\-electron microscopy \(cryoEM\) has become a central tool in structural biology\[[1](https://arxiv.org/html/2609.30569#bib.bib1)\], enabling three\-dimensional reconstruction of macromolecular assemblies, often at near\-atomic resolution\. The Electron Microscopy Data Bank \(EMDB\[[2](https://arxiv.org/html/2609.30569#bib.bib44)\]\) now contains more than 56,000 density maps, providing a large resource for studying macromolecular structure and conformational variability\. However, reconstructed cryoEM maps are three\-dimensional scalar fields that do not by themselves specify atomic models, residue identities, secondary structure, or ligand\-binding sites\. Extracting these annotations typically requires substantial computational modeling and expert validation, and remains a bottleneck in converting maps into biological insight\.
Figure 1:Overview\.Atelier is a self\-supervised framework for learning geometry\-aware implicit neural representations of cryoEM density maps via hypernetworks\. The generated representations expose local features at arbitrary spatial resolution, supporting downstream tasks on unseen protein structures\.CryoEM map interpretation relies on features spanning multiple spatial scales\. Local density patterns support annotations such as secondary structure, nucleotide or protein\-region identification, and residue\-level model building, while larger\-scale context is needed to interpret chain connectivity, molecular surfaces, domain organization, and interfaces between subunits\[[3](https://arxiv.org/html/2609.30569#bib.bib62),[4](https://arxiv.org/html/2609.30569#bib.bib61)\]\. Because the downstream\-relevant features are not known in advance, a useful representation should expose information across scales, combining localized geometric detail with broader volumetric context\.
Current deep learning methods for cryoEM map annotation, most notably 3D U\-Nets, capture multi\-scale spatial context and have been used for secondary\-structure and nucleotide annotation\[[5](https://arxiv.org/html/2609.30569#bib.bib18),[6](https://arxiv.org/html/2609.30569#bib.bib15)\]\. However, these architectures operate on fixed voxel grids\. Their feature fields are tied to the input discretization, and predictions at off\-grid coordinates require interpolation rather than direct coordinate\-conditioned evaluation\.
Implicit neural representations \(INRs\) provide a scale\-agnostic alternative to voxel\-bound representations\. They model a map as a coordinate\-conditioned functionℳθ:\[0,1\]3→ℝ\\mathcal\{M\}\_\{\\theta\}:\[0,1\]^\{3\}\\to\\mathbb\{R\}that can be evaluated at arbitrary spatial coordinates\[[7](https://arxiv.org/html/2609.30569#bib.bib7),[8](https://arxiv.org/html/2609.30569#bib.bib10)\], rather than only on a fixed lattice\. This makes them well suited to cryoEM maps, which are discretized samples of an underlying continuous signal\. Their main limitation is scalability: fitting one INR per map requires hundreds to thousands of optimization steps, and independently fitted INRs do not produce aligned intermediate features across maps\. This prevents their direct use as shared feature extractors for downstream annotation\. Hypernetworks that predict INR parameters from input data\[[9](https://arxiv.org/html/2609.30569#bib.bib36),[10](https://arxiv.org/html/2609.30569#bib.bib37),[11](https://arxiv.org/html/2609.30569#bib.bib12),[12](https://arxiv.org/html/2609.30569#bib.bib40)\]amortize this fitting cost, but cryoEM applications have focused primarily on reconstruction from particle images\[[13](https://arxiv.org/html/2609.30569#bib.bib32),[14](https://arxiv.org/html/2609.30569#bib.bib33)\], rather than transferable feature learning from reconstructed maps\.
We introduce Atelier \(Figure[1](https://arxiv.org/html/2609.30569#S1.F1)\), a self\-supervised framework that learns to map cryoEM volumes to INRs via a hypernetwork architecture\. A transformer hypernetwork tokenizes a 3D density map and outputs the parameters of a compact INR in a single forward pass; this framework amortizes INR fitting across thousands of structures\. By amortizing INR fitting into a feed\-forward prediction, the hypernetwork produces instance\-specific neural fields with implicitly aligned representations\. This allows us to extract localized, coordinate\-conditioned features that reside in a shared latent space across volumes\. Crucially, Atelier is designed to*augment*rather than replace existing voxel grid architectures\. Because the generated INR is queryable at any continuous coordinate, its intermediate activations yield feature vectors that we use as auxiliary channels to a strong 3D nested U\-Net baseline\. This decomposition lets practitioners keep the multi\-scale strengths of established discrete architectures while injecting a continuous, geometry\-aware feature signal\.
To validate the framework, we pretrain on54395439EMDB maps and evaluate the augmented U\-Net on eight voxel\-level annotation tasks, using a capacity\-matched transformer\-only autoencoder as a controlled baseline\. The ablation lets us decompose the source of downstream gains into two components: the contribution of self\-supervised transformer pretraining over a no\-pretraining baseline, and the additional contribution of the INR formulation over the transformer alone\.
In summary, our main contributions are:
- •Amortized continuous representations for cryoEM at scale\.We introduce Atelier, a transformer\-hypernetwork framework that generates per\-volume INRs in a single forward pass\. On689689held\-out EMDB volumes, Atelier achieves an average normalized area under the Fourier shell correlation curve of0\.634±0\.0930\.634\\pm 0\.093\.
- •Coordinate\-queryable features for downstream annotation\.We show that the intermediate activations of the generated INR can be queried at arbitrary spatial coordinates and used as auxiliary inputs to a 3D U\-Net annotation head, improving voxel\-level annotation accuracy on eight out of eight tasks against a volume\-only baseline\.
- •Decomposing the source of the improvement\.A capacity\-matched transformer\-only autoencoder ablation shows that self\-supervised transformer pretraining accounts for the majority of the downstream gain, while the INR formulation typically contributes a smaller additional improvement and uniquely supports coordinate\-conditioned, resolution\-agnostic inference that the transformer alone cannot perform\.
## 2Background and Related Work
#### Grid\-Based Representation Learning in CryoEM\.
Cryo\-electron microscopy determines macromolecular structure by reconstructing 3D maps from images of vitrified specimens acquired with an electron beam\[[15](https://arxiv.org/html/2609.30569#bib.bib38)\]\. For favorable specimens, recent advances routinely produce near\-atomic\-resolution maps, including maps at 3 Å or better\[[16](https://arxiv.org/html/2609.30569#bib.bib39)\]\. Many biologically important maps, however, remain at intermediate or low resolution, especially in heterogeneous assemblies and in situ cryo\-electron tomography, where cellular context is preserved at the cost of lower signal\-to\-noise ratio and resolution\[[17](https://arxiv.org/html/2609.30569#bib.bib64)\]\. Map interpretation therefore spans a broad resolution range: high\-resolution maps can support automated atomic model building with tools such as ModelAngelo\[[18](https://arxiv.org/html/2609.30569#bib.bib63)\], while lower\-resolution maps often require annotation, segmentation, fitting, or validation using a combination of automated methods and interactive tools such as Coot\[[19](https://arxiv.org/html/2609.30569#bib.bib43)\]and ChimeraX\[[20](https://arxiv.org/html/2609.30569#bib.bib17)\]\. Automated annotation methods, including secondary\-structure prediction\[[5](https://arxiv.org/html/2609.30569#bib.bib18),[6](https://arxiv.org/html/2609.30569#bib.bib15)\]and backbone tracing\[[21](https://arxiv.org/html/2609.30569#bib.bib19)\], have historically relied on 3D U\-Nets\. While these methods automate important parts of map interpretation, they operate on discretized voxel grids\. Consequently, their representations are tied to the input grid: increasing spatial resolution increases memory and compute cost, and predictions away from grid points require interpolation rather than direct coordinate\-conditioned evaluation\.
#### Implicit Neural Representations\.
Implicit neural representations \(INRs\) are coordinate\-conditioned neural networks that model signals as continuous functions over space, and have been successfully applied to various domains including 2D images\[[22](https://arxiv.org/html/2609.30569#bib.bib8),[23](https://arxiv.org/html/2609.30569#bib.bib9)\], 3D scenes\[[8](https://arxiv.org/html/2609.30569#bib.bib10),[24](https://arxiv.org/html/2609.30569#bib.bib11)\], and audio\[[7](https://arxiv.org/html/2609.30569#bib.bib7)\]\. In practice, INRs are typically parameterized as lightweight multilayer perceptrons \(MLPs\), often augmented with sinusoidal activations\[[7](https://arxiv.org/html/2609.30569#bib.bib7)\]or Fourier feature embeddings\[[25](https://arxiv.org/html/2609.30569#bib.bib30)\]\.
The continuous formulation of INRs makes them particularly well\-suited for cryoEM density maps, where biologically meaningful signals depend on fine\-grained sub\-voxel geometry\. Prior work has explored INRs as a computational primitive for cryoEM, with\[[26](https://arxiv.org/html/2609.30569#bib.bib14)\]demonstrating their effectiveness for representing electron density\. CryoDRGN\[[14](https://arxiv.org/html/2609.30569#bib.bib33)\]and its extension cryoDRGN\-AI\[[27](https://arxiv.org/html/2609.30569#bib.bib34)\]learn latent\-conditioned INRs from heterogeneous particle images, while cryoAI\[[28](https://arxiv.org/html/2609.30569#bib.bib35)\]fits an INR to reproduce observed projections under varying poses\. However, these approaches are designed for recovering cryoEM densities from raw noisy images rather than learning transferable representations from reconstructed volumes\.
Beyond reconstruction, INRs enable coordinate\-conditioned feature extraction: intermediate activations can be queried at arbitrary spatial locations to produce feature vectors encoding local geometry at sub\-voxel resolution\. This perspective has been explored in recent work on neural fields as continuous feature volumes\[[29](https://arxiv.org/html/2609.30569#bib.bib57),[30](https://arxiv.org/html/2609.30569#bib.bib3)\]\. However, standard INRs are trained independently per datum, making them computationally expensive and resulting in representations that are not aligned across instances, limiting their use as scalable feature extractors\.
#### Hypernetworks and Amortized INRs\.
To circumvent per\-instance INR fitting, hypernetworks generate the parameters of a target network conditioned on an input\[[9](https://arxiv.org/html/2609.30569#bib.bib36),[31](https://arxiv.org/html/2609.30569#bib.bib5),[12](https://arxiv.org/html/2609.30569#bib.bib40)\], enabling amortized inference where a single forward pass produces instance\-specific weights\. In the INR setting, prior work such as MetaSDF\[[10](https://arxiv.org/html/2609.30569#bib.bib37)\], Trans\-INR\[[12](https://arxiv.org/html/2609.30569#bib.bib40)\], and HyperSound\[[32](https://arxiv.org/html/2609.30569#bib.bib20)\]has shown that hypernetworks can efficiently generate high\-fidelity implicit neural representations across diverse domains, including shapes, images, and audio\. More recent work demonstrates their applicability to structured scientific data\[[29](https://arxiv.org/html/2609.30569#bib.bib57),[33](https://arxiv.org/html/2609.30569#bib.bib4)\], while CryoHype\[[13](https://arxiv.org/html/2609.30569#bib.bib32)\]applies hypernetworks to model heterogeneity in cryoEM reconstruction\.
Hypernetworks induce a shared representation across instances by mapping inputs to a common function space\. As a result, hypernetwork\-generated INRs yield features that are implicitly aligned across volumes through the shared hypernetwork mapping, enabling localized, coordinate\-conditioned features that reside in a shared latent space\. This property makes neural fields a practical and scalable foundation for geometry\-aware feature extraction across large collections of cryoEM volumes\.
#### Intermediate Activations as Features\.
Extracting feature representations from intermediate network activations has a long history in computer vision, underpinning transfer learning in pretrained CNNs\[[34](https://arxiv.org/html/2609.30569#bib.bib2)\]and modern self\-supervised learning\[[35](https://arxiv.org/html/2609.30569#bib.bib6),[36](https://arxiv.org/html/2609.30569#bib.bib13)\]\. Recent work has observed that intermediate layers of INRs with sinusoidal activations encode localized spatial frequencies\[[7](https://arxiv.org/html/2609.30569#bib.bib7)\]and geometric structure beyond what is required for pure reconstruction\[[11](https://arxiv.org/html/2609.30569#bib.bib12)\]\. A distinctive property of INR activations is*spatial localization*, where querying the network at coordinatexxproduces features intrinsically tied to the local geometry aroundxx, in contrast with voxel encoders which require pooling for localized predictions\. Atelier harnesses this property at scale, using coordinate\-aligned INR activations from an amortized hypernetwork as features for voxel\-level annotation without task\-specific supervision during pretraining\.
## 3Method
Figure 2:Atelier architecture\.\(a\)A cryoEM volumeVVoccupies a\(D,D,D\)\(D,D,D\)\-shaped bounding box, split into non\-overlapping patches of size\(P,P,P\)\(P,P,P\)\. Since the\(D,D,D\)\(D,D,D\)volume is naturally sparse with most voxels being background \(set to a value of00\), we collect the active patches containing non\-background voxels into a batch and discard the rest\. These patches are fed through a convolutional neural network that transforms each patch into a volume token\. The volume tokens are passed into a transformer along with learnable weight tokens; 3D rotary positional embeddings inform the volume tokens of the patch that they came from\. The contextualized weight tokens output by the transformer are then projected to the input volume\-conditioned weight matricesθV\\theta\_\{V\}of an INR decoderℳ\\mathcal\{M\}, which can be queried at arbitrary spatial points in\[0,1\]3\[0,1\]^\{3\}to produce a reconstructed volumeV^\\widehat\{V\}\. Density map of EMDB\-0406\[[37](https://arxiv.org/html/2609.30569#bib.bib22)\]visualized with ChimeraX\.\(b\)The INR naturally exposes an interface for fine\-grained, coordinate\-conditioned feature extraction\. At inference, the transformerℋφ\\mathcal\{H\}\_\{\\varphi\}produces weight matricesθV\\theta\_\{V\}that parameterize the INRℳθV\\mathcal\{M\}\_\{\\theta\_\{V\}\}\. For any query coordinatex∈\[0,1\]3x\\in\[0,1\]^\{3\}, we extract the post\-activation hidden states from each of theKKINR layers \(pink\) and stack them into a query\-conditioned featurefV,x∈ℝH×Kf\_\{V,x\}\\in\\mathbb\{R\}^\{H\\times K\}\. Becausexxis continuous, features can be queried at sub\-voxel resolution and are spatially localized to the geometry aroundxxwithout interpolation\.### 3\.1Data preprocessing
CryoEM volumes are scalar\-valued voxel grids, where the value at each voxel represents electron density\. Measurements are typically noisy, so we first denoise volumes via thresholding and removal of small connected components, then normalizing all voxel values to\[0,1\]\[0,1\]; this sets all background voxels to 0\. We then pad the denoised volumes to a fixed\(D,D,D\)\(D,D,D\)size, whereD=96D=96; any maps whose foreground voxels exceed this size are discarded\. All maps are resampled to have a uniform resolution of 3 Å per voxel using theresamplecommand in ChimeraX\[[38](https://arxiv.org/html/2609.30569#bib.bib16)\]\. We discuss our data curation and processing pipeline in further detail in §[A](https://arxiv.org/html/2609.30569#A1)\.
### 3\.2Self\-supervised INR pretraining
While a cryoEM volumeV∈ℝD×D×DV\\in\\mathbb\{R\}^\{D\\times D\\times D\}is natively a discrete tensor, we view it as a continuous functionV:\[0,1\]3→ℝ\+V:\[0,1\]^\{3\}\\to\\mathbb\{R\}^\{\+\}, establishing a natural correspondence between the integer lattice\[D\]3\[D\]^\{3\}and the continuous coordinate space\[0,1\]3\[0,1\]^\{3\}\. The goal of our self\-supervised pretraining is to learn a coordinate\-based INRℳθV:\[0,1\]3→ℝ\+\\mathcal\{M\}\_\{\\theta\_\{V\}\}:\[0,1\]^\{3\}\\to\\mathbb\{R\}^\{\+\}such thatℳθV≈V\\mathcal\{M\}\_\{\\theta\_\{V\}\}\\approx Vpointwise for each volumeVV\.
A naïve approach would be to separately train an INR for each volumeVV\. However, this precludes any generalization to new volumes and does not allow for synthesizing information across a diversity of protein structures\. Atelier bypasses this by learning a global hypernetworkℋφ\\mathcal\{H\}\_\{\\varphi\}, parameterized byφ\\varphi\. In other words, we learn a mappingℋφ:V↦θV\\mathcal\{H\}\_\{\\varphi\}:V\\mapsto\\theta\_\{V\}that observes a discrete cryoEM volumeVVand directly predicts the parametersθV\\theta\_\{V\}that renderVVvia the INRℳθV\\mathcal\{M\}\_\{\\theta\_\{V\}\}\.111An*atelier*is a workshop where a master artist trains many apprentices to produce work under the master’s name\. Just as the master produces trainees that render artworks, so too does the hypernetwork produce INRs that render cryoEM electron densities\.
The hypernetwork\-to\-INR formulation has been explored in prior work on 2D images and 3D shapes\[[12](https://arxiv.org/html/2609.30569#bib.bib40),[10](https://arxiv.org/html/2609.30569#bib.bib37),[32](https://arxiv.org/html/2609.30569#bib.bib20)\]\. We adapt this paradigm to volumetric cryoEM data, with a training objective designed to recover the fine structural detail relevant to protein geometry\. We sample the full dense grid ofN=D3N=D^\{3\}query pointsΩ:=\{i/\(D−1\):i∈0,…,D−1\}3⊂\[0,1\]3\\Omega\\vcentcolon=\\left\\\{i/\(D\-1\):i\\in 0,\\dots,D\-1\\right\\\}^\{3\}\\subset\[0,1\]^\{3\}and evaluate the INR over the grid to get a predicted volumeV^=ℳθV\(Ω\)\\widehat\{V\}=\\mathcal\{M\}\_\{\\theta\_\{V\}\}\(\\Omega\)\. In order to encourage the model to learn fine details associated with protein structures, our loss is the sum of the mean squared error loss between the volumes and the mean squared error of the 3D discrete orthonormal Fourier transformℱ\\mathcal\{F\}of the volumes, weighted by the square of the normalized frequenciesξ∈\[−1/2,1/2\)3=:ℱ\(Ω\)\\xi\\in\[\-1/2,1/2\)^\{3\}=\\vcentcolon\\mathcal\{F\}\(\\Omega\)in order to upweight high\-frequency components\.
ℒ\(φ\)=𝔼V∼𝒟\[1N∑x∈Ω\|V\(x\)−V^\(x\)\|2\+1N∑ξ∈ℱ\(Ω\)\|ξ\|2\|ℱ\(V\)\(ξ\)−ℱ\(V^\)\(ξ\)\|2\]\.\\mathcal\{L\}\(\\varphi\)=\\mathbb\{E\}\_\{V\\sim\\mathcal\{D\}\}\\left\[\\frac\{1\}\{N\}\\sum\_\{x\\in\\Omega\}\|V\(x\)\-\\widehat\{V\}\(x\)\|^\{2\}\+\\frac\{1\}\{N\}\\sum\_\{\\xi\\in\\mathcal\{F\}\(\\Omega\)\}\|\\xi\|^\{2\}\|\\mathcal\{F\}\(V\)\(\\xi\)\-\\mathcal\{F\}\(\\widehat\{V\}\)\(\\xi\)\|^\{2\}\\right\]\.\(1\)
Note that without the\|ξ\|2\|\\xi\|^\{2\}term, the Fourier loss would be equal to the spatial loss by Parseval’s theorem\. We use𝒟\\mathcal\{D\}above to denote the data distribution\.
### 3\.3The Atelier architecture
The Atelier architecture, illustrated in Figure[2](https://arxiv.org/html/2609.30569#S3.F2)\(a\), consists of three main components: a tokenizer designed to handle high\-dimensional sparse cryoEM volumes, a transformer encoder that generates positionally\-aware volume and weight tokens, and an INR decoder that uses the transformed weight tokens to render volumes\. Here we briefly describe each component, with further details in §[B](https://arxiv.org/html/2609.30569#A2)\.
#### The convolutional tokenizer\.
Since we pad all volumes to the same\(D,D,D\)\(D,D,D\)shape and cryoEM volumes vary widely in size, most volumes will be sparsely populated with nonzero foreground voxels\. To exploit this sparsity, we split the input volumeVVinto non\-overlapping patches of size\(P,P,P\)\(P,P,P\)withP=6P=6, yielding a sequence of\(D/P\)3\(D/P\)^\{3\}patches\. The patches containing only background voxels are discarded, and the remaining active patches containing foreground voxels are processed in a batch by a 3D convolutional neural network into a sequence ofSSlatent volume tokens\.
#### The hypernetwork𝓗𝝋\\boldsymbol\{\\mathcal\{H\}\_\{\\varphi\}\}\.
Learnable weight tokens are fed into a FlashAttention\-enabled transformer\[[39](https://arxiv.org/html/2609.30569#bib.bib26),[40](https://arxiv.org/html/2609.30569#bib.bib27)\]along with the volume tokens output by the tokenizer\. To retain global spatial context, we apply GPT\-NeoX\-style 3D rotary position embeddings \(RoPE\)\[[41](https://arxiv.org/html/2609.30569#bib.bib28),[42](https://arxiv.org/html/2609.30569#bib.bib29)\]to the volume tokens in the transformer\. The final sequence of weight tokens output by the transformer are imbued with spatial context from the volume tokens; we map the transformed weight tokens to weight matrices for an INR via a learnable linear projection\.
#### The INR𝓜𝜽\\boldsymbol\{\\mathcal\{M\}\_\{\\theta\}\}\.
The INR takes the form of an MLP with ReLU activations andKKhidden layers of uniform widthHH\. The weights are generated by the hypernetwork, while the biases are internal and not data\-dependent\. The INR is tasked with predicting the electron density at any pointx∈\[0,1\]3x\\in\[0,1\]^\{3\}for the given volumeVV\. To enable learning high\-frequency functions, the spatial coordinatexxis first mapped into a higher dimensional space with a fixed Fourier encoding\[[8](https://arxiv.org/html/2609.30569#bib.bib10),[25](https://arxiv.org/html/2609.30569#bib.bib30)\]\.
### 3\.4Local Feature Extraction
Typical self\-supervised tasks map input sequences to tokens that are used directly to reconstruct a corrupted sequence, fundamentally limiting downstream reasoning to the resolution of the tokens\. However, since our reconstruction is done with an INRℳθ\\cal\{M\}\_\{\\theta\}that can be queried at any continuous point in\[0,1\]3\[0,1\]^\{3\}, we naturally are able to obtain latent representations of a volumeVVat arbitrary resolution\. Sinceℋφ\\mathcal\{H\}\_\{\\varphi\}is trained to produce highly accurate structural reconstructions, the intermediate activations ofℳθV\\mathcal\{M\}\_\{\\theta\_\{V\}\}must encode rich, geometry\-aware information about the local density neighborhood around any query pointxx, as illustrated in Figure[2](https://arxiv.org/html/2609.30569#S3.F2)\(b\)\. Given any query pointxx, we can obtain a query\-conditioned featurefV,x∈ℝH×Kf\_\{V,x\}\\in\\mathbb\{R\}^\{H\\times K\}by extracting the post\-activation hidden states from the INR evaluated at coordinatexx\. We can also obtain a feature for an arbitrary set of pointsX=\{xi\}i=1NX=\\\{x\_\{i\}\\\}\_\{i=1\}^\{N\}by concatenating the activations at each point to get a region\-conditioned featurefV,X∈ℝN×H×Kf\_\{V,X\}\\in\\mathbb\{R\}^\{N\\times H\\times K\}\.
This highly flexible extraction mechanism yields several critical advantages for structural biology\. First, the featurefV,xf\_\{V,x\}is resolution\-independent, meaning it can be queried at sub\-voxel resolution regardless of the voxel spacing of the original mapVV\. Second, it is inherently geometrically localized, as the spatial coordinates directly condition the activation cascade without requiring interpolation heuristics\.
## 4Experiments
### 4\.1Pretraining
We pretrain Atelier on a library of 5439 cryoEM volumes, with validation and test sets of size 684 and 689 respectively, following the method described in §[3\.2](https://arxiv.org/html/2609.30569#S3.SS2)\. Our volumes are derived from the Cryo2StructData dataset\[[43](https://arxiv.org/html/2609.30569#bib.bib41)\], which consists only of cryoEM volumes with resolved atomic structure \(i\.e\., with corresponding fitted PDB maps\)\. Following\[[44](https://arxiv.org/html/2609.30569#bib.bib31),[13](https://arxiv.org/html/2609.30569#bib.bib32)\], we evaluate the quality of our reconstruction by computing the area under the Fourier shell correlation curve \(AUFSC\) between ground truthVVand predictionV^\\widehat\{V\}\. We also report thenormalizedAUFSC, which is always between 0 and 1 \(1\.0 indicating perfect reconstruction\) and defined by us as
nAUFSC\(V,V^\):=AUFSC\(V,V^\)AUFSC\(V,V\)\.\\nAUFSC\(V,\\widehat\{V\}\)\\vcentcolon=\\frac\{\\AUFSC\(V,\\widehat\{V\}\)\}\{\\AUFSC\(V,V\)\}\.\(2\)
On 689 test volumes, we are able to obtain a mean±\\pmstd nAUFSC of0\.634±0\.0930\.634\\pm 0\.093\(AUFSC:0\.180±0\.0260\.180\\pm 0\.026\)\. We show qualitative and quantitative results of reconstructed test volumes in Figure[3](https://arxiv.org/html/2609.30569#S4.F3)\. In particular, as one of the examples, we show Atelier’s rendering of a fully assembled T\-cell receptor \(TCR\), CD3, and peptide\-MHC complex \(EMDB\-28571\[[45](https://arxiv.org/html/2609.30569#bib.bib21)\]\)\. Many of the most biologically important macromolecules are large, multi\-chain assemblies whose function depends on the relative arrangement of distinct subunits, often spanning both soluble and membrane\-embedded regions\. Faithfully reconstructing such complexes is a demanding test for a volumetric model, since the network must simultaneously recover distant subunits, asymmetric chain organization, and the contrast change between protein and detergent or lipid density\. For EMDB\-28571, Atelier recovers both the membrane\-distal TCRαβ\\alpha\\beta/pMHC recognition interface and the membrane\-proximal CD3 signaling subunits in a single forward pass, despite the complex’s asymmetric multi\-chain architecture and embedded transmembrane region\. The achieved nAUFSC is 0\.621, comparable to the test set mean, demonstrating that reconstruction fidelity holds on large, heterogenous assemblies\.
Figure 3:Atelier successfully generalizes to unseen volumes\.Here we show ground truth \(orange\) and INR\-rendered \(blue\) volumes for four representative samples in the test set, as well as the Fourier shell correlation curves between them\. We also show close ups for the predicted densities overlaid over ground truth atomic models to demonstrate that Atelier is able to reconstruct near atomic\-level detail\. Note the biological diversity of structures; from left to right, the structures are a membrane\-bound enzyme \(EMDB\-25368\[[46](https://arxiv.org/html/2609.30569#bib.bib23)\]\), a nanobody bound to an HIV envelope glycoprotein \(EMDB\-23480\[[47](https://arxiv.org/html/2609.30569#bib.bib24)\]\), a nucleosome–kinetochore complex \(EMDB\-7293\[[48](https://arxiv.org/html/2609.30569#bib.bib25)\]\), and a fully assembled TCR/CD3/pMHC immune\-recognition complex \(EMDB\-28571\[[45](https://arxiv.org/html/2609.30569#bib.bib21)\]\)\.We perform two experiments to demonstrate that our pretraining results in geometrically meaningful local features via the extraction method described in §[3\.4](https://arxiv.org/html/2609.30569#S3.SS4)\. During training, we augment all volumes by a random element ing∈Ohg\\in O\_\{h\}, the group of all 48 symmetries of the cube \(generated by90∘90^\{\\circ\}axis\-aligned rotations and reflections\)\. To demonstrate invariance of our local features, we sample a volumeV∗V^\{\*\}and a foreground pointx∈\[0,1\]x\\in\[0,1\], which we use to compute a localized featurefV∗,xf\_\{V^\{\*\},x\}from the pretrained model\. We then compute the cosine similarity betweenfV∗,xf\_\{V^\{\*\},x\}andfg\(V∗\),g\(x\)f\_\{g\(V^\{\*\}\),g\(x\)\}for allg∈Ohg\\in O\_\{h\}; we expect the features to be aligned if the model has learned to generate local representations that are invariant to global symmetries\. As negative controls, we also compute, for allg∈Ohg\\in O\_\{h\}, the cosine similarities betweenfV∗,xf\_\{V^\{\*\},x\}andfg\(V∗\),g\(y\)f\_\{g\(V^\{\*\}\),g\(y\)\}for a decoy foreground pointy≠xy\\neq x, as well as betweenfV∗,xf\_\{V^\{\*\},x\}andfg\(V†\),g\(x\)f\_\{g\(V^\{\\dagger\}\),g\(x\)\}for decoy test volumeV†≠V∗V^\{\\dagger\}\\neq V^\{\*\}\. We visualize our results in Figure[4](https://arxiv.org/html/2609.30569#S4.F4)\(a\) and observe that our model has learned approximately invariant local representations; we argue that this is a natural consequence of the pretraining task, which forces the model to learn geometrically informed features that are locally meaningful\.
We also visualize the actual activations for two test proteins in Figure[4](https://arxiv.org/html/2609.30569#S4.F4)\(b\) by computing features at all foreground voxels, projecting down to three dimensions and normalizing to\[0,1\]3\[0,1\]^\{3\}\(which we interpret as RGB space\) via principal components analysis, coloring each atom in the corresponding atomic model by trilinear interpolation of the RGB values on the voxel grid, and visualizing in PyMol\[[49](https://arxiv.org/html/2609.30569#bib.bib58)\]\. We observe that the feature field is, in many cases, approximately constant within domains while varying between them, suggesting that our pretrained fine features carry some biologically useful signal and can partially segment novel proteins\.
Figure 4:Atelier generates meaningful local continuous features\.\(a\)After pretraining, Atelier is able to generate local features invariant to global symmetries due to its geometrically\-driven pretraining task in two structures from Figure[3](https://arxiv.org/html/2609.30569#S4.F3)\. Box denotes minimum, quartiles, and maximum\.\(b\)PyMol visualization of the same two structures evaluated in part \(a\), with atoms colored by trilinearly interpolating voxel features projected and scaled to\[0,1\]3\[0,1\]^\{3\}via PCA\. Observe in the nucleosome \(EMDB\-7293\) that the histones \(center, green\) primarily occupy a different color space than the surrounding DNA \(purple\) and centromere protein \(upper left, teal\)\. In the TCR/CD3/pMHC complex \(EMDB\-28571\), there is considerably less variation within chains than across chains\. Note, for example, the consistent coloring of the transmembrane region \(bottom, purple\)\.
### 4\.2Demonstrating the Locality of Atelier Fine Features
To demonstrate the locality of the features that Atelier is able to extract from cryo\-EM volumes, we performed the following study on the effect of local versus distant corruptions to a protein on Atelier activations\. We collected 16 held\-out test structures and independently sampled 16 foreground voxel coordinates each; the foreground voxels within a single structure are at least 12 voxels \(= 36 Å\) apart\. For each of the 16 volumes, we generate 16 different corruptions where we zero out all voxels within anR=6R=6voxel radius, centered at one of the foreground pointsxx, and measure the effect of this deletion on the post\-ReLU activations at each layer of the INR generated by the trained Atelier model\. As a matched control, we repeat the deletion at a second foreground pointyyin the same structure, chosen from the remaining 15 points as the one whose enclosed mass most closely matches that of the ball atxx, subject to the two balls not overlapping; that the total mass in the ball around our chosenyyis always within 25% of the mass within the ball centered atxx\. Our hypothesis is that deleting the far neighborhood centered atyyhas much less of an effect than deleting the local neighborhood atxx\. Note that the points within a volume are spaced far enough so that deleting a neighborhood aroundy≠xy\\neq xwill not delete the pointxxitself\.
Our measure of how much deletions change the activations is given by:
Eℓ\(x,y,R\)=‖aℓ\(x,VR,y\)−aℓ\(x,V\)‖‖aℓ\(x,∅\)−aℓ\(x,V\)‖,E\_\{\\ell\}\(x;y,R\)=\\frac\{\\\|a\_\{\\ell\}\(x,V\_\{R,y\}\)\-a\_\{\\ell\}\(x,V\)\\\|\}\{\\\|a\_\{\\ell\}\(x,\\varnothing\)\-a\_\{\\ell\}\(x,V\)\\\|\},\(3\)whereaℓ\(x,V\)a\_\{\\ell\}\(x,V\)is the layer\-ℓ\\ellpost\-ReLU activation of an INR generated from a cryo\-EM volumeVV, andVR,yV\_\{R,y\}is the volumeVVwith a ball of radiusRRcentered atyyzeroed out\. We use∅\\varnothingto denote the all\-zeros volume\. The numerator is therefore the difference in the activations between the deleted and original volume, while the denominator normalizes so that deleting the entire volume would have score 1; a larger score means a stronger perturbation in thexx\-localized features from deleting pointyy\.
For each volume and each hidden layer, we compute the ratio of the average \(over all 16 points\)nearperturbationEℓ\(x,x,R\)E\_\{\\ell\}\(x;x,R\)and the averagefarperturbationEℓ\(x,y,R\)E\_\{\\ell\}\(x;y,R\)\. We report the median value over all 16 volumes at all layers in Table[1](https://arxiv.org/html/2609.30569#S4.T1)\. We also report the same ratio for anuntrainedinstantiation of our model\. As seen in Table[1](https://arxiv.org/html/2609.30569#S4.T1), local perturbations have between an18×18\\timesand34×34\\timeslarger effect than distant ones for the trained model\. Furthermore, for the untrained instantiation of Atelier, local and distant perturbations have nearly the same effect magnitude, indicating that this locality islearnedrather than an inherent property of the architecture\.
Table 1:Per\-layer median near/far average corruption effect \(see Equation \([3](https://arxiv.org/html/2609.30569#S4.E3)\)\) over 16 holdout volumes \(95% bootstrap confidence interval\) for trained and untrained Atelier model\. For the trained model, perturbations local to a point change intermediate activations significantly more than distant perturbations, while for the untrained model, local and distant perturbations have nearly identical effect\. This indicates that Atelier activations capture local information, and that this locality is learned\.
### 4\.3Using Local Features for Downstream Property Prediction
To demonstrate the usefulness of the local features described in §[3\.4](https://arxiv.org/html/2609.30569#S3.SS4), we perform an experiment where we adapt EMNUSS\[[6](https://arxiv.org/html/2609.30569#bib.bib15)\]to predict various local properties for cryoEM volumes and augment the input to EMNUSS with our query\-conditioned features\. EMNUSS is a 3D nested U\-Net\[[50](https://arxiv.org/html/2609.30569#bib.bib42)\]architecture for secondary structure prediction that takes a\(D,D,D\)\(D,D,D\)\-shaped 3D cryoEM volume as input and was originally designed to densely predict a\(D,D,D,3\)\(D,D,D,3\)\-shaped prediction of whether each voxel containing a backbone atom is part of a coil, helix, or strand \(voxels not containing backbone atoms do not contribute to the loss\)\.
Since our data consists only of cryoEM volumes with resolved atomic structure, we are able to extract eight different types of voxel\-localized labels for our volumes, four of which are classification tasks \(secondary structure, amino acid, molecule type, molecule class\) and four of which are regression tasks \(hydrophobicity, absolute SASA, relative SASA, crystallographic B\-factor\); we describe these labels in more detail in §[A\.2](https://arxiv.org/html/2609.30569#A1.SS2)\. We adopt the exact same architecture as EMNUSS, only changing the number of output channels depending on the task\. Classification tasks are trained with standard cross\-entropy loss and regression tasks are trained with mean squared error\.
For each voxel containing an alpha\-carbon, nucleotide glycosidic carbon, or other heavy non\-water heteroatom, we extract two types of embeddings\. The first type of embedding we extract for voxels with defined annotations is a fine\-grained localized embedding obtained by sampling all voxels contained in the sphere of radius 3 voxels centered at the annotated voxel, extracting their corresponding hidden activations as described in §[3\.4](https://arxiv.org/html/2609.30569#S3.SS4), mean pooling over theKKintermediate activations, and performing a Gaussian\-weighted mean pooling over the voxels in the sphere so that features corresponding to voxels further away from the center of the sphere contribute less to the finalHH\-dimensional feature; we pool activations from the sphere instead of only the single feature at the voxel itself in order to incorporate some local context\. As these features are extracted from an INR trained to render the volume at fine detail, we expect them to encode highly localized information useful for biological annotation\. The features are extracted from a frozen Atelier after pretraining\.
To demonstrate the importance of these fine\-grained local features from the INR, we extract a second type of embedding; these are coarse, patch\-level embeddings for each annotated voxel for comparison\. To do this, we retrain an autoencoder to perform the self\-supervised volume reconstruction task without the INR\. This transformer\-only autoencoder uses the same volume tokenizer, but now has a transformer that only takes in volume tokens \(without any weight tokens\); the transformed volume tokens are projected to directly reconstruct the input volume at the corresponding patch; see §[B\.3](https://arxiv.org/html/2609.30569#A2.SS3)for details\. For each annotated voxel, we can associate with it theHH\-dimensional transformed volume token corresponding to the patch that the voxel lives in as a coarse representation at the patch level\. Since all annotated voxels within a patch will have identical representations, this is a much less fine\-grained representation\. For fair comparison, the latent dimensionHHof the transformer\-only autoencoder is the same as the hidden layer sizeHHof the INR in Atelier, and once again our features are extracted from the frozen transformer\-only autoencoder after pretraining\.
For each task, we retrain EMNUSS from scratch with four different types of inputs\.222Molecule type and molecule class are two evaluations of the same model; see §[A\.2](https://arxiv.org/html/2609.30569#A1.SS2)\.As a baseline, we input only the\(D,D,D,1\)\(D,D,D,1\)volume into EMNUSS\. To illustrate the efficacy of our INR\-derived local features, we also retrain with the input augmented to a\(D,D,D,1\+H\)\(D,D,D,1\+H\)multi\-channel volume, withHHrepresenting either the Atelier INR\-derived local feature or the transformer\-only autoencoder’s coarse feature\. These augmentations are applied to voxels with annotations; for voxels without annotations, we set the extraHHchannels to zero\. Since these augmented features, appended mostly at alpha\-carbons, may provide a signal as to the shape of the protein backbone, as an additional control we also run an experiment where we appendrandomfeatures to each annotated voxel\. The random features have the same shape as the fine features and are generated by computing the global per\-channel mean/std of the re\-normalized frozen fine features over all voxels in all volumes, and attaching a random vector to each annotated voxel following this distribution; this guarantees that the first two moments of the random features match those of the Atelier fine features\. Note that other than the dimension of the input channel, we use the exact same training setup for all three types of inputs to ensure a fair comparison\.
We use the same train/validation/test splits as in the pretraining task, and report test set balanced accuracy \(arithmetic mean of per\-class recall\) or relative error for each of the three inputs in Table[2](https://arxiv.org/html/2609.30569#S4.T2)\. We also report the Matthews Correlation Coefficient \(MCC\)\[[51](https://arxiv.org/html/2609.30569#bib.bib59)\], another robust metric for imbalanced multiclass classification, in Table[4](https://arxiv.org/html/2609.30569#A2.T4)in the Appendix\. We observe that augmenting with our local features from the INR pretraining outperforms the augmentation with coarse patch\-level features in seven out of eight tasks, and outperforms the volume\-only vanilla EMNUSS baseline in eight out of eight tasks\. We attribute this to the fact that the pretraining task forces the learning of rich representations capturing local geometry, which improves downstream localized property prediction\. Our results suggest that the INR activations from Atelier contain additional structural information orthogonal to that which is contained in the cryoEM density alone\. We emphasize once more that we are augmenting with*frozen*features after pretraining, indicating that purely geometric information is able to improve performance on biological annotation\. Furthermore, our transformer\-only autoencoder achieves an average per\-volume nAUFSC of0\.870±0\.0610\.870\\pm 0\.061\(AUFSC:0\.248±0\.0170\.248\\pm 0\.017\); despite outperforming Atelier on reconstruction, it typically underperforms on property prediction, highlighting the importance of the fine features extracted by the INR in Atelier\.
Table 2:Comparison on property prediction performance on the test set for \(a\) vanilla EMNUSS that takes the single\-channel volume alone as input, \(b\) EMNUSS with inputs augmented with coarse features as additional channels, and \(c\) EMNUSS with inputs augmented with fine features as extra channels\. Metrics are averaged over all labeled voxels in the test set; a more detailed, per\-volume statistical analysis is reported in §[B\.4](https://arxiv.org/html/2609.30569#A2.SS4)\. Augmentation with fine features outperforms the volume\-alone model in eight out of eight tasks, and outperforms the coarse model in seven out of eight tasks\.
### 4\.4Case Study: Comparing Atelier Features Across Related Structures
To investigate if the hypernetwork nature of Atelier is able to learn representations that are consistent across different volumes, we ran the following case study on the following quartet of four G\-protein coupled receptors \(GPCRs\), all of which were not present in the training data, to examine if Atelier features are conserved across evolutionarily related structures:
1. 1\.EMDB\-62586\[[52](https://arxiv.org/html/2609.30569#bib.bib65)\]: we call this thebasestructure, a mouse TLQP21 bound to mouse C3aR in complex with Go;
2. 2\.EMDB\-62654\[[53](https://arxiv.org/html/2609.30569#bib.bib66)\]: this is the exact same structure just with a different bound peptide, EP67 bound mouse C3aR in complex with Go; we call this thesamestructure \(only the peptide differs\);
3. 3\.EMDB\-62651\[[53](https://arxiv.org/html/2609.30569#bib.bib66)\]: human structure with homologous protein sequences, structure of EP67 bound human C3aR in complex with Go; we call this thehomologousstructure;
4. 4\.EMDB\-64912\[[54](https://arxiv.org/html/2609.30569#bib.bib67)\]: a human Neurotensin Receptor 1 \(hNTSR1\)\-Gi1 complex in nucleotide\-free NC state 3; the NTSR1 and C3aR receptors are both G protein coupled receptors, but are more distantly related than human and mouse C3aR; we call this thedifferentstructure\.
We aligned all four complexes \(pictured in Figure[5](https://arxiv.org/html/2609.30569#S4.F5), resampled to 3 Å voxels, and selected 80 coordinates that are foreground \(nonzero\) in all four structures simultaneously and mutually at least 4 voxels \(12 Å\) apart\. At each point we computed a localized feature exactly as in our EMNUSS experiments, using the same pretrained Atelier model, and then subtracted each volume’s mean feature as a control for volume identity\. In Table[3](https://arxiv.org/html/2609.30569#S4.T3), we report the cosine similarity at corresponding points between the base and same structures, base and homologous structures, and base and different structures\. We observe that the cosine similarity decreases as structural relatedness increases, with the closely related structures \(base, same, and homologous\) clearly separated from the distantly related receptor \(different\)\.
Figure 5:Illustration of the family of four complexes compared in Table[3](https://arxiv.org/html/2609.30569#S4.T3)\.Table 3:Cosine similarities between Atelier fine features at corresponding coordinates between four different structures \(mean±\\pmstd\. over 80 sampled points\)\. The results suggest that at corresponding locations, the activations for more similar structures are more similar than those of distantly related structures, even after spatial alignment\. As negative controls, we also computed the cosine similarities between the 80 points in the base structure and the same 80 points in a completely unrelated decoy volume \(EMDB\-28571, a TCR/CD3 complex, cosine similarities 0\.450±\\pm0\.417, as well as computing the cosine similarity between a base point and one mismatched point each from the same/homologous/different structures \(0\.021±\\pm0\.441\)\.
## 5Conclusion
By training a transformer hypernetwork to generate implicit neural representations of high\-resolution cryoEM volumes, Atelier enables the extraction of geometrically\-rich, localized feature vectors that can be queried at any spatial coordinate\. We demonstrate the efficacy of these features by augmenting a baseline secondary structure prediction model with our coordinate\-conditioned activations, showing consistent improvements on eight out of eight voxel\-level property prediction tasks\. We note several boundaries of our current evaluation that offer clear directions for future work in §[C](https://arxiv.org/html/2609.30569#A3)\.
Beyond the tasks studied here, Atelier’s localized features are naturally suited to annotating functional interfaces in large protein complexes\. In assemblies like the immune synapse, distinguishing the precise contacts between TCR, peptide, and MHC subunits requires sub\-voxel, side\-chain\-level precision\. Discrete voxel grids risk losing fine geometric signal at this scale\. Atelier’s continuous features provide a strong foundation for future work in interface prediction, binding\-site identification, structure\-guided protein engineering, and the next generation of learned representations of cryoEM density maps\.
## References
- \[1\]S\. Subramaniam\(2019\)The cryo\-EM revolution: fueling the next phase\.IUCrJ6\(Pt 1\),pp\. 1–2\(en\)\.Cited by:[§1](https://arxiv.org/html/2609.30569#S1.p1.1)\.
- \[2\]wwPDB Consortium\(2024\)EMDB\-the electron microscopy data bank\.Nucleic Acids Res\.52\(D1\),pp\. D456–D465\(en\)\.Cited by:[Appendix A](https://arxiv.org/html/2609.30569#A1.p1.1),[§1](https://arxiv.org/html/2609.30569#S1.p1.1)\.
- \[3\]N\. Giri, L\. Wang, and J\. Cheng\(2024\)Cryo2structdata: a large labeled cryo\-em density map dataset for ai\-based modeling of protein structures\.Scientific Data11\(1\),pp\. 458\.Cited by:[§1](https://arxiv.org/html/2609.30569#S1.p2.1)\.
- \[4\]K\. Jamali, D\. Kimanius, and S\. H\. Scheres\(2022\)A graph neural network approach to automated model building in cryo\-em maps\.arXiv preprint arXiv:2210\.00006\.Cited by:[§1](https://arxiv.org/html/2609.30569#S1.p2.1)\.
- \[5\]P\. Mostosi, H\. Schindelin, P\. Kollmannsberger, and A\. Thorn\(2020\)Haruspex: a neural network for the automatic identification of oligonucleotides and protein secondary structure in cryo\-electron microscopy maps\.Angew\. Chem\. Int\. Ed Engl\.59\(35\),pp\. 14788–14795\(en\)\.Cited by:[§1](https://arxiv.org/html/2609.30569#S1.p3.1),[§2](https://arxiv.org/html/2609.30569#S2.SS0.SSS0.Px1.p1.1)\.
- \[6\]J\. He and S\. Huang\(2021\)EMNUSS: a deep learning framework for secondary structure annotation in cryo\-EM maps\.Brief\. Bioinform\.22\(6\) \(en\)\.Cited by:[§B\.4](https://arxiv.org/html/2609.30569#A2.SS4.SSS0.Px1.p1.1),[Appendix C](https://arxiv.org/html/2609.30569#A3.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.30569#S1.p3.1),[§2](https://arxiv.org/html/2609.30569#S2.SS0.SSS0.Px1.p1.1),[§4\.3](https://arxiv.org/html/2609.30569#S4.SS3.p1.1)\.
- \[7\]V\. Sitzmann, J\. N\. P\. Martel, A\. W\. Bergman, D\. B\. Lindell, and G\. Wetzstein\(2020\)Implicit neural representations with periodic activation functions\.InProceedings of the 34th International Conference on Neural Information Processing Systems,NeurIPS ’20,Red Hook, NY, USA\.External Links:ISBN 9781713829546Cited by:[§1](https://arxiv.org/html/2609.30569#S1.p4.1),[§2](https://arxiv.org/html/2609.30569#S2.SS0.SSS0.Px2.p1.1),[§2](https://arxiv.org/html/2609.30569#S2.SS0.SSS0.Px4.p1.1)\.
- \[8\]B\. Mildenhall, P\. P\. Srinivasan, M\. Tancik, J\. T\. Barron, R\. Ramamoorthi, and R\. Ng\(2020\)NeRF: representing scenes as neural radiance fields for view synthesis\.InECCV,Cited by:[§1](https://arxiv.org/html/2609.30569#S1.p4.1),[§2](https://arxiv.org/html/2609.30569#S2.SS0.SSS0.Px2.p1.1),[§3\.3](https://arxiv.org/html/2609.30569#S3.SS3.SSS0.Px3.p1.1)\.
- \[9\]D\. Ha, A\. M\. Dai, and Q\. V\. Le\(2017\)HyperNetworks\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=rkpACe1lx)Cited by:[§1](https://arxiv.org/html/2609.30569#S1.p4.1),[§2](https://arxiv.org/html/2609.30569#S2.SS0.SSS0.Px3.p1.1)\.
- \[10\]V\. Sitzmann, E\. R\. Chan, R\. Tucker, N\. Snavely, and G\. Wetzstein\(2020\)MetaSDF: meta\-learning signed distance functions\.InProceedings of the 34th International Conference on Neural Information Processing Systems,NeurIPS ’20,Red Hook, NY, USA\.External Links:ISBN 9781713829546Cited by:[§1](https://arxiv.org/html/2609.30569#S1.p4.1),[§2](https://arxiv.org/html/2609.30569#S2.SS0.SSS0.Px3.p1.1),[§3\.2](https://arxiv.org/html/2609.30569#S3.SS2.p3.1)\.
- \[11\]E\. Dupont, H\. Kim, S\. M\. A\. Eslami, D\. J\. Rezende, and D\. Rosenbaum\(2022\)From data to functa: your data point is a function and you can treat it like one\.InInternational Conference on Machine Learning, ICML 2022, 17\-23 July 2022, Baltimore, Maryland, USA,K\. Chaudhuri, S\. Jegelka, L\. Song, C\. Szepesvári, G\. Niu, and S\. Sabato \(Eds\.\),Proceedings of Machine Learning Research,pp\. 5694–5725\.External Links:[Link](https://proceedings.mlr.press/v162/dupont22a.html)Cited by:[§1](https://arxiv.org/html/2609.30569#S1.p4.1),[§2](https://arxiv.org/html/2609.30569#S2.SS0.SSS0.Px4.p1.1)\.
- \[12\]Y\. Chen and X\. Wang\(2022\)Transformers as meta\-learners for implicit neural representations\.InEuropean Conference on Computer Vision,Cited by:[§B\.1](https://arxiv.org/html/2609.30569#A2.SS1.p1.1),[§1](https://arxiv.org/html/2609.30569#S1.p4.1),[§2](https://arxiv.org/html/2609.30569#S2.SS0.SSS0.Px3.p1.1),[§3\.2](https://arxiv.org/html/2609.30569#S3.SS2.p3.1)\.
- \[13\]J\. Gu, M\. Jeon, A\. Ma, S\. Yeung\-Levy, and E\. D\. Zhong\(2025\)CryoHype: reconstructing a thousand cryo\-em structures with transformer\-based hypernetworks\.External Links:arXiv:2512\.06332Cited by:[§1](https://arxiv.org/html/2609.30569#S1.p4.1),[§2](https://arxiv.org/html/2609.30569#S2.SS0.SSS0.Px3.p1.1),[§4\.1](https://arxiv.org/html/2609.30569#S4.SS1.p1.1)\.
- \[14\]E\. D\. Zhong, T\. Bepler, B\. Berger, and J\. H\. Davis\(2021\)CryoDRGN: reconstruction of heterogeneous cryo\-EM structures using neural networks\.Nat\. Methods18\(2\),pp\. 176–185\(en\)\.Cited by:[§1](https://arxiv.org/html/2609.30569#S1.p4.1),[§2](https://arxiv.org/html/2609.30569#S2.SS0.SSS0.Px2.p2.1)\.
- \[15\]E\. Nogales\(2016\)The development of cryo\-EM into a mainstream structural biology technique\.Nat\. Methods13\(1\),pp\. 24–27\(en\)\.Cited by:[§2](https://arxiv.org/html/2609.30569#S2.SS0.SSS0.Px1.p1.1)\.
- \[16\]C\. L\. Lawson, A\. Kryshtafovych, G\. D\. Pintilie, S\. K\. Burley, J\. Černý, V\. B\. Chen, P\. Emsley, A\. Gobbi, A\. Joachimiak, S\. Noreng, M\. G\. Prisant, R\. J\. Read, J\. S\. Richardson, A\. L\. Rohou, B\. Schneider, B\. D\. Sellers, C\. Shao, E\. Sourial, C\. I\. Williams, C\. J\. Williams, Y\. Yang, V\. Abbaraju, P\. V\. Afonine, M\. L\. Baker, P\. S\. Bond, T\. L\. Blundell, T\. Burnley, A\. Campbell, R\. Cao, J\. Cheng, G\. Chojnowski, K\. D\. Cowtan, F\. DiMaio, R\. Esmaeeli, N\. Giri, H\. Grubmüller, S\. W\. Hoh, J\. Hou, C\. F\. Hryc, C\. Hunte, M\. Igaev, A\. P\. Joseph, W\. Kao, D\. Kihara, D\. Kumar, L\. Lang, S\. Lin, S\. R\. Maddhuri Venkata Subramaniya, S\. Mittal, A\. Mondal, N\. W\. Moriarty, A\. Muenks, G\. N\. Murshudov, R\. A\. Nicholls, M\. Olek, C\. M\. Palmer, A\. Perez, E\. Pohjolainen, K\. R\. Pothula, C\. N\. Rowley, D\. Sarkar, L\. U\. Schäfer, C\. J\. Schlicksup, G\. F\. Schröder, M\. Shekhar, D\. Si, A\. Singharoy, O\. V\. Sobolev, G\. Terashi, A\. C\. Vaiana, S\. C\. Vedithi, J\. Verburgt, X\. Wang, R\. Warshamanage, M\. D\. Winn, S\. Weyand, K\. Yamashita, M\. Zhao, M\. F\. Schmid, H\. M\. Berman, and W\. Chiu\(2024\)Outcomes of the EMDataResource cryo\-EM ligand modeling challenge\.Nat\. Methods21\(7\),pp\. 1340–1348\(en\)\.Cited by:[§2](https://arxiv.org/html/2609.30569#S2.SS0.SSS0.Px1.p1.1)\.
- \[17\]C\. Berger, N\. Premaraj, R\. B\. Ravelli, K\. Knoops, C\. López\-Iglesias, and P\. J\. Peters\(2023\)Cryo\-electron tomography on focused ion beam lamellae transforms structural cell biology\.Nature Methods20\(4\),pp\. 499–511\.Cited by:[§2](https://arxiv.org/html/2609.30569#S2.SS0.SSS0.Px1.p1.1)\.
- \[18\]K\. Jamali, L\. Käll, R\. Zhang, A\. Brown, D\. Kimanius, and S\. H\. Scheres\(2024\)Automated model building and protein identification in cryo\-em maps\.Nature628\(8007\),pp\. 450–457\.Cited by:[§2](https://arxiv.org/html/2609.30569#S2.SS0.SSS0.Px1.p1.1)\.
- \[19\]A\. Casañal, B\. Lohkamp, and P\. Emsley\(2020\)Current developments in coot for macromolecular model building of electron cryo\-microscopy and crystallographic data\.Protein Sci\.29\(4\),pp\. 1069–1078\(en\)\.Cited by:[§2](https://arxiv.org/html/2609.30569#S2.SS0.SSS0.Px1.p1.1)\.
- \[20\]E\. C\. Meng, T\. D\. Goddard, E\. F\. Pettersen, G\. S\. Couch, Z\. J\. Pearson, J\. H\. Morris, and T\. E\. Ferrin\(2023\)UCSF ChimeraX: tools for structure building and analysis\.Protein Sci\.32\(11\),pp\. e4792\(en\)\.Cited by:[§2](https://arxiv.org/html/2609.30569#S2.SS0.SSS0.Px1.p1.1)\.
- \[21\]J\. Pfab, N\. M\. Phan, and D\. Si\(2021\)DeepTracer for fast de novo cryo\-EM protein structure modeling and special studies on CoV\-related complexes\.Proc\. Natl\. Acad\. Sci\. U\. S\. A\.118\(2\),pp\. e2017525118\(en\)\.Cited by:[§2](https://arxiv.org/html/2609.30569#S2.SS0.SSS0.Px1.p1.1)\.
- \[22\]I\. Skorokhodov, S\. Ignatyev, and M\. Elhoseiny\(2021\)Adversarial generation of continuous images\.In2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\)2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),Cited by:[§2](https://arxiv.org/html/2609.30569#S2.SS0.SSS0.Px2.p1.1)\.
- \[23\]Y\. Chen, S\. Liu, and X\. Wang\(2021\)Learning continuous image representation with local implicit image function\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 8628–8638\.Cited by:[§2](https://arxiv.org/html/2609.30569#S2.SS0.SSS0.Px2.p1.1)\.
- \[24\]V\. Sitzmann, M\. Zollhöfer, and G\. Wetzstein\(2019\)Scene representation networks: continuous 3d\-structure\-aware neural scene representations\.InAdvances in Neural Information Processing Systems,Cited by:[§2](https://arxiv.org/html/2609.30569#S2.SS0.SSS0.Px2.p1.1)\.
- \[25\]M\. Tancik, P\. Srinivasan, B\. Mildenhall, S\. Fridovich\-Keil, N\. Raghavan, U\. Singhal, R\. Ramamoorthi, J\. Barron, and R\. Ng\(2020\)Fourier features let networks learn high frequency functions in low dimensional domains\.InAdvances in Neural Information Processing Systems,H\. Larochelle, M\. Ranzato, R\. Hadsell, M\.F\. Balcan, and H\. Lin \(Eds\.\),Vol\.33,pp\. 7537–7547\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2020/file/55053683268957697aa39fba6f231c68-Paper.pdf)Cited by:[§2](https://arxiv.org/html/2609.30569#S2.SS0.SSS0.Px2.p1.1),[§3\.3](https://arxiv.org/html/2609.30569#S3.SS3.SSS0.Px3.p1.1)\.
- \[26\]N\. Ranno and D\. Si\(2022\)Neural representations of cryo\-EM maps and a graph\-based interpretation\.BMC Bioinformatics23\(Suppl 3\),pp\. 397\(en\)\.Cited by:[§2](https://arxiv.org/html/2609.30569#S2.SS0.SSS0.Px2.p2.1)\.
- \[27\]A\. Levy, R\. Raghu, J\. R\. Feathers, M\. Grzadkowski, F\. Poitevin, J\. D\. Johnston, F\. Vallese, O\. B\. Clarke, G\. Wetzstein, and E\. D\. Zhong\(2025\)CryoDRGN\-AI: neural ab initio reconstruction of challenging cryo\-EM and cryo\-ET datasets\.Nat\. Methods22\(7\),pp\. 1486–1494\(en\)\.Cited by:[§2](https://arxiv.org/html/2609.30569#S2.SS0.SSS0.Px2.p2.1)\.
- \[28\]A\. Levy, F\. Poitevin, J\. Martel, Y\. Nashed, A\. Peck, N\. Miolane, D\. Ratner, M\. Dunne, and G\. Wetzstein\(2022\)CryoAI: amortized inference of poses for ab initio reconstruction of 3d molecular volumes from real cryo\-em images\.External Links:arXiv:2203\.08138Cited by:[§2](https://arxiv.org/html/2609.30569#S2.SS0.SSS0.Px2.p2.1)\.
- \[29\]S\. Babu, P\. Lo, X\. Zhang, A\. Srivastava, A\. Davariashtiyani, J\. Perera, M\. Maire, and A\. A\. Khan\(2025\)HyperDiffusionFields \(hydif\): diffusion\-guided hypernetworks for learning implicit molecular neural fields\.External Links:arXiv:2510\.18122Cited by:[§2](https://arxiv.org/html/2609.30569#S2.SS0.SSS0.Px2.p3.1),[§2](https://arxiv.org/html/2609.30569#S2.SS0.SSS0.Px3.p1.1)\.
- \[30\]S\. Kobayashi, E\. Matsumoto, and V\. Sitzmann\(2022\)Decomposing nerf for editing via feature field distillation\.Advances in neural information processing systems35,pp\. 23311–23330\.Cited by:[§2](https://arxiv.org/html/2609.30569#S2.SS0.SSS0.Px2.p3.1)\.
- \[31\]S\. Babu, R\. Liu, A\. Zhou, M\. Maire, G\. Shakhnarovich, and R\. Hanocka\(2023\)HyperFields: towards zero\-shot generation of nerfs from text\.External Links:arXiv:2310\.17075Cited by:[§2](https://arxiv.org/html/2609.30569#S2.SS0.SSS0.Px3.p1.1)\.
- \[32\]F\. Szatkowski, K\. J\. Piczak, P\. Spurek, J\. Tabor, and T\. Trzciński\(2023\)Hypernetworks build implicit neural representations of sounds\.InMachine Learning and Knowledge Discovery in Databases: Research Track: European Conference, ECML PKDD 2023, Turin, Italy, September 18–22, 2023, Proceedings, Part IV,Berlin, Heidelberg,pp\. 661–676\.External Links:[Document](https://dx.doi.org/10.1007/978-3-031-43421-1%5F39),ISBN 978\-3\-031\-43420\-4,[Link](https://doi.org/10.1007/978-3-031-43421-1_39)Cited by:[§2](https://arxiv.org/html/2609.30569#S2.SS0.SSS0.Px3.p1.1),[§3\.2](https://arxiv.org/html/2609.30569#S3.SS2.p3.1)\.
- \[33\]S\. Babu\(2025\)Acquiring and adapting priors for novel tasks via neural meta\-architectures\.External Links:arXiv:2507\.10446Cited by:[§2](https://arxiv.org/html/2609.30569#S2.SS0.SSS0.Px3.p1.1)\.
- \[34\]J\. Donahue, Y\. Jia, O\. Vinyals, J\. Hoffman, N\. Zhang, E\. Tzeng, and T\. Darrell\(2014\)DeCAF: a deep convolutional activation feature for generic visual recognition\.InProceedings of the 31st International Conference on Machine Learning,E\. P\. Xing and T\. Jebara \(Eds\.\),Proceedings of Machine Learning Research, Vol\.32\(1\),Bejing, China,pp\. 647–655\.Cited by:[§2](https://arxiv.org/html/2609.30569#S2.SS0.SSS0.Px4.p1.1)\.
- \[35\]T\. Chen, S\. Kornblith, M\. Norouzi, and G\. Hinton\(2020\)A simple framework for contrastive learning of visual representations\.InProceedings of the 37th International Conference on Machine Learning,H\. D\. III and A\. Singh \(Eds\.\),Proceedings of Machine Learning Research, Vol\.119,pp\. 1597–1607\.Cited by:[§2](https://arxiv.org/html/2609.30569#S2.SS0.SSS0.Px4.p1.1)\.
- \[36\]K\. He, X\. Chen, S\. Xie, Y\. Li, P\. Dollar, and R\. Girshick\(2022\)Masked autoencoders are scalable vision learners\.In2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\)2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 15979–15988\.Cited by:[§2](https://arxiv.org/html/2609.30569#S2.SS0.SSS0.Px4.p1.1)\.
- \[37\]M\. A\. Herzik, M\. Wu, and G\. C\. Lander\(2019\)High\-resolution structure determination of sub\-100 kda complexes using conventional cryo\-EM\.Nat\. Commun\.10\(1\),pp\. 1032\(en\)\.Cited by:[Figure 2](https://arxiv.org/html/2609.30569#S3.F2)\.
- \[38\]E\. F\. Pettersen, T\. D\. Goddard, C\. C\. Huang, E\. C\. Meng, G\. S\. Couch, T\. I\. Croll, J\. H\. Morris, and T\. E\. Ferrin\(2021\)UCSF ChimeraX: structure visualization for researchers, educators, and developers\.Protein Sci\.30\(1\),pp\. 70–82\(en\)\.Cited by:[§3\.1](https://arxiv.org/html/2609.30569#S3.SS1.p1.1)\.
- \[39\]T\. Dao, D\. Y\. Fu, S\. Ermon, A\. Rudra, and C\. Ré\(2022\)FlashAttention: fast and memory\-efficient exact attention with IO\-awareness\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§3\.3](https://arxiv.org/html/2609.30569#S3.SS3.SSS0.Px2.p1.1)\.
- \[40\]A\. Vaswani, N\. Shazeer, N\. Parmar, J\. Uszkoreit, L\. Jones, A\. N\. Gomez, Ł\. Kaiser, and I\. Polosukhin\(2017\)Attention is all you need\.InAdvances in Neural Information Processing Systems,I\. Guyon, U\. V\. Luxburg, S\. Bengio, H\. Wallach, R\. Fergus, S\. Vishwanathan, and R\. Garnett \(Eds\.\),Vol\.30,pp\.\.Cited by:[§3\.3](https://arxiv.org/html/2609.30569#S3.SS3.SSS0.Px2.p1.1)\.
- \[41\]J\. Su, Y\. Lu, S\. Pan, A\. Murtadha, B\. Wen, and Y\. Liu\(2021\)RoFormer: enhanced transformer with rotary position embedding\.External Links:arXiv:2104\.09864Cited by:[§3\.3](https://arxiv.org/html/2609.30569#S3.SS3.SSS0.Px2.p1.1)\.
- \[42\]A\. Andonian, Q\. Anthony, S\. Biderman, S\. Black, P\. Gali, L\. Gao, E\. Hallahan, J\. Levy\-Kramer, C\. Leahy, L\. Nestler, K\. Parker, M\. Pieler, J\. Phang, S\. Purohit, H\. Schoelkopf, D\. Stander, T\. Songz, C\. Tigges, B\. Thérien, P\. Wang, and S\. Weinbach\(2023\)GPT\-NeoX: Large Scale Autoregressive Language Modeling in PyTorch\.External Links:[Document](https://dx.doi.org/10.5281/zenodo.5879544),[Link](https://www.github.com/eleutherai/gpt-neox)Cited by:[§3\.3](https://arxiv.org/html/2609.30569#S3.SS3.SSS0.Px2.p1.1)\.
- \[43\]N\. Giri, L\. Wang, and J\. Cheng\(2024\)Cryo2StructData: a large labeled cryo\-EM density map dataset for AI\-based modeling of protein structures\.Sci\. Data11\(1\),pp\. 458\(en\)\.Cited by:[Appendix A](https://arxiv.org/html/2609.30569#A1.p1.1),[Appendix C](https://arxiv.org/html/2609.30569#A3.SS0.SSS0.Px1.p1.1),[§4\.1](https://arxiv.org/html/2609.30569#S4.SS1.p1.1)\.
- \[44\]M\. Jeon, R\. Raghu, M\. Astore, G\. Woollard, R\. Feathers, A\. Kaz, S\. M\. Hanson, P\. Cossio, and E\. Zhong\(2024\)CryoBench: diverse and challenging datasets for the heterogeneity problem in cryo\-em\.InAdvances in Neural Information Processing Systems,A\. Globerson, L\. Mackey, D\. Belgrave, A\. Fan, U\. Paquet, J\. Tomczak, and C\. Zhang \(Eds\.\),Vol\.37,pp\. 89468–89512\.External Links:[Document](https://dx.doi.org/10.52202/079017-2840)Cited by:[§4\.1](https://arxiv.org/html/2609.30569#S4.SS1.p1.1)\.
- \[45\]K\. Saotome, D\. Dudgeon, K\. Colotti, M\. J\. Moore, J\. Jones, Y\. Zhou, A\. Rafique, G\. D\. Yancopoulos, A\. J\. Murphy, J\. C\. Lin, W\. C\. Olson, and M\. C\. Franklin\(2023\)Structural analysis of cancer\-relevant TCR\-CD3 and peptide\-MHC complexes by cryoEM\.Nat\. Commun\.14\(1\),pp\. 2401\(en\)\.Cited by:[Figure 3](https://arxiv.org/html/2609.30569#S4.F3),[§4\.1](https://arxiv.org/html/2609.30569#S4.SS1.p2.1)\.
- \[46\]F\. P\. Maloney, J\. Kuklewicz, R\. A\. Corey, Y\. Bi, R\. Ho, L\. Mateusiak, E\. Pardon, J\. Steyaert, P\. J\. Stansfeld, and J\. Zimmer\(2022\)Structure, substrate recognition and initiation of hyaluronan synthase\.Nature604\(7904\),pp\. 195–201\(en\)\.Cited by:[Figure 3](https://arxiv.org/html/2609.30569#S4.F3)\.
- \[47\]T\. Zhou, L\. Chen, J\. Gorman, S\. Wang, Y\. D\. Kwon, B\. C\. Lin, M\. K\. Louder, R\. Rawi, E\. D\. Stancofski, Y\. Yang, B\. Zhang, A\. F\. Quigley, L\. E\. McCoy, L\. Rutten, T\. Verrips, R\. A\. Weiss, VRC Production Program, N\. A\. Doria\-Rose, L\. Shapiro, and P\. D\. Kwong\(2022\)Structural basis for llama nanobody recognition and neutralization of HIV\-1 at the CD4\-binding site\.Structure30\(6\),pp\. 862–875\.e4\(en\)\.Cited by:[Figure 3](https://arxiv.org/html/2609.30569#S4.F3)\.
- \[48\]S\. Chittori, J\. Hong, H\. Saunders, H\. Feng, R\. Ghirlando, A\. E\. Kelly, Y\. Bai, and S\. Subramaniam\(2018\)Structural mechanisms of centromeric nucleosome recognition by the kinetochore protein CENP\-N\.Science359\(6373\),pp\. 339–343\(en\)\.Cited by:[Figure 3](https://arxiv.org/html/2609.30569#S4.F3)\.
- \[49\]Schrödinger, LLC\(2026\)The PyMOL molecular graphics system, version 3\.8\.1\.Note:The PyMOL Molecular Graphics System, Version 3\.8\.1, Schrödinger, LLC\.Cited by:[§4\.1](https://arxiv.org/html/2609.30569#S4.SS1.p4.1)\.
- \[50\]Z\. Zhou, M\. M\. R\. Siddiquee, N\. Tajbakhsh, and J\. Liang\(2020\)UNet\+\+: redesigning skip connections to exploit multiscale features in image segmentation\.IEEE Trans\. Med\. Imaging39\(6\),pp\. 1856–1867\(en\)\.Cited by:[§B\.4](https://arxiv.org/html/2609.30569#A2.SS4.SSS0.Px1.p1.1),[§4\.3](https://arxiv.org/html/2609.30569#S4.SS3.p1.1)\.
- \[51\]P\. Stoica and P\. Babu\(2023\)Pearson\-matthews correlation coefficients for binary and multinary classification and hypothesis testing\.External Links:arXiv:2305\.05974Cited by:[§B\.4](https://arxiv.org/html/2609.30569#A2.SS4.SSS0.Px5.p1.1),[§4\.3](https://arxiv.org/html/2609.30569#S4.SS3.p6.1)\.
- \[52\]S\. Mishra, M\. K\. Yadav, A\. Dalal, M\. Ganguly, R\. Yadav, K\. Sawada, D\. Tiwari, N\. Roy, N\. Banerjee, J\. N\. Fung, J\. Marallag, C\. S\. Cui, X\. X\. Li, J\. D\. Lee, C\. A\. Dsouza, S\. Saha, P\. Sarma, G\. Rawat, H\. Zhu, H\. A\. Khant, R\. J\. Clark, F\. K\. Sano, R\. Banerjee, T\. M\. Woodruff, O\. Nureki, C\. Gati, and A\. K\. Shukla\(2025\)Molecular fingerprints of a convergent mechanism orchestrating diverse ligand recognition and species\-specific pharmacology at the complement anaphylatoxin receptors\.bioRxivorg\(en\)\.Cited by:[item 1](https://arxiv.org/html/2609.30569#S4.I1.i1.p1.1)\.
- \[53\]A\. Dalal, M\. K\. Yadav, M\. Ganguly, S\. Mishra, R\. Yadav, S\. Sinha, N\. Roy, D\. Tiwari, D\. Mukherjee, A\. Reyaz, C\. A\. Dsouza, A\. Nigam, N\. Banerjee, X\. X\. Li, R\. J\. Clark, T\. M\. Woodruff, R\. Banerjee, C\. Gati, and A\. K\. Shukla\(2026\)Structural basis of complement anaphylatoxin receptor activation by an immunostimulant lead candidate\.Proc\. Natl\. Acad\. Sci\. U\. S\. A\.123\(28\),pp\. e2614459123\(en\)\.Cited by:[item 2](https://arxiv.org/html/2609.30569#S4.I1.i2.p1.1),[item 3](https://arxiv.org/html/2609.30569#S4.I1.i3.p1.1)\.
- \[54\]K\. Kobayashi, K\. Kawakami, T\. E\. Matsui, S\. Yokoi, M\. Fukuda, T\. J\. Narita, H\. Arai, M\. Tambo, T\. Sumikama, M\. Tatsumi, K\. Yamashita, J\. Koyanagi, M\. Kugawa, H\. Ikeda, A\. Sumino, A\. Mitsutake, B\. K\. Kobilka, A\. Inoue, and H\. E\. Kato\(2026\)The dynamic basis of g\-protein recognition and activation by a GPCR\.Nature652\(8110\),pp\. 812–821\(en\)\.Cited by:[item 4](https://arxiv.org/html/2609.30569#S4.I1.i4.p1.1)\.
- \[55\]H\. M\. Berman, J\. Westbrook, Z\. Feng, G\. Gilliland, T\. N\. Bhat, H\. Weissig, I\. N\. Shindyalov, and P\. E\. Bourne\(2000\)The protein data bank\.Nucleic Acids Res\.28\(1\),pp\. 235–242\(en\)\.Cited by:[Appendix A](https://arxiv.org/html/2609.30569#A1.p1.1)\.
- \[56\]A\. Singer\(2018\)Mathematics for cryo\-electron microscopy\.External Links:arXiv:1803\.06714Cited by:[§A\.1](https://arxiv.org/html/2609.30569#A1.SS1.p1.1)\.
- \[57\]R\. Sanchez\-Garcia, J\. Gomez\-Blanco, A\. Cuervo, J\. M\. Carazo, C\. O\. S\. Sorzano, and J\. Vargas\(2021\)DeepEMhancer: a deep learning solution for cryo\-EM volume post\-processing\.Commun\. Biol\.4\(1\),pp\. 874\(en\)\.Cited by:[§A\.1](https://arxiv.org/html/2609.30569#A1.SS1.p1.1)\.
- \[58\]J\. He, T\. Li, and S\. Huang\(2023\)Improvement of cryo\-EM maps by simultaneous local and non\-local deep learning\.Nat\. Commun\.14\(1\),pp\. 3217\(en\)\.Cited by:[§A\.1](https://arxiv.org/html/2609.30569#A1.SS1.p1.1)\.
- \[59\]I\. Agarwal, J\. Kaczmar\-Michalska, S\. F\. Nørrelykke, and A\. J\. Rzepiela\(2024\)Refinement of cryo\-EM 3D maps with a self\-supervised denoising model: crefdenoiser\.IUCrJ11\(Pt 5\),pp\. 821–830\(en\)\.Cited by:[§A\.1](https://arxiv.org/html/2609.30569#A1.SS1.p1.1)\.
- \[60\]K\. Ramlaul, C\. M\. Palmer, and C\. H\. S\. Aylett\(2019\)A local agreement filtering algorithm for transmission EM reconstructions\.J\. Struct\. Biol\.205\(1\),pp\. 30–40\(en\)\.Cited by:[§A\.1](https://arxiv.org/html/2609.30569#A1.SS1.p1.1)\.
- \[61\]N\. Otsu\(1979\)A threshold selection method from gray\-level histograms\.IEEE Trans\. Syst\. Man Cybern\.9\(1\),pp\. 62–66\.Cited by:[item 1](https://arxiv.org/html/2609.30569#A1.I1.i1.p1.1)\.
- \[62\]J\. Kyte and R\. F\. Doolittle\(1982\)A simple method for displaying the hydropathic character of a protein\.J\. Mol\. Biol\.157\(1\),pp\. 105–132\(en\)\.Cited by:[5th item](https://arxiv.org/html/2609.30569#A1.I2.i5.p1.1)\.
- \[63\]S\. Mitternacht\(2016\)FreeSASA: an open source C library for solvent accessible surface area calculations\.F1000Res\.5,pp\. 189\(en\)\.Cited by:[6th item](https://arxiv.org/html/2609.30569#A1.I2.i6.p1.1)\.
- \[64\]M\. Z\. Tien, A\. G\. Meyer, D\. K\. Sydykova, S\. J\. Spielman, and C\. O\. Wilke\(2013\)Maximum allowed solvent accessibilites of residues in proteins\.PLoS One8\(11\),pp\. e80635\(en\)\.Cited by:[7th item](https://arxiv.org/html/2609.30569#A1.I2.i7.p1.1)\.
- \[65\]J\. J\. G\. Ortiz, J\. Guttag, and A\. V\. Dalca\(2024\)Magnitude invariant parametrizations improve hypernetwork learning\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=fJNnerz6iH)Cited by:[§B\.1](https://arxiv.org/html/2609.30569#A2.SS1.SSS0.Px2.p1.1)\.
- \[66\]I\. Loshchilov and F\. Hutter\(2019\)Decoupled weight decay regularization\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=Bkg6RiCqY7)Cited by:[§B\.2](https://arxiv.org/html/2609.30569#A2.SS2.p1.1),[§B\.4](https://arxiv.org/html/2609.30569#A2.SS4.SSS0.Px2.p1.1)\.
- \[67\]D\. Chicco and G\. Jurman\(2020\)The advantages of the matthews correlation coefficient \(MCC\) over F1 score and accuracy in binary classification evaluation\.BMC Genomics21\(1\),pp\. 6\(en\)\.Cited by:[§B\.4](https://arxiv.org/html/2609.30569#A2.SS4.SSS0.Px5.p1.1)\.
## Appendix AData Preparation
We base our data curation pipeline heavily on the method in Cryo2StructData\[[43](https://arxiv.org/html/2609.30569#bib.bib41)\], starting with a list of 7392 curated EMDB IDs with corresponding fitted PDB maps\. We downloaded all density maps from the EMDB\[[2](https://arxiv.org/html/2609.30569#bib.bib44)\]and corresponding\.pdbfiles from the RCSB Protein Data Bank\[[55](https://arxiv.org/html/2609.30569#bib.bib45)\]\. All density maps are resampled to a uniform voxel size of 3 Å×\\times3 Å×\\times3 Å using theresamplecommand in ChimeraX\.
### A\.1Denoising CryoEM Volumes
Reconstruction of 3D electron densities from raw cryoEM images is a notoriously difficult problem, with the raw measurements extremely low signal\-to\-noise ratio\[[56](https://arxiv.org/html/2609.30569#bib.bib50)\]\. As a result, even deposited maps can be highly noisy and contain large amounts of “dust”\. In order to allow our tokenizer to ignore background patches that contain only noise, we wish to threshold out all noise to a background value of zero\. Various competing methods have been developed before, some deep\-learning based\[[57](https://arxiv.org/html/2609.30569#bib.bib49),[58](https://arxiv.org/html/2609.30569#bib.bib47),[59](https://arxiv.org/html/2609.30569#bib.bib46)\]and some using classical signal processing techniques\[[60](https://arxiv.org/html/2609.30569#bib.bib48)\]\. For simplicity, we devise our own denoising algorithm, which performs the following steps on the 3 Å resampled map:
1. 1\.We first apply Otsu thresholding\[[61](https://arxiv.org/html/2609.30569#bib.bib51)\]to mask out the background, then compute the volume reduction of the tight bounding box around the foreground\. If the reduction is less than 10% \(indicating an unusually noisy background\), we instead sweep 50 threshold values between the minimum and the 50th percentile voxel value and select the one that maximizes the discontinuity in data reduction to threshold out the background\.
2. 2\.To further remove “dust”, we compute the size of all connected components \(using 26\-connectivity\) in the volume after thresholding and remove anything smaller in size than 1% of the largest connected component\.
Since the voxel values in cryoEM maps are unitless, we linearly scale all voxel values to\[0,1\]\[0,1\]after denoising\. In Figure[6](https://arxiv.org/html/2609.30569#A1.F6), we show four representative samples before and after applying our denoising method\. Denoised volumes are saved to disk and padded to size\(96,96,96\)\(96,96,96\)by the dataloader during training and inference; any volumes exceeding this size are discarded, leaving a total of 5439/684/689 usable train/validation/test samples\. Splits were generated randomly before filtering for size\.
Figure 6:Illustration of the effect of our denoising method on four uncurated cryoEM volumes\. By thresholding out the “dust” present in raw deposited maps, it becomes straightforward for the tokenizer component of Atelier to only tokenize patches containing real signal\.
### A\.2Label Generation
Since all cryoEM volume in our dataset have corresponding fitted PDB atomic structures, we are able to align atomic coordinates with the cryoEM voxel grid and label individual voxels with properties derived from the PDB files\. For each cryoEM volume, we compute the following labels\.
- •Secondary structure \(classification, 3 labels\): at voxels containing alpha carbons, we label the secondary structure corresponding that residue as determined by ChimeraX\.
- •Amino acid \(classification, 20 labels\): at voxels containing alpha carbons, we label the identity of the corresponding residue as read from the PDB file\.
- •Molecule type \(classification, 22 labels\): an extension of secondary structure; at voxels containing either alpha carbons, nucleotide glycosidic carbons, or other heavy non\-water heteroatoms, we label them accordingly \(20 labels for amino acids, 1 for any nucleotide, and 1 for any heteroatom\)\.
- •Molecule class \(classification, 3 labels, same model as molecule type\): We do not train a separate model for this task; instead, we remap the molecule type predictions by grouping all amino acid labels into a single label, so that we may evaluate the molecule type\-trained model’s ability to distinguish between protein, nucleotide, and other\.
- •Hydrophobicity \(regression\): at voxels containing alpha carbons, we label with the Kyte\-Doolittle hydrophobicity\[[62](https://arxiv.org/html/2609.30569#bib.bib52)\]of the corresponding residue\.
- •Absolute SASA \(regression\): at voxels containing alpha carbons, we label with the absolute solvent accessible surface area of the corresponding residue as computed by FreeSASA\[[63](https://arxiv.org/html/2609.30569#bib.bib53)\]\.
- •Relative SASA \(regression\): at voxels containing alpha carbons, we label with the absolute solvent accessible surface area of the corresponding residue as computed by FreeSASA, divided by the empirical maximum possible SASA for each residue type as given by\[[64](https://arxiv.org/html/2609.30569#bib.bib54)\]\.
- •B\-factor \(regression\): at voxels containing alpha carbons, we label with the mean crystallographic B\-factor, as read from the PDB file\.
For each label, all voxels that do not contain alpha\-carbons \(or nucleotides and heteroatoms, for molecule type\) are assigned a background label\.
## Appendix BArchitecture and Training Details
### B\.1Atelier architectural details
The total number of trainable parameters of Atelier with the hyperparameters described below is 38,332,806\. Our architecture is adapted in part from that of Trans\-INR\[[12](https://arxiv.org/html/2609.30569#bib.bib40)\]\.
#### The convolutional tokenizer\.
Our dataloader pads all volumes to a size of\(96,96,96\)\(96,96,96\), but most proteins do not take up this whole volume\. To exploit this sparsity, we divide the voxel grid into patches of size\(6,6,6\)\(6,6,6\)for a total of 4096 patches\. Patches which only contain background voxels are discarded, while the remaining active patches are collected into a single batch and processed by a shared 3D convolutional neural network that outputs ad/2d/2\-dimensional token representing each active patch, whered=256d=256is the latent embedding dimension of the transformer\. Across the dataset, the median number of active patches is 236, meaning that 95% of patches are discarded\.
#### Magnitude Invariant Parametrization encoder\.
Before entering the transformer, both the volume tokens and the weight tokens are passed through a Magnitude Invariant Parametrization encoder \(MIP\[[65](https://arxiv.org/html/2609.30569#bib.bib55)\]\):
MIP:x↦\[cos\(πx/2\),sin\(πx/2\)\]\.\\MIP:x\\mapsto\[\\cos\(\\pi x/2\),\\sin\(\\pi x/2\)\]\.\(4\)We found in early experiments that this improved training convergence\. Note that this doubles the dimension of each token before feeding into the transformer\.
#### The hypernetworkHφH\_\{\\varphi\}\.
Our hypernetwork is a transformer encoder with latent dimension 256, 24 attention layers, 8 attention heads, a head dimension of 64, and a feedforward hidden dimension of 2048 with GeLU activations and a dropout rate of0\.10\.1\. Rotary positional embeddings are applied independently per spatial axis using the integer patch grid coordinates\(i,j,k\)\(i,j,k\)\. Each head dimension is divided into three equal groups\(x,y,z\)\(x,y,z\)\. Each group uses 20 rotary dimensions \(60 total out of 64; the remaining 4 are left unrotated\), with base frequency 10,000\.
#### INR weight generation\.
For each of theK\+1K\+1INR layers, we maintain a set of learned*weight tokens*: one\(d/2\)\(d/2\)\-dimensional token per output neuron in that layer, initialized from𝒩\(0,1\)\\mathcal\{N\}\(0,1\)\. These tokens are data\-independent and shared across all input volumes\. At each forward pass they are concatenated with the volume tokens along the sequence dimension, jointly MIP\-encoded \(Equation[4](https://arxiv.org/html/2609.30569#A2.E4)\), and processed by the transformer\. Because the weight tokens carry no spatial meaning, they are assigned position\(0,0,0\)\(0,0,0\)for the purpose of computing rotary embeddings\.
After the transformer, the output positions corresponding to the weight tokens are extracted\. For INR layeriiwithnioutn\_\{i\}^\{\\mathrm\{out\}\}output neurons andniinn\_\{i\}^\{\\mathrm\{in\}\}input features, thenioutn\_\{i\}^\{\\mathrm\{out\}\}extracteddd\-dimensional tokens are passed through a dedicated learnable linear projector𝐏i∈ℝniin×d\\mathbf\{P\}\_\{i\}\\in\\mathbb\{R\}^\{n\_\{i\}^\{\\mathrm\{in\}\}\\times d\}\. The resulting rows areℓ2\\ell\_\{2\}\-normalized and scaled by a learned per\-layer scalarαi\\alpha\_\{i\}\(initialized to0\.10\.1\) to decouple weight direction from magnitude:
Wi\(j\)=αi⋅𝐏i𝐡j\(i\)‖𝐏i𝐡j\(i\)‖2,W\_\{i\}^\{\(j\)\}=\\alpha\_\{i\}\\cdot\\frac\{\\mathbf\{P\}\_\{i\}\\mathbf\{h\}\_\{j\}^\{\(i\)\}\}\{\\bigl\\\|\\mathbf\{P\}\_\{i\}\\mathbf\{h\}\_\{j\}^\{\(i\)\}\\bigr\\\|\_\{2\}\},\(5\)where𝐡j\(i\)∈ℝd\\mathbf\{h\}\_\{j\}^\{\(i\)\}\\in\\mathbb\{R\}^\{d\}is the transformer output for thejj\-th weight token of layerii\. Thenioutn\_\{i\}^\{\\mathrm\{out\}\}rowsWi\(j\)W\_\{i\}^\{\(j\)\}are stacked to form the weight matrix for layerii\.
Biases are not generated by the hypernetwork\. Each INR layer instead has a learned bias vector shared across all inputs, initialized uniformly in\[−1/niin,1/niin\]\\bigl\[\-1/\\sqrt\{n\_\{i\}^\{\\mathrm\{in\}\}\},\\,1/\\sqrt\{n\_\{i\}^\{\\mathrm\{in\}\}\}\\bigr\]\.
#### The INR decoderℳθ\\mathcal\{M\}\_\{\\theta\}
Instead of taking raw spatial coordinates as input, the INR first encodes spatial query pointsx∈\[0,1\]3x\\in\[0,1\]^\{3\}with a fixed Fourier feature encoding:
γ:x↦\[sin\(2πfkxi\),cos\(2πfkxi\)\]k=1,…,16i∈\{1,2,3\}∈ℝ96,\\gamma:x\\mapsto\\bigl\[\\sin\(2\\pi f\_\{k\}x\_\{i\}\),\\,\\cos\(2\\pi f\_\{k\}x\_\{i\}\)\\bigr\]\_\{\\begin\{subarray\}\{c\}k=1,\\dots,16\\\\ i\\,\\in\\,\\\{1,2,3\\\}\\end\{subarray\}\}\\;\\in\\;\\mathbb\{R\}^\{96\},\(6\)where\{fk\}k=116\\\{f\_\{k\}\\\}\_\{k=1\}^\{16\}are log\-spaced in\[1,256\]\[1,256\]\. With each forward pass, the INR is evaluated at all96396^\{3\}grid points to render the full volume\.
### B\.2Atelier pretraining details\.
Atelier was trained using AdamW\[[66](https://arxiv.org/html/2609.30569#bib.bib56)\]with a\(β1,β2\)=\(0\.90,0\.95\)\(\\beta\_\{1\},\\beta\_\{2\}\)=\(0\.90,0\.95\)and a weight decay of0\.010\.01\. We trained for 5000 total epochs with a linear warmup of 10 epochs to a learning rate of 0\.001 followed by a cosine decay to 0\.0001\. Training was distributed over 8 H100 GPUs with an effective batch size of 64 volumes per batch\. Gradients were clipped at 1\.0\. During training, we augment the volumes by a random element of the octahedral groupOhO\_\{h\}\. Total wall time for pretraining was approximately 90 hours\.
### B\.3Transformer\-only ablation
To highlight the importance of the fine\-grained features exposed by the hypernetwork structure of Atelier, we perform an ablation where we train a transformer\-only variant of our architecture to reconstruct cryoEM densities, illustrated in Figure[7](https://arxiv.org/html/2609.30569#A2.F7)\(a\)\. The volume tokenizer and transformer remain exactly the same, except the transformer no longer accepts weight tokens\. Instead, the transformed volume tokens are directly used to reconstruct the volume at the corresponding patch using a learned linear projection\. The patches are then reassembled and the loss is computed as in Equation[1](https://arxiv.org/html/2609.30569#S3.E1)\. We train this ablated model with exactly the same hyperparameters as Atelier\.
Figure 7:The transformer\-only ablation architecture\.\(a\)As an ablation, we train a transformer\-only model that directly maps per\-patch volume tokens to patch reconstructions through a learnable linear projection\.\(b\)Coarse patch\-level features can be extracted by returning the transformed volume token corresponding to the patchpxp\_\{x\}containing a query pointxx\. Note that all query points within the same patch will have the same coarse feature\.
### B\.4EMNUSS property prediction experiment details
#### Architecture\.
We use EMNUSS\[[6](https://arxiv.org/html/2609.30569#bib.bib15)\], a 3D nested U\-Net\[[50](https://arxiv.org/html/2609.30569#bib.bib42)\]architecture for per\-voxel property prediction\. It takes aCC\-channel\(D,D,D,C\)\(D,D,D,C\)volume as input and outputs a\(D,D,D,k\)\(D,D,D,k\)volume forkk\-class classification or\(D,D,D,1\)\(D,D,D,1\)for regression\. For all three variants, we use the exact same implementation as the original EMNUSS codebase, only changing the input and output dimensions\.
#### Training\.
The training setup is identical for all three input variants \(volume\-only, augmentation with coarse features, augmentation with fine features\) in the property prediction experiment from §[4\.3](https://arxiv.org/html/2609.30569#S4.SS3)\. We use AdamW\[[66](https://arxiv.org/html/2609.30569#bib.bib56)\]with\(β1,β2\)=\(0\.90,0\.999\)\(\\beta\_\{1\},\\beta\_\{2\}\)=\(0\.90,0\.999\)and a weight decay of 0\.01\. We train for 300 total epochs with a linear warmup of 10 epochs to a learning rate of 0\.001 followed by a cosine decay to 0\.0001\. Remaining training hyperparameters \(weight decay, batch size, gradient clipping\) are exactly the same as in Atelier pretraining as described in §[B\.2](https://arxiv.org/html/2609.30569#A2.SS2), with the exception that we do not train with random rotations fromOhO\_\{h\}\. Classification tasks are trained with cross\-entropy loss with the background class excluded and evaluated by accuracy over non\-background voxels\. Regression tasks are trained with MSE computed only over non\-background voxels and evaluated by the mean absolute relative error over non\-background voxels, reported as “relative error” in Table[2](https://arxiv.org/html/2609.30569#S4.T2)\. We use the same train/validation/test splits as Atelier pretraining\.
#### Fine feature extraction \(Atelier\)\.
For each annotated voxel at positionxxwe:
1. 1\.collect all voxel positions within a sphere of radius 3 voxels centered atxx;
2. 2\.query the*frozen*best\-checkpoint Atelier INR at each sphere voxel and extract theK=4K\{=\}4post\-ReLU hidden\-layer activations, eachH=256H\{=\}256\-dimensional;
3. 3\.mean\-pool across theKKlayers to obtain oneHH\-dimensional vector per sphere voxel;
4. 4\.perform a Gaussian\-weighted mean\-pooling across sphere voxels withσ=1\.5\\sigma=1\.5voxels, so voxels further fromxxcontribute less to the voxel\-pooled representation\.
This yields one 256\-dimensional feature per annotated voxel, precomputed offline from the frozen Atelier checkpoint\.
#### Coarse feature extraction \(transformer\-only ablation\)\.
For each annotated voxel at positionxx, retrieve theH=256H=256\-dimensional transformed volume token of the6×6×66\{\\times\}6\{\\times\}6voxel patchpxp\_\{x\}containingxxfrom the frozen best\-checkpoint transformer\-only autoencoder \(Figure[7](https://arxiv.org/html/2609.30569#A2.F7)\(b\)\)\. Consequently, all voxels within the same patch share an identical representation\.
#### MCC for Classification Tasks
In Table[4](https://arxiv.org/html/2609.30569#A2.T4), we report the Matthews correlation coefficient, a robust metric for highly imbalanced classification tasks\[[51](https://arxiv.org/html/2609.30569#bib.bib59),[67](https://arxiv.org/html/2609.30569#bib.bib60)\]for classification experiments reported in Table[2](https://arxiv.org/html/2609.30569#S4.T2)\.
Table 4:Matthews correlation coefficient \(MCC↑\\uparrow\) on the classification tasks from Table[2](https://arxiv.org/html/2609.30569#S4.T2)\. Metrics are averaged over all labeled voxels in the test set\.
#### Per\-volume property prediction results\.
The property prediction metrics reported in Table[2](https://arxiv.org/html/2609.30569#S4.T2)are pooled over all labeled voxels in the test set\. However, since voxels within a volume may be correlated, pooled results do not ensure consistency of improvements across volumes\. We therefore run the following analysis:
1. 1\.For each of the tasks, we compute theper\-volumemetric with the fine features and the three baselines \(volume\-alone, coarse features, random features\)\. We then compute the difference between the fine metric and three baselines for each volume; differences are oriented so that a positive value means the fine model is better\.
2. 2\.In Table[5](https://arxiv.org/html/2609.30569#A2.T5), for each comparison within each task, we report the Wilcoxon signed\-rank testpp\-value, the Holm\-Bonferonni correctedpp\-value to account for multiple testing, the win rate \(fraction of maps on which the final map is better\), and the Hodges\-Lehmann effect size with a 95% distribution\-free confidence interval \(positive means fine features are better\)\. We see that for most comparisons, the Atelier fine features significantly outperform baselines\.
Table 5:Per\-volume comparison on fine Atelier features against baselines\. We see that for amino acid prediction, coarse significantly outperforms fine \(row 4, note the negative effect size\), while for B\-factor, fine does not significantly outperform volume\-alone \(row 24\)\. For all other comparisons, the improvement is modest but significant, supporting the efficacy of Atelier\-derived features\.
## Appendix CLimitations and Future Work
We note several boundaries of our current evaluation that offer clear directions for future work\.
#### Generalization to non\-redundant structures\.
Our train, validation, and test splits are random over EMDB IDs, following prior cryoEM benchmarks\[[43](https://arxiv.org/html/2609.30569#bib.bib41),[6](https://arxiv.org/html/2609.30569#bib.bib15)\]\. All models in our comparison are trained and evaluated on the same splits, so the relative improvements in Table[2](https://arxiv.org/html/2609.30569#S4.T2)are not affected\. However, random splitting can place homologous proteins or alternate conformations on both sides, which limits our ability to claim absolute generalization to entirely unseen folds\. Re\-evaluation on splits stratified by sequence identity or CATH topology is a natural next step\.
#### Pretraining cost and high\-frequency representation\.
Pretraining Atelier requires roughly 720 H100\-hours\. Once trained, the model is fully amortized: a single forward pass producesθV\\theta\_\{V\}for an unseen volume, and feature extraction is a deterministic offline precomputation\. The pretraining cost remains a barrier to iteration and to scaling the framework to larger volumes or higher\-resolution grids, and reducing it through better tokenization, sparser attention, or distillation is a promising direction\. Separately, the highest frequencies in our Fourier coordinate encoding can exceed the Nyquist limit of the96396^\{3\}grid; the spectral bias of ReLU MLPs acts as an implicit regularizer in practice, but principled bandlimiting is a logical refinement\.相似文章
CryoACE:一种以原子为中心的冷冻电镜精确自动模型构建框架
CryoACE是一种以原子为中心的框架,用于从冷冻电镜密度图自动构建模型,具有迭代精炼和利用局部分辨率先验的无训练引导功能。它在静态基准测试中显著优于现有方法,并在复杂数据集上揭示了原子级别的动态构象。
3D Masked Autoencoders是显微镜下体积和多模态细胞表示的鲁棒学习器
本文提出了用于体积显微镜数据的3D Masked Autoencoders,并展示了在下游单细胞任务中,3D建模优于2D最大投影和基于切片的变体,而通过与蛋白质语言模型的跨模态对齐进一步提升了性能。
解决AI中叠加问题以实现可解释性与患者神经元图像的跨模态对齐
本文引入稀疏自编码器来解决神经网络中的叠加问题,提高潜在空间的可解释性和几何保真度,并提出GW-map实现图像表示与单细胞RNA测序数据之间的跨模态对齐。
基于混合潜空间建模的结构连接组获取变异无监督学习
本文提出了一种无监督框架,通过混合潜空间建模来模拟结构连接组中与获取相关的变异,利用架构退火编码器输出消除了手动容量调优的需求。
深度强化学习中冻结随机CNN特征提取器的涌现稀疏性
本文报告发现:使用冻结的随机初始化CNN特征提取器的深度强化学习智能体,在没有任何稀疏性诱导目标的情况下,自发地产生极端稀疏的全连接表示,通过极少数神经元压缩任务相关信息。