ED-CSP: Crystal Structure Prediction from Electron Diffraction

arXiv cs.LG Papers

Summary

ED-CSP is a machine learning model that predicts crystal structures from electron diffraction patterns, achieving improved match rates over the PXRD-based PXRDGen and demonstrating the value of multi-view diffraction data.

arXiv:2608.06448v1 Announce Type: new Abstract: Recovering a periodic 3D crystal structure from sparse, unindexed electron diffraction (ED) observations is a challenging generative inverse problem. Existing ED-based learning methods mainly predict crystallographic labels, reconstruct structures from indexed reflections, or retrieve candidates from finite structure libraries. Here, we introduce ED-CSP, a machine learning framework that predicts crystal structures from chemical composition, atom count, and multiple detector-plane ED spot sets. ED-CSP combines a relational set encoder, permutation-invariant multi-view aggregation, and a periodic flow generator to jointly predict lattice parameters and fractional atomic coordinates. To train the model, we construct ED-CS, a dataset of 4.85 million simulated multi-view ED crystal structures, deduplicated across seven materials repositories and filtered to exclude CHILI-100K overlaps. On 2,075 held-out CHILI-100K materials, ED-CSP trained only on CHILI achieves a structural match rate of 57.49% MR@5, outperforming PXRDGen (52.92%), a state-of-the-art crystal structure prediction model conditioned on powder X-ray diffraction. Scaling training data further improves performance: initializing from a one-million-structure precursor raises MR@5 to 66.27%. On 1,024 compositions absent from the training retrieval library, the model still achieves 53.52% MR@5, demonstrating true generative capability beyond exact-formula retrieval. Replacing target ED observations with diffraction from non-isomorphic structures of identical composition decreases MR@5 by 22.09 percentage points, confirming that predictions depend on the input diffraction patterns rather than composition alone. ED-CSP and ED-CS establish a benchmark for generative crystal structure prediction from sparse ED observations and provide a foundation for future transfer to experimental data.
Original Article
View Cached Full Text

Cached at: 08/10/26, 08:01 AM

# Crystal Structure Prediction from Electron Diffraction
Source: [https://arxiv.org/html/2608.06448](https://arxiv.org/html/2608.06448)
###### Abstract

Recovering a periodic 3D crystal structure from sparse, unindexed detector\-plane observations is a challenging generative inverse problem\. Prior electron diffraction \(ED\) learning methods largely predict crystallographic labels, reconstruct structures from indexed reflections, or retrieve structures from finite libraries\. In this paper, we consider the task of crystal structure prediction from known composition and atom count, together with multiple detector\-plane ED spot sets, and introduce ED\-CSP, a machine learning model that combines a relational set encoder, permutation\-invariant multi\-view aggregation, and a periodic flow generator to generate the lattice and fractional atomic coordinates\. To train ED\-CSP, we construct Electron Diffraction Crystal Structures \(ED\-CS\), a4\.854\.85\-million\-structure resource with simulated multi\-view ED, deduplicated across seven materials repositories and filtered to exclude CHILI\-100K matches under a fixed structural matcher\. On 2,075 held\-out CHILI\-100K materials, CHILI\-only ED\-CSP achieves a structural match rate \(MR\) of57\.49%57\.49\\text\{\\,\}\\mathrm\{\\char 37\\relax\}at five candidates per query \(MR@5\), versus52\.92%52\.92\\text\{\\,\}\\mathrm\{\\char 37\\relax\}for PXRDGen, a state\-of\-the\-art crystal structure prediction \(CSP\) model conditioned on powder X\-ray diffraction \(PXRD\); both use the same periodic\-generator architecture but modality\-specific encoders and inputs\. To demonstrate that increasing the dataset size improves model performance, we warm\-start the full model from a separate one\-million\-structure precursor and show that this raises MR@5 to66\.27%66\.27\\text\{\\,\}\\mathrm\{\\char 37\\relax\}\. On 1,024 queries whose reduced formulas are absent from the train/validation retrieval library, this registry\-1M\-initialized model retains53\.52%53\.52\\text\{\\,\}\\mathrm\{\\char 37\\relax\}MR@5, demonstrating recovery where exact\-formula lookup has no candidate\. For the same model, replacing target ED observations with those from a non\-isomorphic same\-formula donor on 67 queries lowers mean MR@5 by22\.0922\.09percentage points across five generation seeds, providing evidence of query\-specific diffraction use\. ED\-CSP and ED\-CS provide a controlled benchmark for generative inference from sparse simulated ED and future experimental\-transfer studies\.

## Introduction

Many crystalline materials cannot be grown as crystals large enough for conventional single\-crystal X\-ray diffraction, whereas ED can collect structural signal from individual nanocrystals\(Gemmiet al\.[2019](https://arxiv.org/html/2608.06448#bib.bib11); Ungeet al\.[2025](https://arxiv.org/html/2608.06448#bib.bib12)\)\. Powder X\-ray diffraction \(PXRD\) remains broadly accessible for bulk powders composed of randomly oriented crystallites, making the two modalities complementary rather than interchangeable\(Segalet al\.[2026](https://arxiv.org/html/2608.06448#bib.bib21); Gemmiet al\.[2019](https://arxiv.org/html/2608.06448#bib.bib11)\)\. Unlike the orientation\-averaged PXRD profile, each ED view samples an oriented region of reciprocal space and preserves detector\-plane relationships among scattering vectors\. Both 3D ED and scanning diffraction experiments produce orientation\-dependent reciprocal\-space observations, although their acquisition geometries differ from the discrete simulated views studied here\(Gemmiet al\.[2019](https://arxiv.org/html/2608.06448#bib.bib11); Savitzkyet al\.[2021](https://arxiv.org/html/2608.06448#bib.bib24)\)\. This orientation dependence motivates a multi\-view inverse problem in which geometric information is available while simulation\-to\-experiment transfer and incomplete orientation coverage remain explicit limitations\. Figure[1](https://arxiv.org/html/2608.06448#Sx1.F1)contrasts these observation regimes\.

Diffraction\-conditioned CSP is now established for PXRD through contrastive pretraining, diffusion and flow models, and autoregressive generation\(Laiet al\.[2025](https://arxiv.org/html/2608.06448#bib.bib15); Guoet al\.[2025](https://arxiv.org/html/2608.06448#bib.bib18); Liet al\.[2025](https://arxiv.org/html/2608.06448#bib.bib16); Johansenet al\.[2026](https://arxiv.org/html/2608.06448#bib.bib17); Segalet al\.[2026](https://arxiv.org/html/2608.06448#bib.bib21)\)\. In ED, however, machine learning has mainly predicted crystallographic labels or retrieved structures from finite databases\(Gleasonet al\.[2024](https://arxiv.org/html/2608.06448#bib.bib25); Nathaniet al\.[2026](https://arxiv.org/html/2608.06448#bib.bib26); Penget al\.[2026](https://arxiv.org/html/2608.06448#bib.bib27)\)\. These tasks show that sparse ED patterns contain learnable structural signal, but they stop short of predicting lattice and fractional atomic coordinates from unindexed multi\-view spot lists given composition\. ED\-CSP targets this gap with composition\-conditioned crystal\-geometry prediction\.

ED\-CSP encodes each view with a shared relational spot encoder, aggregates information across views, and jointly optimizes the resulting representation with a periodic flow generator\. Known composition fixes the atom types and count, while the ED branch receives detector\-plane spot coordinates and intensities\.

On CHILI\-100K, we use a common held\-out split to evaluate signal use, registry\-scale transfer, finite\-library coverage, same\-dataset PXRD generation, and indexed\-reflection reconstruction\(Friis\-Jensenet al\.[2024](https://arxiv.org/html/2608.06448#bib.bib22)\)\. ED input interventions and a converged composition\-only reference support signal use beyond composition; registry scaling improves held\-out recovery, and library\-coverage stratification separates generation from finite\-database lookup\. The study uses simulated detector\-plane ED with known composition\.

#### Contributions\.

This paper makes the following focused contributions:

- •A formulation of composition\-conditioned CSP from unindexed sparse multi\-view ED, instantiated by ED\-CSP\.
- •Controlled ED\-input interventions and a converged composition\-only reference testing use of the ED conditioning signal\.
- •ED\-CS, a CHILI\-disjoint corpus of 4,852,131 deduplicated structures with simulated multi\-view ED, together with registry\-scale transfer and comparisons across generation, retrieval, and indexed reconstruction\.

![Refer to caption](https://arxiv.org/html/2608.06448v1/assets/edcsp_diffraction_regimes.png)Figure 1:Diffraction observation regimes\.ED\-CSP conditions on discrete sparse ED views from multiple orientations, whereas PXRD aggregates randomly oriented crystallites into a one\-dimensional radial profile\.

## Related Work

#### Diffraction\-conditioned generation\.

PXRD\-conditioned CSP has progressed from establishing diffraction as a generative condition to testing which auxiliary information is available at inference and whether diffraction resolves structural ambiguity\. XtalNet combines contrastive PXRD–structure pretraining with equivariant generation; PXRDnet and PXRDGen condition diffusion or flow generators on formula and PXRD, with PXRDGen additionally supporting lattice inference and Rietveld refinement; deCIFer generates crystallographic information file \(CIF\) sequences autoregressively\(Laiet al\.[2025](https://arxiv.org/html/2608.06448#bib.bib15); Liet al\.[2025](https://arxiv.org/html/2608.06448#bib.bib16); Johansenet al\.[2026](https://arxiv.org/html/2608.06448#bib.bib17); Guoet al\.[2025](https://arxiv.org/html/2608.06448#bib.bib18)\)\. PXRDGen evaluates its contrastively pretrained XRD encoders through retrieval and then uses them, either frozen or trainable, to condition structure generation\(Liet al\.[2025](https://arxiv.org/html/2608.06448#bib.bib16)\)\. More recent systems emphasize experimental transfer and varying chemical or crystallographic inputs: XRDSol receives stoichiometry and unit\-cell parameters, RealPXRD\-Solver supports lattice\-conditioned and lattice\-free inference after large\-scale simulated pretraining, and XRDiff evaluates full and partial composition with composition\-grouped polymorph splits\(Yuet al\.[2026](https://arxiv.org/html/2608.06448#bib.bib19); Liet al\.[2026](https://arxiv.org/html/2608.06448#bib.bib20); Segalet al\.[2026](https://arxiv.org/html/2608.06448#bib.bib21)\)\. Outside diffraction\-conditioned CSP, Atomistic Language Models couple a language backbone to an atomistic diffusion decoder and report strong composition\-conditioned crystal recovery, emphasizing the importance of isolating diffraction\-specific gains from learned structural priors\(Edamadakaet al\.[2026](https://arxiv.org/html/2608.06448#bib.bib32)\)\. ED\-CSP addresses the complementary setting of sparse multi\-view ED spot lists\.

#### Sparse multi\-view ED learning\.

ED representation learning provides the closest architectural precedent\. RF\-ED predicts crystal systems, space groups, and lattice parameters from one or more simulated two\-dimensional ED patterns, while PE\-AG\-GMoE processes Bragg spots as variable\-size relational sets and aggregates predictions across orientations for crystallographic classification\(Gleasonet al\.[2024](https://arxiv.org/html/2608.06448#bib.bib25); Nathaniet al\.[2026](https://arxiv.org/html/2608.06448#bib.bib26)\)\. Recent multi\-view selected\-area electron diffraction \(SAED\) learning fuses two views for symmetry prediction and formula\-constrained retrieval from a finite structure database\(Penget al\.[2026](https://arxiv.org/html/2608.06448#bib.bib27)\)\. ED\-CSP changes the output from labels or database identities to lattice and fractional atomic coordinates given composition\.

#### Indexed ED solution and refinement\.

Learned ED inverse methods also operate after crystallographic preprocessing\. GraPhAI phases indexed three\-dimensional reflection amplitudes, while hybrid physics–ML refinement optimizes an existing structural model against integrated per\-HKL intensities using differentiable dynamical simulation\(Melgalvis and Rekis[2026](https://arxiv.org/html/2608.06448#bib.bib28); Maliket al\.[2026](https://arxiv.org/html/2608.06448#bib.bib29)\)\. Conventional continuous\-rotation 3D ED likewise combines indexing and integration with structure solution and refinement\(Klaret al\.[2023](https://arxiv.org/html/2608.06448#bib.bib30)\)\. These methods complement ED\-CSP but solve phasing or refinement after indexing, rather than generation from unindexed detector\-plane spot lists\.

## Method

### Problem Setting and Inputs

Given a compositionAAandKKED viewsS1:KS\_\{1:K\}, ED\-CSP models candidate latticesLLand fractional coordinatesFFthroughpθ​\(L,F∣A,S1:K\)p\_\{\\theta\}\(L,F\\mid A,S\_\{1:K\}\)\. Each view is a variable\-length spot listSk=\{\(qx,qy,log⁡\(1\+I\)\)j\}jS\_\{k\}=\\\{\(q\_\{x\},q\_\{y\},\\log\(1\+I\)\)\_\{j\}\\\}\_\{j\}, padded only at batching time\. The composition fixes the atom types and atom count, while the ED branch receives only the sampled detector\-plane spot lists\. Indexed Miller labels, zone\-axis vectors, and crystallographic labels are not provided to the ED branch\. Figure[2](https://arxiv.org/html/2608.06448#Sx3.F2)summarizes the resulting pipeline\.

![Refer to caption](https://arxiv.org/html/2608.06448v1/assets/edcsp_pipeline.png)Figure 2:ED\-CSP pretraining, inference, and optional post\-processing\.The optional uMLIP branch receives generated candidates but neither ED observations nor ground truth; structure matching is evaluation only\.
### Sparse Multi\-View ED Encoder

The CHILI\-100K ED\-CSP runs reported here use an ED encoder adapted from the EDiffCrystals PE\-AG\-GMoE sparse diffraction backbone\(Nathaniet al\.[2026](https://arxiv.org/html/2608.06448#bib.bib26)\)\. Each ED view is encoded by a shared PE\-AG\-GMoE\-style graph\-attention module over the raw py4DSTEM detector\-plane spot list\. The resulting per\-view representations are then aggregated across the sampled ED views with a mean–max pooling head and projected into the conditioning space of the generator\.

### Training and Initialization

The ED encoder is optimized jointly with the periodic generator in all reported ED\-CSP runs\. For the CHILI\-only checkpoint, its initial weights come from a separate ED–structure contrastive model that aligns paired multi\-view observations and crystal structures in a normalized embedding space; the transferred ED encoder remains trainable and the structure encoder is discarded\. Following the PXRDGen training design, this stage serves two roles: ED\-to\-structure retrieval benchmarks the learned diffraction representation, and the pretrained ED weights initialize the conditioning encoder before generator training\(Liet al\.[2025](https://arxiv.org/html/2608.06448#bib.bib16)\)\. It is particularly useful for screening encoder and view\-aggregation choices: retrieval isolates the ED representation from composition conditioning and requires neither periodic\-generator optimization nor iterative structure sampling\. We evaluate the first role directly; without a matched randomly initialized generator, the downstream contribution of the second is not isolated\. The registry\-1M experiment instead warm\-starts the complete registry\-trained ED\-CSP model before CHILI finetuning, so its gain measures full\-model transfer rather than ED\-CL alone\. This completed one\-million\-structure precursor predates the final ED\-CS filtering and is therefore reported separately from the 4\.85\-million\-structure corpus\.

### Periodic Flow Generation

A six\-layer CSPNet\-style periodic graph decoder\(Jiaoet al\.[2023](https://arxiv.org/html/2608.06448#bib.bib31); Liet al\.[2025](https://arxiv.org/html/2608.06448#bib.bib16)\)conditions on flow time, the current lattice and fractional coordinates, and the aggregated ED state to predict lattice and periodic\-coordinate vector fields\. Training minimizes their weighted mean\-squared errors,ℒ=ℒlat\+100​ℒcoord\\mathcal\{L\}=\\mathcal\{L\}\_\{\\mathrm\{lat\}\}\+100\\mathcal\{L\}\_\{\\mathrm\{coord\}\}\. ED\-CSP and PXRDGen share the same CSPFlow/CSPNet structure generator but use modality\-specific encoders; their comparison therefore evaluates the complete ED\- and PXRD\-conditioned systems rather than isolating diffraction modality alone\.

## Experimental Setup

### Datasets and ED Simulation

We use CHILI\-100K, a KDD graph\-ML benchmark derived from experimentally determined inorganic structures\(Friis\-Jensenet al\.[2024](https://arxiv.org/html/2608.06448#bib.bib22)\)\. We additionally construct Electron Diffraction Crystal Structures \(ED\-CS\) v1, a frozen snapshot comprising 4,852,131 structures selected from AFLOW, Alexandria, the Crystallography Open Database, GNoME, Materials Project, OQMD, and JARVIS\-DFT\(Curtaroloet al\.[2012](https://arxiv.org/html/2608.06448#bib.bib7); Gražuliset al\.[2012](https://arxiv.org/html/2608.06448#bib.bib5); Schmidtet al\.[2024](https://arxiv.org/html/2608.06448#bib.bib6); Merchantet al\.[2023](https://arxiv.org/html/2608.06448#bib.bib8); Jainet al\.[2013](https://arxiv.org/html/2608.06448#bib.bib1); Saalet al\.[2013](https://arxiv.org/html/2608.06448#bib.bib9); Choudharyet al\.[2020](https://arxiv.org/html/2608.06448#bib.bib10)\)\. ED\-CS v1 is a curated construction snapshot, not an exhaustive mirror of any upstream repository; each source contribution is defined by its versioned eligibility, simulation, deduplication, and exclusion manifests, and later additions require a separately versioned expansion\. The v1 candidate snapshot contains entries with at most 100 sites; we merge canonical identifiers, deduplicate same\-formula structures with StructureMatcher, require a certified payload with at least ten valid simulated views, and exclude every match to CHILI\-100K under the same strict matcher settings\. The resulting certificate contains zero CHILI matches; source counts and the full construction record are reported in the supplementary material\. For every retained identifier, a reconstruction registry links the stored canonical structure, exact orientations, and ED arrays to record\-level hashes; all 4,852,131 entries pass this technical completeness check\. The Code and Data Supplement provides the construction code, provenance and integrity schemas, and a source\-stratified subset of 256 real ED\-CS records; source\-specific redistribution terms for the full staged corpus are detailed in the supplementary material\. Table[1](https://arxiv.org/html/2608.06448#Sx4.T1)records the attrition at each certified construction stage\.

Table 1:ED\-CS construction record; bold marks the final v1 snapshot\.ED\-CS v1 is a frozen curated snapshot, not an exhaustive export of its upstream repositories\. The final strict StructureMatcher certificate reports zero CHILI matches and zero matcher errors\.

For each retained structure, we precompute dynamical ED spot patterns using the py4DSTEM simulation pipeline and parameters adopted by EDiffCrystals\(Savitzkyet al\.[2021](https://arxiv.org/html/2608.06448#bib.bib24); Gleasonet al\.[2024](https://arxiv.org/html/2608.06448#bib.bib25); Nathaniet al\.[2026](https://arxiv.org/html/2608.06448#bib.bib26)\), at300keV300\\text\{\\,\}\\mathrm\{keV\}electron energy,20nm20\\text\{\\,\}\\mathrm\{nm\}thickness, and a reciprocal\-space cutoff of2\.0Å−12\.0\\text\{\\,\}\{\\mathrm\{\\text\{Å\}\}\}^\{\-1\}\. This multi\-orientation protocol follows recent ML electron\-diffraction benchmarks based on py4DSTEM or Bloch\-wave point\-list patterns\(Gleasonet al\.[2024](https://arxiv.org/html/2608.06448#bib.bib25); Nathaniet al\.[2026](https://arxiv.org/html/2608.06448#bib.bib26)\)\. Simulation starts from ten random orientations\. A throughput\-oriented online policy extends a case to at most 100 views only when its ten\-view pilot runtime falls below a break\-even threshold estimated from recent simulation throughput and extension cost; the complete rule is given in the supplementary material\. A structure is retained only when at least ten valid views are available\. Each accepted view has at least ten spots before retaining its 16 strongest intensities\. The completed million\-structure transfer curriculum uses an earlier pool drawn from Materials Project, the Crystallography Open Database, and Alexandria while retaining source provenance; it is a precursor rather than a claimed subset of the final ED\-CS corpus\. The main CHILI benchmark uses 16,611 training structures and the same 2,075 ED\-valid held\-out queries across its reported comparisons\. ED\-CSP samples ten views per structure and retains the 16 strongest spots per view\.

### Training Protocol

The CHILI\-100K protocol is frozen before model evaluation\. ED\-CSP is trained on the CHILI\-100K train split with Adam optimization, gradient clipping, and mixed precision\. Validation match rate with one candidate selects the reported ED\-CSP checkpoints, with the policy fixed before test\-set evaluation\.

### Baselines and Comparison Settings

We evaluate three comparison settings with distinct inputs\. PXRDGen, XRDSol, and deCIFer provide same\-dataset diffraction\-conditioned generative comparisons; ED library matching measures finite\-database retrieval with explicit coverage; and Superflip/EDMA measures reconstruction when simulator\-indexed reflections are supplied\(Palatinus and Chapuis[2007](https://arxiv.org/html/2608.06448#bib.bib14); Liet al\.[2025](https://arxiv.org/html/2608.06448#bib.bib16); Yuet al\.[2026](https://arxiv.org/html/2608.06448#bib.bib19); Johansenet al\.[2026](https://arxiv.org/html/2608.06448#bib.bib17)\)\. For the matched deCIFer adaptation, we train from scratch on the same 16,611 CHILI structures, provide only composition and a clean PXRD profile at inference, and select the checkpoint by validation loss before test evaluation\. For the XRDSol adaptation, we use the same CHILI split, its published 1,000\-step training budget, five full 1,000\-step diffusion samples, and its native PXRD, composition, and ground\-truth unit\-cell inputs\.

The retrieval library is restricted to train and validation materials; test structures are excluded\. It searches all 18,688 ED\-valid train/validation reference structures and returns five nearest neighbors\. Radial retrieval compares normalized reciprocal\-radius histograms; canonical\-Chamfer retrieval compares sparse two\-dimensional spot sets after radial prefiltering by matching each query view to its closest library view\. Formula\-aware variants restrict candidates by anonymous formula, chemical system, or exact reduced formula before ranking\. If no train/validation candidate passes a formula filter, that protocol has no valid candidate for the query\.

The Palatinus\-style control exports simulator\-indexed reflection lists to Superflip/EDMA and uses five fixed reconstruction restarts\. It is evaluated separately from methods that receive detector\-plane spot lists\.

### Evaluation Protocol

Structure generation is evaluated with the shared structural matcher implemented through pymatgen\(Onget al\.[2013](https://arxiv.org/html/2608.06448#bib.bib23)\), with the same matcher settings for ED\-CSP, library retrieval, and all evaluations\. We report match rate \(MR\): the fraction of test materials for which at least one candidate structure matches the ground truth under this fixed matcher\. For CHILI\-100K, ED\-CSP, PXRDGen, and deCIFer are reported at both MR@1 and MR@5; XRDSol is reported at MR@5 over five independent samples; ED library matching is reported at MR@1 and MR@5 for the exact\-formula control and at MR@5 for the unfiltered full\-coverage control; Superflip/EDMA is reported as MR@5 over five fixed restarts\. We compute each confidence interval \(CI\) by non\-parametric bootstrap over held\-out materials and assess paired significance with a sign\-flip test on the per\-query ED\-CSP\-minus\-library outcomes\.

## Results

Table[2](https://arxiv.org/html/2608.06448#Sx5.T2)summarizes the CHILI\-100K benchmark\. Its blocks answer different questions under a common test split and matcher; they are not an input\-equivalent leaderboard\.

Table 2:CHILI\-100K benchmark on the same 2,075 held\-out queries\. Coverage is the fraction of queries for which the method has a valid input candidate set; dashes denote unmeasured metrics\. Rows are grouped by their available inputs and are not a single input\-equivalent leaderboard\. Bold marks the strongest measured MR within the first two ED\-CSP/generator blocks; control rows are not ranked\.![Refer to caption](https://arxiv.org/html/2608.06448v1/x1.png)Figure 3:Effect of pretraining scale and exact\-formula retrieval availability\.\(a\) Sequential ED\-to\-structure stages: each nested\-pool checkpoint resumes its converged predecessor, so the points are not independent fits\. \(b\) Registry\-initialized ED\-CSP recovery with and without an exact\-formula train/validation candidate; exact\-formula lookup has no candidate in the absent stratum\. Error bars are query\-bootstrap 95% CIs\.### Registry Scaling and Signal Use

Full\-model registry\-1M initialization is strongest on the aligned ED\-valid split: MR@1 reaches51\.66%51\.66\\text\{\\,\}\\mathrm\{\\char 37\\relax\}versus42\.12%42\.12\\text\{\\,\}\\mathrm\{\\char 37\\relax\}, and MR@5 reaches66\.27%66\.27\\text\{\\,\}\\mathrm\{\\char 37\\relax\}versus57\.49%57\.49\\text\{\\,\}\\mathrm\{\\char 37\\relax\}for CHILI\-only initialization\. The paired gains are9\.549\.54percentage points \(95% CI \[7\.577\.57,11\.5211\.52\]\) at MR@1 and8\.788\.78percentage points \(\[7\.047\.04,10\.5110\.51\]\) at MR@5\. Registry\-1M ED\-CSP achieves66\.27%66\.27\\text\{\\,\}\\mathrm\{\\char 37\\relax\}MR@5 for the designated headline seed; across three inference seeds, mean MR@5 is66\.28%66\.28\\text\{\\,\}\\mathrm\{\\char 37\\relax\}, with a sample standard deviation of0\.120\.12percentage points\. We retain the reference CHILI\-100K split for comparability; a post\-hoc sensitivity analysis excluding 888 queries with a strict StructureMatcher near\-duplicate still ranks registry\-1M ED\-CSP first at38\.84%38\.84\\text\{\\,\}\\mathrm\{\\char 37\\relax\}/53\.92%53\.92\\text\{\\,\}\\mathrm\{\\char 37\\relax\}, versus27\.38%27\.38\\text\{\\,\}\\mathrm\{\\char 37\\relax\}/42\.46%42\.46\\text\{\\,\}\\mathrm\{\\char 37\\relax\}for CHILI\-only ED\-CSP and23\.67%23\.67\\text\{\\,\}\\mathrm\{\\char 37\\relax\}/38\.75%38\.75\\text\{\\,\}\\mathrm\{\\char 37\\relax\}for PXRDGen\. The separate representation diagnostic in Figure[3](https://arxiv.org/html/2608.06448#Sx5.F3)a follows a progressive nested\-pool curriculum: each scale resumes the converged checkpoint from the preceding scale, expands the training pool, and continues until validation retrieval plateaus\. Its monotonic gains show that adding registry structures improves representation retrieval along this curriculum\. Figure[3](https://arxiv.org/html/2608.06448#Sx5.F3)b separately evaluates full\-model ED\-CSP transfer\.

The full\-split signal ablation changes only the ED input\. Removing the ED spots reduces MR@1 from51\.66%51\.66\\text\{\\,\}\\mathrm\{\\char 37\\relax\}to17\.35%17\.35\\text\{\\,\}\\mathrm\{\\char 37\\relax\}, a paired drop of34\.3134\.31percentage points \(95% CI \[32\.1432\.14,36\.4836\.48\]\)\. On the stricter 67\-query subset with a non\-isomorphic same\-formula donor, swapping in donor ED reduces mean MR@5 by22\.0922\.09percentage points across five generation seeds \(95% CI \[10\.4510\.45,34\.0334\.03\]\)\. A separately optimized converged composition\-only reference reaches50\.94%50\.94\\text\{\\,\}\\mathrm\{\\char 37\\relax\}MR@5, compared with57\.49%57\.49\\text\{\\,\}\\mathrm\{\\char 37\\relax\}for CHILI\-only ED\-CSP on the same queries and candidate budget; because their optimization and sampling layouts differ, this6\.556\.55\-point gap is descriptive rather than a paired causal estimate\. Together, the converged reference and input interventions show that ED\-CSP uses ED beyond composition and learned priors on the full split, and query\-specific ED on the same\-formula donor subset\.

### Input Sensitivity

In a separate single\-seed evaluation, we apply paired corruptions to the cached CHILI\-only ED inputs while fixing the checkpoint, queries, compositions, matcher, five\-candidate budget, and inference setting\. These interventions measure sensitivity to perturbed model inputs, not transfer to a new physical simulation regime\. Table[3](https://arxiv.org/html/2608.06448#Sx5.T3)shows a graded response to missing spots: 10% dropout lowers MR@5 by2\.072\.07percentage points, while 25% lowers it by7\.717\.71percentage points\. Intensity noise atσ=0\.50\\sigma=0\.50produces a2\.022\.02\-percentage\-point decrease, whereas perturbing stored view angles by up to5∘5^\{\\circ\}has no resolved effect under this cached\-input protocol\.

Table 3:CHILI\-100K sensitivity to cached ED input corruptions\.Single\-seed cached\-input evaluation\. All deltas are paired percentage\-point changes from this table’s 57\.40% clean row, not from the 57\.49% primary evaluation\.

A separate single\-seed detector\-frame intervention starts from the registry\-1M checkpoint and fine\-tunes with random shared in\-plane rotations\. It raises MR@5 under shared and independent rotations by7\.287\.28and5\.935\.93percentage points, leaving a1\.351\.35\-percentage\-point gap to clean inputs in both cases, at a2\.312\.31\-percentage\-point clean\-input cost \(Supplementary Material\)\.

### Benchmark Comparisons

The exact\-formula library control reaches42\.60%42\.60\\text\{\\,\}\\mathrm\{\\char 37\\relax\}MR@1 and43\.47%43\.47\\text\{\\,\}\\mathrm\{\\char 37\\relax\}MR@5 at50\.65%50\.65\\text\{\\,\}\\mathrm\{\\char 37\\relax\}coverage, whereas unfiltered full\-coverage retrieval reaches18\.31%18\.31\\text\{\\,\}\\mathrm\{\\char 37\\relax\}MR@5\. The apparent strength of formula\-filtered lookup is therefore tied to analogue availability: among the 1,051 queries with a same\-formula train/validation candidate, retrieval reaches85\.82%85\.82\\text\{\\,\}\\mathrm\{\\char 37\\relax\}MR@5 versus78\.69%78\.69\\text\{\\,\}\\mathrm\{\\char 37\\relax\}for registry\-pretrained ED\-CSP; on the remaining 1,024 queries it has no candidate, while ED\-CSP reaches53\.52%53\.52\\text\{\\,\}\\mathrm\{\\char 37\\relax\}\. This stratification separates phase lookup from out\-of\-library generation rather than averaging the two regimes into a misleading leaderboard\. Using exact\-formula retrieval when it has coverage and ED\-CSP otherwise reaches69\.88%69\.88\\text\{\\,\}\\mathrm\{\\char 37\\relax\}MR@5 at the same five\-candidate budget, a paired gain of3\.613\.61percentage points over ED\-CSP \(95% CI \[2\.172\.17,5\.065\.06\]\)\.

The indexed\-reflection Superflip/EDMA control reaches10\.51%10\.51\\text\{\\,\}\\mathrm\{\\char 37\\relax\}MR@5, with99\.76%99\.76\\text\{\\,\}\\mathrm\{\\char 37\\relax\}valid\-CIF query coverage and96\.40%96\.40\\text\{\\,\}\\mathrm\{\\char 37\\relax\}valid\-candidate coverage\. It consumes simulator\-indexed reflections rather than detector\-plane spots; its role is to measure reconstruction performance when indexed reflections are provided\.

Under CHILI\-only training and identical query IDs and candidate budgets, ED\-CSP reaches42\.12%42\.12\\text\{\\,\}\\mathrm\{\\char 37\\relax\}MR@1 and57\.49%57\.49\\text\{\\,\}\\mathrm\{\\char 37\\relax\}MR@5\. PXRDGen reaches34\.02%34\.02\\text\{\\,\}\\mathrm\{\\char 37\\relax\}and52\.92%52\.92\\text\{\\,\}\\mathrm\{\\char 37\\relax\}, respectively\. At MR@5, ED\-CSP alone solves 241 queries, PXRDGen alone solves 146, both solve 952, and both miss 736\. The matched autoregressive deCIFer adaptation reaches23\.28%23\.28\\text\{\\,\}\\mathrm\{\\char 37\\relax\}MR@1 and34\.99%34\.99\\text\{\\,\}\\mathrm\{\\char 37\\relax\}MR@5\. ED\-CSP’s paired gains over deCIFer are18\.8418\.84percentage points \(95% CI \[16\.4816\.48,21\.1621\.16\]\) at MR@1 and22\.5122\.51percentage points \(\[20\.1920\.19,24\.8224\.82\]\) at MR@5; at MR@5, ED\-CSP alone solves 584 queries and deCIFer alone solves 117\. XRDSol, retrained on the same CHILI split and additionally given the ground\-truth unit cell, reaches16\.00%16\.00\\text\{\\,\}\\mathrm\{\\char 37\\relax\}MR@5; CHILI\-only ED\-CSP’s paired gain is41\.4941\.49percentage points \(95% CI \[39\.1839\.18,43\.8143\.81\]\)\. These paired results compare specific systems rather than establish intrinsic ED superiority: ED\-CSP receives multiple ED spot\-list views through its encoder; PXRDGen uses a one\-dimensional powder profile with a convolutional neural network \(CNN\) encoder but shares ED\-CSP’s CSPFlow/CSPNet generator; XRDSol additionally receives the unit cell; and deCIFer uses an autoregressive CIF\-generation architecture\.

### Multi\-View Diagnostic

The supplementary train\-time view\-count diagnostic keeps a 100\-view simulation pool fixed and retrains an ED–structure retrieval encoder for each input count\. Increasing the consumed views from one to twenty improves Top\-5 ED\-to\-structure retrieval from1\.95%1\.95\\text\{\\,\}\\mathrm\{\\char 37\\relax\}to13\.38%13\.38\\text\{\\,\}\\mathrm\{\\char 37\\relax\}, supporting multi\-view representation learning in this encoder\-level benchmark without claiming a downstream generation optimum\.

![Refer to caption](https://arxiv.org/html/2608.06448v1/x2.png)Figure 4:Train\-timeKmodelK\_\{\\mathrm\{model\}\}ablation\.Single\-seed checkpoints; error bars are exact binomial 95% CIs over 1,024 queries and exclude training\-run variation\.
### Post\-Generation Relaxation

We test whether a target\-free interatomic potential can stabilize and rank a frozen five\-candidate CHILI\-only ED\-CSP payload\. ORB\-v3\(Rhodeset al\.[2025](https://arxiv.org/html/2608.06448#bib.bib33)\), MACE\-MPA\-0\(Batatia and others[2025](https://arxiv.org/html/2608.06448#bib.bib2)\), and eSEN\-30M\-OAM\(Fuet al\.[2025](https://arxiv.org/html/2608.06448#bib.bib4); Barroso\-Luqueet al\.[2024](https://arxiv.org/html/2608.06448#bib.bib3)\)perform up to 100 FIRE steps, while CHGNet\(Denget al\.[2023](https://arxiv.org/html/2608.06448#bib.bib34)\)performs 30 relaxation steps; all four relax the cell and rank candidates by final energy per atom\. None of the potentials receives the ED observations or ground\-truth structure\.

Table 4:Post\-generation relaxation and energy ranking on CHILI\-100K using a separately sampled, fixed candidate set\. Deltas are paired gains over that set in percentage points; bold marks the validation\-selected potential\.Every paired delta uses the same fixed candidate set, sampled separately from the candidates used for the 42\.12%/57\.49% primary evaluation\.

ORB\-v3 improves top\-1 recovery by13\.1613\.16percentage points \(95% CI \[11\.3711\.37,14\.9414\.94\]\) and the relaxed candidate\-pool MR@5 by4\.634\.63percentage points \(\[3\.423\.42,5\.885\.88\]\)\. MACE\-MPA\-0 provides a near\-identical independent check, with gains of13\.2013\.20and4\.534\.53percentage points \(\[11\.4711\.47,14\.9914\.99\] and \[3\.333\.33,5\.735\.73\]\), respectively; ORB\-v3 remains the validation\-selected potential\. eSEN\-30M\-OAM independently gives gains of12\.8712\.87and4\.104\.10percentage points \(\[11\.0811\.08,14\.6514\.65\] and \[2\.892\.89,5\.355\.35\]\), respectively, without exceeding MACE\-MPA\-0 or ORB\-v3\. CHGNet independently gives gains of10\.6510\.65and3\.283\.28percentage points \(\[8\.968\.96,12\.3412\.34\] and \[2\.272\.27,4\.344\.34\]\), respectively\. The separately sampled frozen payload supplies its own raw baseline, which differs slightly from the designated headline evaluation\. The gains show that both candidate ordering and local geometry limit recovery, while the61\.69%61\.69\\text\{\\,\}\\mathrm\{\\char 37\\relax\}relaxed\-pool ceiling leaves substantial room for ED\-aware refinement rather than energy\-only post\-processing\.

## Discussion and Limitations

The signal\-use interventions show that the generator responds to query\-specific diffraction geometry rather than treating ED as an optional auxiliary input\. Together with the coverage\-stratified retrieval results, this supports generative recovery as a complement to analogue lookup and indexed\-reflection workflows\.

The candidate analyses identify generation quality and selection as immediate bottlenecks: relaxation and energy ranking improve top\-1 recovery, yet the remaining candidate\-pool ceiling indicates room for ED\-aware refinement\. A natural next step is differentiable dynamical Bloch\-wave refinement of generated candidates against observed ED views, jointly regularized by crystallographic or learned energy priors\(Maliket al\.[2026](https://arxiv.org/html/2608.06448#bib.bib29)\)\. Beginning with indexed, orientation\-aware upper bounds, this would provide an ED\-specific post\-generation analogue to Rietveld refinement while modeling thickness\-dependent intensities\. The current ED\-CSP generator consumes ten ED views, while the encoder\-level diagnostic in Figure[4](https://arxiv.org/html/2608.06448#Sx5.F4)improves Top\-5 retrieval from7\.81%7\.81\\text\{\\,\}\\mathrm\{\\char 37\\relax\}atKmodel=10K\_\{\\mathrm\{model\}\}=10to13\.38%13\.38\\text\{\\,\}\\mathrm\{\\char 37\\relax\}atKmodel=20K\_\{\\mathrm\{model\}\}=20, suggesting that the present input regime does not saturate multi\-view representation learning\. Evaluating larger view sets in full generator training is therefore a promising direction\. ED\-CS makes the scale and provenance of this direction explicit, while its simulation cost motivates adaptive allocation of orientations\. The registry results also motivate continuing to expand the unique\-structure pool, which already spans millions of structures, while reducing its ED simulation cost\. Rather than assigning every structure the same simulation budget, future work should test whether Křivovichev\-style Shannon structural complexity\(Krivovichev[2014](https://arxiv.org/html/2608.06448#bib.bib13)\)can help estimate the number of informative orientations required per structure\. The principal external\-validity gap remains the transition from calibrated simulated point lists with known composition to experimental 3D ED, where detector calibration, uncertain spot finding, background, missing reflections, thickness variation, indexing, and expert refinement all affect the observed data\.

## Conclusion

ED\-CSP predicts lattices and fractional atomic coordinates from known composition and sparse simulated multi\-view ED\. The benchmark separates generation from finite\-library retrieval and indexed\-reflection preprocessing, while registry\-scale full\-model transfer further improves recovery\. It provides a reproducible basis for developing ED\-consistent candidate ranking and refinement\.

## Acknowledgments

The authors gratefully acknowledge GENCI/IDRIS for providing high\-performance computing resources on the Jean Zay supercomputer, which supported the computational aspects of this work\. This work was also supported by MAIA \(“Maitrise des Applications de l’IA”\) project from alliance A2U \(Université d’Artois, UPJV et ULCO\) and Hauts\-de\-France \(HdF\) region\.

## References

- L\. Barroso\-Luque, M\. Shuaibi, X\. Fu, B\. M\. Wood, M\. Dzamba, M\. Gao, A\. Rizvi, C\. L\. Zitnick, and Z\. W\. Ulissi \(2024\)Open materials 2024 \(OMat24\) inorganic materials dataset and models\.arXiv preprint arXiv:2410\.12771\.Cited by:[Post\-Generation Relaxation](https://arxiv.org/html/2608.06448#Sx5.SSx5.p1.1)\.
- I\. Batatiaet al\.\(2025\)A foundation model for atomistic materials chemistry\.The Journal of Chemical Physics163\(18\),pp\. 184110\.External Links:[Document](https://dx.doi.org/10.1063/5.0297006)Cited by:[Post\-Generation Relaxation](https://arxiv.org/html/2608.06448#Sx5.SSx5.p1.1)\.
- K\. Choudhary, K\. F\. Garrity, A\. C\. E\. Reid, B\. DeCost, A\. J\. Biacchi,et al\.\(2020\)The joint automated repository for various integrated simulations \(JARVIS\) for data\-driven materials design\.npj Computational Materials6,pp\. 173\.External Links:[Document](https://dx.doi.org/10.1038/s41524-020-00440-1)Cited by:[Datasets and ED Simulation](https://arxiv.org/html/2608.06448#Sx4.SSx1.p1.1)\.
- S\. Curtarolo, W\. Setyawan, S\. Wang, J\. Xue, K\. Yang, R\. H\. Taylor, L\. J\. Nelson, G\. L\. W\. Hart, S\. Sanvito, M\. Buongiorno\-Nardelli, N\. Mingo, and O\. Levy \(2012\)AFLOWLIB\.ORG: a distributed materials properties repository from high\-throughput ab initio calculations\.Computational Materials Science58,pp\. 227–235\.External Links:[Document](https://dx.doi.org/10.1016/j.commatsci.2012.02.002)Cited by:[Datasets and ED Simulation](https://arxiv.org/html/2608.06448#Sx4.SSx1.p1.1)\.
- B\. Deng, P\. Zhong, K\. Jun, J\. Riebesell, K\. Han, C\. J\. Bartel, and G\. Ceder \(2023\)CHGNet as a pretrained universal neural network potential for charge\-informed atomistic modelling\.Nature Machine Intelligence5\(9\),pp\. 1031–1041\.External Links:[Document](https://dx.doi.org/10.1038/s42256-023-00716-3)Cited by:[Post\-Generation Relaxation](https://arxiv.org/html/2608.06448#Sx5.SSx5.p1.1)\.
- S\. Edamadaka, K\. Ramesh, J\. Li, and R\. Gómez\-Bombarelli \(2026\)Atomistic language models understand and generate materials\.arXiv preprint arXiv:2606\.21395\.Cited by:[Diffraction\-conditioned generation\.](https://arxiv.org/html/2608.06448#Sx2.SS0.SSS0.Px1.p1.1)\.
- U\. Friis\-Jensen, F\. L\. Johansen, A\. S\. Anker, E\. B\. Dam, K\. M\. Ø\. Jensen, and R\. Selvan \(2024\)CHILI: chemically\-informed large\-scale inorganic nanomaterials dataset for advancing graph machine learning\.InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining,pp\. 4962–4973\.External Links:[Document](https://dx.doi.org/10.1145/3637528.3671538)Cited by:[Introduction](https://arxiv.org/html/2608.06448#Sx1.p4.1),[Datasets and ED Simulation](https://arxiv.org/html/2608.06448#Sx4.SSx1.p1.1)\.
- X\. Fu, B\. M\. Wood, L\. Barroso\-Luque, D\. S\. Levine, M\. Gao, M\. Dzamba, and C\. L\. Zitnick \(2025\)Learning smooth and expressive interatomic potentials for physical property prediction\.arXiv preprint arXiv:2502\.12147\.Cited by:[Post\-Generation Relaxation](https://arxiv.org/html/2608.06448#Sx5.SSx5.p1.1)\.
- M\. Gemmi, E\. Mugnaioli, T\. E\. Gorelik, U\. Kolb, L\. Palatinus, P\. Boullay, S\. Hovmöller, and J\. P\. Abrahams \(2019\)3D electron diffraction: the nanocrystallography revolution\.ACS Central Science5\(8\),pp\. 1315–1329\.External Links:[Document](https://dx.doi.org/10.1021/acscentsci.9b00394)Cited by:[Introduction](https://arxiv.org/html/2608.06448#Sx1.p1.1)\.
- S\. P\. Gleason, A\. Rakowski, S\. M\. Ribet, S\. E\. Zeltmann, B\. H\. Savitzky, M\. Henderson, J\. Ciston, and C\. Ophus \(2024\)Random forest prediction of crystal structure from electron diffraction patterns incorporating multiple scattering\.Physical Review Materials8,pp\. 093802\.External Links:[Document](https://dx.doi.org/10.1103/PhysRevMaterials.8.093802)Cited by:[Introduction](https://arxiv.org/html/2608.06448#Sx1.p2.1),[Sparse multi\-view ED learning\.](https://arxiv.org/html/2608.06448#Sx2.SS0.SSS0.Px2.p1.1),[Datasets and ED Simulation](https://arxiv.org/html/2608.06448#Sx4.SSx1.p2.3)\.
- S\. Gražulis, A\. Daškevič, A\. Merkys, D\. Chateigner, L\. Lutterotti, M\. Quirós, N\. R\. Serebryanaya, P\. Moeck, R\. T\. Downs, and A\. Le Bail \(2012\)Crystallography open database \(COD\): an open\-access collection of crystal structures and platform for world\-wide collaboration\.Nucleic Acids Research40\(D1\),pp\. D420–D427\.External Links:[Document](https://dx.doi.org/10.1093/nar/gkr900)Cited by:[Datasets and ED Simulation](https://arxiv.org/html/2608.06448#Sx4.SSx1.p1.1)\.
- G\. Guo, T\. L\. Saidi, M\. W\. Terban, M\. Valsecchi, S\. J\. L\. Billinge, and H\. Lipson \(2025\)Ab initio structure solutions from nanocrystalline powder diffraction data via diffusion models\.Nature Materials24,pp\. 1726–1734\.External Links:[Document](https://dx.doi.org/10.1038/s41563-025-02220-y)Cited by:[Introduction](https://arxiv.org/html/2608.06448#Sx1.p2.1),[Diffraction\-conditioned generation\.](https://arxiv.org/html/2608.06448#Sx2.SS0.SSS0.Px1.p1.1)\.
- A\. Jain, S\. P\. Ong, G\. Hautier, W\. Chen, W\. D\. Richards, S\. Dacek, S\. Cholia, D\. Gunter, D\. Skinner, G\. Ceder, and K\. A\. Persson \(2013\)Commentary: the materials project: a materials genome approach to accelerating materials innovation\.APL Materials1\(1\),pp\. 011002\.External Links:[Document](https://dx.doi.org/10.1063/1.4812323)Cited by:[Datasets and ED Simulation](https://arxiv.org/html/2608.06448#Sx4.SSx1.p1.1)\.
- R\. Jiao, W\. Huang, P\. Lin, J\. Han, P\. Chen, Y\. Lu, and Y\. Liu \(2023\)Crystal structure prediction by joint equivariant diffusion\.InAdvances in Neural Information Processing Systems,Vol\.36,pp\. 17464–17497\.External Links:[Document](https://dx.doi.org/10.52202/075280-0767)Cited by:[Periodic Flow Generation](https://arxiv.org/html/2608.06448#Sx3.SSx4.p1.1)\.
- F\. L\. Johansen, U\. Friis\-Jensen, E\. B\. Dam, K\. M\. Ø\. Jensen, R\. Mercado, and R\. Selvan \(2026\)deCIFer: crystal structure prediction from powder diffraction data using autoregressive language models\.Transactions on Machine Learning Research\.Cited by:[Introduction](https://arxiv.org/html/2608.06448#Sx1.p2.1),[Diffraction\-conditioned generation\.](https://arxiv.org/html/2608.06448#Sx2.SS0.SSS0.Px1.p1.1),[Baselines and Comparison Settings](https://arxiv.org/html/2608.06448#Sx4.SSx3.p1.1)\.
- P\. B\. Klar, Y\. Krysiak, H\. Xu, G\. Steciuk, J\. Cho, X\. Zou, and L\. Palatinus \(2023\)Accurate structure models and absolute configuration determination using dynamical effects in continuous\-rotation 3D electron diffraction data\.Nature Chemistry15,pp\. 848–855\.External Links:[Document](https://dx.doi.org/10.1038/s41557-023-01186-1)Cited by:[Indexed ED solution and refinement\.](https://arxiv.org/html/2608.06448#Sx2.SS0.SSS0.Px3.p1.1)\.
- S\. V\. Krivovichev \(2014\)Which inorganic structures are the most complex?\.Angewandte Chemie International Edition53\(3\),pp\. 654–661\.External Links:[Document](https://dx.doi.org/10.1002/anie.201304374)Cited by:[Discussion and Limitations](https://arxiv.org/html/2608.06448#Sx6.p2.4)\.
- Q\. Lai, F\. Xu, L\. Yao, Z\. Gao, S\. Liu, H\. Wang, S\. Lu, D\. He, L\. Wang, L\. Zhang, C\. Wang, and G\. Ke \(2025\)End\-to\-end crystal structure prediction from powder X\-ray diffraction\.Advanced Science12\(8\),pp\. 2410722\.External Links:[Document](https://dx.doi.org/10.1002/advs.202410722)Cited by:[Introduction](https://arxiv.org/html/2608.06448#Sx1.p2.1),[Diffraction\-conditioned generation\.](https://arxiv.org/html/2608.06448#Sx2.SS0.SSS0.Px1.p1.1)\.
- Q\. Li, M\. Guo, R\. Jiao, J\. Gao, F\. Xu, H\. Xue, W\. Zhang, W\. Huang, J\. Yan, L\. Zhang, C\. Wang, Z\. Yan, G\. Ke, W\. E, Z\. Tang, S\. Jin, and L\. Yao \(2026\)Experimental powder X\-ray diffraction crystal structure determination with RealPXRD\-Solver\.arXiv preprint arXiv:2603\.00965\.Cited by:[Diffraction\-conditioned generation\.](https://arxiv.org/html/2608.06448#Sx2.SS0.SSS0.Px1.p1.1)\.
- Q\. Li, R\. Jiao, L\. Wu, T\. Zhu, W\. Huang, S\. Jin, Y\. Liu, H\. Weng, and X\. Chen \(2025\)Powder diffraction crystal structure determination using generative models\.Nature Communications16,pp\. 7428\.External Links:[Document](https://dx.doi.org/10.1038/s41467-025-62708-8)Cited by:[Introduction](https://arxiv.org/html/2608.06448#Sx1.p2.1),[Diffraction\-conditioned generation\.](https://arxiv.org/html/2608.06448#Sx2.SS0.SSS0.Px1.p1.1),[Training and Initialization](https://arxiv.org/html/2608.06448#Sx3.SSx3.p1.1),[Periodic Flow Generation](https://arxiv.org/html/2608.06448#Sx3.SSx4.p1.1),[Baselines and Comparison Settings](https://arxiv.org/html/2608.06448#Sx4.SSx3.p1.1)\.
- S\. A\. Malik, T\. A\. S\. Doherty, B\. Colmey, S\. J\. Roberts, Y\. Gal, and P\. A\. Midgley \(2026\)Hybrid physics\-machine learning models for quantitative electron diffraction refinements\.Nature Communications17,pp\. 5056\.External Links:[Document](https://dx.doi.org/10.1038/s41467-026-71673-9)Cited by:[Indexed ED solution and refinement\.](https://arxiv.org/html/2608.06448#Sx2.SS0.SSS0.Px3.p1.1),[Discussion and Limitations](https://arxiv.org/html/2608.06448#Sx6.p2.4)\.
- D\. M\. Melgalvis and T\. Rekis \(2026\)GraPhAI: neural networks for solving centrosymmetric crystal structures\.Journal of the American Chemical Society148\(27\),pp\. 28754–28763\.External Links:[Document](https://dx.doi.org/10.1021/jacs.6c05607)Cited by:[Indexed ED solution and refinement\.](https://arxiv.org/html/2608.06448#Sx2.SS0.SSS0.Px3.p1.1)\.
- A\. Merchant, S\. Batzner, S\. S\. Schoenholz, M\. Aykol, G\. Cheon, and E\. D\. Cubuk \(2023\)Scaling deep learning for materials discovery\.Nature624,pp\. 80–85\.External Links:[Document](https://dx.doi.org/10.1038/s41586-023-06735-9)Cited by:[Datasets and ED Simulation](https://arxiv.org/html/2608.06448#Sx4.SSx1.p1.1)\.
- A\. Nathani, A\. R\. C\. McCray, Y\. Liu, H\. Ding, P\. Kazempoor, S\. Xu, C\. Ophus, and I\. Ghamarian \(2026\)Accelerating electron diffraction analysis using graph neural networks and attention mechanisms\.npj Computational Materials12,pp\. 56\.External Links:[Document](https://dx.doi.org/10.1038/s41524-025-01927-5)Cited by:[Introduction](https://arxiv.org/html/2608.06448#Sx1.p2.1),[Sparse multi\-view ED learning\.](https://arxiv.org/html/2608.06448#Sx2.SS0.SSS0.Px2.p1.1),[Sparse Multi\-View ED Encoder](https://arxiv.org/html/2608.06448#Sx3.SSx2.p1.1),[Datasets and ED Simulation](https://arxiv.org/html/2608.06448#Sx4.SSx1.p2.3)\.
- S\. P\. Ong, W\. D\. Richards, A\. Jain, G\. Hautier, M\. Kocher, S\. Cholia, D\. Gunter, V\. L\. Chevrier, K\. A\. Persson, and G\. Ceder \(2013\)Python Materials Genomics \(pymatgen\): a robust, open\-source python library for materials analysis\.Computational Materials Science68,pp\. 314–319\.External Links:[Document](https://dx.doi.org/10.1016/j.commatsci.2012.10.028)Cited by:[Evaluation Protocol](https://arxiv.org/html/2608.06448#Sx4.SSx4.p1.1)\.
- L\. Palatinus and G\. Chapuis \(2007\)SUPERFLIP—a computer program for the solution of crystal structures by charge flipping in arbitrary dimensions\.Journal of Applied Crystallography40\(4\),pp\. 786–790\.External Links:[Document](https://dx.doi.org/10.1107/S0021889807029238)Cited by:[Baselines and Comparison Settings](https://arxiv.org/html/2608.06448#Sx4.SSx3.p1.1)\.
- Q\. Peng, X\. Han, Z\. Wang, Y\. Hong, F\. Meng, Z\. Zhang, J\. Zhang, and X\. Zhao \(2026\)Diffraction\-native representation learning for automatic symmetry recognition and structural retrieval\.Advanced Functional Materials36\(50\),pp\. e75949\.External Links:[Document](https://dx.doi.org/10.1002/adfm.75949)Cited by:[Introduction](https://arxiv.org/html/2608.06448#Sx1.p2.1),[Sparse multi\-view ED learning\.](https://arxiv.org/html/2608.06448#Sx2.SS0.SSS0.Px2.p1.1)\.
- B\. Rhodes, S\. Vandenhaute, V\. Šimkus, J\. Gin, J\. Godwin, T\. Duignan, and M\. Neumann \(2025\)Orb\-v3: atomistic simulation at scale\.arXiv preprint arXiv:2504\.06231\.Cited by:[Post\-Generation Relaxation](https://arxiv.org/html/2608.06448#Sx5.SSx5.p1.1)\.
- J\. E\. Saal, S\. Kirklin, M\. Aykol, B\. Meredig, and C\. Wolverton \(2013\)Materials design and discovery with high\-throughput density functional theory: the open quantum materials database \(OQMD\)\.JOM65,pp\. 1501–1509\.External Links:[Document](https://dx.doi.org/10.1007/s11837-013-0755-4)Cited by:[Datasets and ED Simulation](https://arxiv.org/html/2608.06448#Sx4.SSx1.p1.1)\.
- B\. H\. Savitzky, S\. E\. Zeltmann, L\. A\. Hughes, H\. G\. Brown, S\. Zhao, P\. M\. Pelz, T\. C\. Pekin, E\. S\. Barnard, J\. Donohue, L\. Rangel DaCosta, E\. Kennedy, Y\. Xie, M\. T\. Janish, M\. M\. Schneider, P\. Herring, C\. Gopal, A\. Anapolsky, R\. Dhall, K\. C\. Bustillo, P\. Ercius, M\. C\. Scott, J\. Ciston, A\. M\. Minor, and C\. Ophus \(2021\)py4DSTEM: a software package for four\-dimensional scanning transmission electron microscopy data analysis\.Microscopy and Microanalysis27\(4\),pp\. 712–743\.External Links:[Document](https://dx.doi.org/10.1017/S1431927621000477)Cited by:[Introduction](https://arxiv.org/html/2608.06448#Sx1.p1.1),[Datasets and ED Simulation](https://arxiv.org/html/2608.06448#Sx4.SSx1.p2.3)\.
- J\. Schmidt, T\. F\. T\. Cerqueira, A\. H\. Romero, A\. Loew, F\. Jäger, H\. Wang, S\. Botti, and M\. A\. L\. Marques \(2024\)Improving machine\-learning models in materials science through large datasets\.Materials Today Physics48,pp\. 101560\.External Links:[Document](https://dx.doi.org/10.1016/j.mtphys.2024.101560)Cited by:[Datasets and ED Simulation](https://arxiv.org/html/2608.06448#Sx4.SSx1.p1.1)\.
- N\. Segal, M\. Li, B\. K\. Miller, and R\. Gómez\-Bombarelli \(2026\)XRDiff: crystal structure prediction from powder X\-ray diffraction data using diffusion models\.arXiv preprint arXiv:2606\.14003\.Cited by:[Introduction](https://arxiv.org/html/2608.06448#Sx1.p1.1),[Introduction](https://arxiv.org/html/2608.06448#Sx1.p2.1),[Diffraction\-conditioned generation\.](https://arxiv.org/html/2608.06448#Sx2.SS0.SSS0.Px1.p1.1)\.
- J\. Unge, B\. L\. Nannenga, A\. G\. Oliver, and T\. Gonen \(2025\)Standards for MicroED\.Acta Crystallographica Section C81,pp\. 376–390\.External Links:[Document](https://dx.doi.org/10.1107/S2053229625004875)Cited by:[Introduction](https://arxiv.org/html/2608.06448#Sx1.p1.1)\.
- D\. Yu, Z\. Zhu, F\. Leng, and Y\. Zhu \(2026\)Equivariant diffusion solution for inorganic crystal structure determination from powder X\-ray diffraction data\.Nature Communications17,pp\. 3274\.External Links:[Document](https://dx.doi.org/10.1038/s41467-026-70035-9)Cited by:[Diffraction\-conditioned generation\.](https://arxiv.org/html/2608.06448#Sx2.SS0.SSS0.Px1.p1.1),[Baselines and Comparison Settings](https://arxiv.org/html/2608.06448#Sx4.SSx3.p1.1)\.

Similar Articles