PhononBench-MP40: a spectrum-resolved benchmark dataset for phonon stability

arXiv cs.AI Papers

Summary

PhononBench-MP40 is a benchmark dataset of phonon stability labels and spectra for over 46,000 Materials Project-derived crystals, designed to evaluate workflow-defined phonon stability and support materials screening.

arXiv:2607.22573v1 Announce Type: new Abstract: Imaginary phonon modes remain a practical bottleneck in computational materials screening because otherwise plausible structures can be locally dynamically unstable under a chosen workflow. Here we present PhononBench-MP40, a spectrum-resolved benchmark dataset of Materials Project-derived crystals for workflow-defined phonon stability. The dataset starts from 47,969 MP40 workflow tasks and provides 46,899 completed records with paired stability labels and local phonopy YAML spectra, including 16,683 Stable records and 30,216 completed-phonon unstable records. A further 1,067 relaxation failures are reported separately rather than merged into the completed phonon denominator. The release centers on the local YAML spectrum: the stability label, the lowest sampled frequency and any threshold-dependent relabeling are derived from that spectrum. The dataset is openly available through Science Data Bank at https://doi.org/10.57760/sciencedb.38735. A companion GitHub repository provides the calculation code and lightweight access utilities. PhononBench-MP40 provides an auditable reference for workflow-defined stability classification, minimum-frequency analysis, threshold studies and failure-aware triage, while keeping the reference workflow, data schema and interpretation boundaries explicit.
Original Article
View Cached Full Text

Cached at: 07/28/26, 06:24 AM

# PhononBench-MP40: a spectrum-resolved benchmark dataset for phonon stability
Source: [https://arxiv.org/html/2607.22573](https://arxiv.org/html/2607.22573)
Wen\-Kao Li1,†, Ze\-Feng Gao1,†,∗, Zhong\-Yi Lu1,∗ 1School of Physics and Key Laboratory of Quantum State Construction and Manipulation \(Ministry of Education\), Renmin University of China, Beijing 100872, China †These authors contributed equally\. ∗e\-mail: zfgao@ruc\.edu\.cn; zlu@ruc\.edu\.cn

###### Abstract

Imaginary phonon modes remain a practical bottleneck in computational materials screening because otherwise plausible structures can be locally dynamically unstable under a chosen workflow\. Here we present PhononBench\-MP40, a spectrum\-resolved benchmark dataset of Materials Project\-derived crystals for workflow\-defined phonon stability\. The dataset starts from 47,969 MP40 workflow tasks and provides 46,899 completed records with paired stability labels and local phonopy YAML spectra, including 16,683 Stable records and 30,216 completed\-phonon unstable records\. A further 1,067 relaxation failures are reported separately rather than merged into the completed phonon denominator\. The release centers on the local YAML spectrum: the stability label, the lowest sampled frequency and any threshold\-dependent relabeling are derived from that spectrum\. The dataset is openly available through Science Data Bank at[https://doi\.org/10\.57760/sciencedb\.38735](https://doi.org/10.57760/sciencedb.38735)\. A companion GitHub repository provides the calculation code and lightweight access utilities\. PhononBench\-MP40 provides an auditable reference for workflow\-defined stability classification, minimum\-frequency analysis, threshold studies and failure\-aware triage, while keeping the reference workflow, data schema and interpretation boundaries explicit\.

Keywords:phonon stability; benchmark dataset; phonon spectra; materials informatics

## 1\.Introduction

Large\-scale crystal generation and screening have changed the rate at which candidate materials are proposed\[[1](https://arxiv.org/html/2607.22573#bib.bib1),[2](https://arxiv.org/html/2607.22573#bib.bib2),[3](https://arxiv.org/html/2607.22573#bib.bib3),[4](https://arxiv.org/html/2607.22573#bib.bib4)\]\. The bottleneck is no longer only how to generate structures, but how to decide which of them deserve expensive follow\-up\. Formation energy, charge balance, symmetry and structural validity are useful filters, but they do not by themselves establish local dynamical stability\. Within a specified interatomic\-potential and phonon workflow, the relevant question is sharper: does the relaxed structure have a sampled phonon mode with negative curvature, i\.e\. an imaginary frequency? This makes phonon stability a natural target for reusable computed data and for model evaluation\.

The field already has extensive computed\-data infrastructure\. Materials Project, AFLOW, OQMD, JARVIS, NOMAD and Materials Cloud organize large computational materials spaces\[[5](https://arxiv.org/html/2607.22573#bib.bib5),[6](https://arxiv.org/html/2607.22573#bib.bib6),[7](https://arxiv.org/html/2607.22573#bib.bib7),[8](https://arxiv.org/html/2607.22573#bib.bib8),[9](https://arxiv.org/html/2607.22573#bib.bib9),[10](https://arxiv.org/html/2607.22573#bib.bib10),[11](https://arxiv.org/html/2607.22573#bib.bib11)\]\. Workflow and analysis tools such as pymatgen, ASE, FireWorks, atomate and AiiDA support reproducible workflows\[[12](https://arxiv.org/html/2607.22573#bib.bib12),[13](https://arxiv.org/html/2607.22573#bib.bib13),[14](https://arxiv.org/html/2607.22573#bib.bib14),[15](https://arxiv.org/html/2607.22573#bib.bib15),[16](https://arxiv.org/html/2607.22573#bib.bib16)\]\. High\-throughput phonon resources and software have established practical routes to vibrational and thermal properties, including DFPT phonons, AFLOW\-AAPL, ShengBTE, ALAMODE, TDEP, SCAILD and recent anharmonic\-phonon databases\[[17](https://arxiv.org/html/2607.22573#bib.bib17),[18](https://arxiv.org/html/2607.22573#bib.bib18),[19](https://arxiv.org/html/2607.22573#bib.bib19),[20](https://arxiv.org/html/2607.22573#bib.bib20),[21](https://arxiv.org/html/2607.22573#bib.bib21),[22](https://arxiv.org/html/2607.22573#bib.bib22),[23](https://arxiv.org/html/2607.22573#bib.bib23),[24](https://arxiv.org/html/2607.22573#bib.bib24),[25](https://arxiv.org/html/2607.22573#bib.bib25)\]\. Most existing open phonon datasets contain spectra or vibrational properties for thousands to tens of thousands of materials\. PhononBench\-MP40 extends this scale to 46,899 completed MP\-derived phonon records, while preserving the spectral object behind each stability label\.

This design matters because a binary stability flag is convenient but incomplete\. A shallow soft mode, a deep imaginary branch and a failed calculation are different records for both physics and workflow auditing\. Collapsing them into one undifferentiated “unstable” class hides the evidence behind the decision and obscures the denominator used in later statistics\. PhononBench\-MP40 keeps these cases separate\. It pairs labels with local phonopy YAML spectra, fixes the completed label\+YAML cohort, and reports missing\-spectrum relaxation failures outside the completed phonon denominator\. The result is both a data release and a benchmark target: users can work with a fixed PhononBench\-defined label, but can also inspect the spectrum, extract a scalar frequency diagnostic, or move the stability threshold for their own application\.

## 2\.Dataset construction and workflow

Figure[1](https://arxiv.org/html/2607.22573#S2.F1)summarizes the construction\. The starting point is an audited MP40 task cohort containing 47,969 recorded tasks within the localNatoms≤40N\_\{\\mathrm\{atoms\}\}\\leq 40workflow scope\. Atom\-count windows drawn from Materials Project are common in crystal\-generation benchmarks, including MP\-20, MP\-21–40 and MP\-40 settings\[[1](https://arxiv.org/html/2607.22573#bib.bib1),[26](https://arxiv.org/html/2607.22573#bib.bib26),[27](https://arxiv.org/html/2607.22573#bib.bib27)\]\. Here the cutoff also has a direct phonon\-workflow meaning\. Under the2×2×22\\times 2\\times 2finite\-displacement supercell used in the workflow, a 40\-atom recorded cell corresponds to an up\-to\-320\-atom force\-evaluation cell before symmetry reduction\. The MP40 cutoff therefore defines a practical release domain: it includes many inorganic prototypes, ordered derivatives and multicomponent cells, while keeping the large\-scale finite\-displacement calculation computationally controlled\. It is a benchmark\-domain definition rather than a physical stability boundary\.

Each task follows the PhononBench route\. Structures are relaxed, finite\-displacement supercells are constructed, forces are evaluated with MatterSim/uMLIP, and phonopy is used to build the dynamical matrix and phonon outputs\[[28](https://arxiv.org/html/2607.22573#bib.bib28),[29](https://arxiv.org/html/2607.22573#bib.bib29),[30](https://arxiv.org/html/2607.22573#bib.bib30),[31](https://arxiv.org/html/2607.22573#bib.bib31)\]\. High\-symmetry paths are generated through the seekpath/spglib ecosystem\[[32](https://arxiv.org/html/2607.22573#bib.bib32),[33](https://arxiv.org/html/2607.22573#bib.bib33)\]\. The top row of Fig\.[1](https://arxiv.org/html/2607.22573#S2.F1)is therefore shown as an explicit workflow rather than as a stand\-alone label source: the reference target is a computed phonon result under a specified route\.

The primary release object is the local phonopy YAML file\. It contains the spectral information needed to reconstruct the dispersion, extract the lowest sampled phonon frequency and derive a Stable/unstable decision\. The label is thus a derived view of the spectrum\. A completed record in PhononBench\-MP40 is defined by the intersection of two local objects: a recovered workflow stability label and a matching local YAML output\. This rule yields 46,899 completed label\+YAML records from 47,966 audited local labels and 46,902 local YAML outputs\. Within the completed cohort, 16,683 records are Stable and 30,216 records are completed\-phonon unstable\. The audit also reports 1,067 relaxation failures separately, because no matching local YAML spectrum is available for those records\. This separation is essential for denominator\-aware reuse: a completed imaginary\-mode spectrum and a missing\-spectrum relaxation failure answer different scientific and computational questions\. Release\-file definitions, common fields, workflow settings and the label rule are provided in the Supplementary Information, Secs\. S1–S3\.

![Refer to caption](https://arxiv.org/html/2607.22573v1/x1.png)Figure 1:Workflow and release object of PhononBench\-MP40\.The dataset starts from 47,969 MP40 tasks and follows the PhononBench route through structural relaxation, finite\-displacement supercell construction, MatterSim/uMLIP force evaluation and phonopy processing\. The local YAML output is the primary release object because it enables dispersion reconstruction, minimum\-frequency extraction and stability\-label derivation\. The audit summary reports 47,966 audited local labels, 46,902 local YAML outputs, 46,899 completed label\+YAML entries, 16,683 Stable entries, 30,216 completed\-phonon unstable entries and 1,067 relaxation failures\.Table 1:Cohort accounting\.Counts for the MP40 audit and completed spectral cohort\.
## 3\.Data records

The released data are organized to make the Stable/unstable decision traceable\. The archival Science Data Bank record contains the completed label table, the failure\-accounting table, the reference split table, release metadata and the phonopy YAML spectra for completed records\. Table[2](https://arxiv.org/html/2607.22573#S3.T2)summarizes the data objects that define the public release\. The completed table is the main machine\-learning table: it links each completed record to a structure identifier, formula, space group, unit\-cell atom count, workflow label, completed label, YAML path,ωmin\\omega\_\{\\min\}value and split assignment\. The failure table is kept separate because a relaxation failure without a YAML spectrum is a workflow outcome, not a completed phonon\-instability label\.

This separation is important for reuse\. Users who need a fixed binary target can train on the completed label\+YAML cohort\. Users who need more physically resolved targets can reconstruct the band frequencies from the YAML files, calculateωmin\\omega\_\{\\min\}, or impose an application\-specific threshold\. Users interested in high\-throughput workflow robustness can include the 1,067 relaxation failures in a separate failure\-aware triage task\. The file schema, field definitions, split construction and example parsing workflow are given in the Supplementary Information, Secs\. S1, S2, S5 and S9\.

Table 2:Released data records\.The public release is defined by tabular metadata plus phonopy YAML spectra\. Counts refer to the release described in this work\.
## 4\.Spectra behind the labels

The completed stability label is derived from high\-symmetry\-path band frequencies\. A completed entry is labeled unstable if any sampled band frequency is below−10−3\-10^\{\-3\}THz; otherwise it is labeled Stable\. This threshold turns a phonon dispersion into a machine\-readable target, but the spectrum is the more informative object\. We useωmin\\omega\_\{\\min\}as the lowest sampled frequency along the released path\. For negative values,ωmin\\omega\_\{\\min\}measures the severity of an imaginary branch under the reference workflow; near\-zero values identify threshold\-sensitive cases\. It is a compact diagnostic of the YAML spectrum, not a replacement for the full dispersion and not a model\-independent stability law\.

Figure[2](https://arxiv.org/html/2607.22573#S4.F2)illustrates why the spectral object matters\. MgB2is labeled Stable within the adopted tolerance, withωmin=0\.0000\\omega\_\{\\min\}=0\.0000THz\. RbPbCl3is a completed\-phonon unstable example with a soft mode andωmin=−1\.4984\\omega\_\{\\min\}=\-1\.4984THz\. Such shallow imaginary modes are often the cases most sensitive to relaxation details, threshold choice, supercell construction and low\-symmetry distortions\. BaTiO3is also unstable, but with a much deeper imaginary branch,ωmin=−7\.9102\\omega\_\{\\min\}=\-7\.9102THz, indicating a stronger negative\-curvature signal under the same reference route\. These records share the same binary decision, yet the underlying spectra carry different physical and computational information\. A label table alone would hide this distinction; a YAML\-backed release keeps it available for inspection, threshold analysis and learning tasks\. The example metadata are listed in the Supplementary Information, Sec\. S4\.

![Refer to caption](https://arxiv.org/html/2607.22573v1/x2.png)Figure 2:Representative phonon spectra behind the stability labels\.Three YAML\-derived dispersions illustrate the label rule and the information retained beyond the binary decision\.a,MgB2is labeled Stable, withωmin=0\.0000\\omega\_\{\\min\}=0\.0000THz within the adopted tolerance\.b,RbPbCl3is a soft\-mode completed\-phonon unstable example withωmin=−1\.4984\\omega\_\{\\min\}=\-1\.4984THz\.c,BaTiO3is a deep\-imaginary completed\-phonon unstable example withωmin=−7\.9102\\omega\_\{\\min\}=\-7\.9102THz\. Grey branches denote non\-negative modes, colored branches denote imaginary modes and the dashed line marks zero frequency\.The spectrum\-resolved design also defines how the data can be reused\. A classifier may learn the PhononBench\-defined Stable/unstable target\. A regression model may predictωmin\\omega\_\{\\min\}to distinguish near\-threshold modes from deep imaginary branches\. A calibration study may move the−10−3\-10^\{\-3\}THz threshold and quantify label changes\. These tasks all refer back to the same released YAML objects rather than to disconnected post\-processing tables\.

## 5\.Chemical and structural coverage

Figure[3](https://arxiv.org/html/2607.22573#S5.F3)summarizes the completed cohort rather than the planned task list\. All panels use the same denominator: the 46,899 completed label\+YAML records\. The completed cohort spans 87 observed elements and 182 observed space groups\. Oxygen appears in 37\.2% of completed records, making it the dominant element in the occurrence count\. This O\-rich character is chemically natural for MP\-derived inorganic crystals: O2\-forms robust coordination polyhedra with many metal cations, stabilizes a wide range of oxide and oxyanion frameworks, and supports charge\-balanced multication chemistries\. The extended element, space\-group and chemical\-complexity statistics are reported in the Supplementary Information, Sec\. S6\.

The chemical\-complexity distribution is dominated by ternary and quaternary compounds: 50\.5% of records contain three elements and 24\.0% contain four\. This distribution is relevant for reuse and model evaluation\. Binary compounds provide important prototypes, but additional cation or anion species expand charge\-compensation routes, site\-substitution degrees of freedom and coordination environments\. In oxides and oxyanion frameworks, for example, robust oxygen coordination can coexist with multiple metal sublattices, substitutions or charge\-balancing species\. For machine learning, the dataset therefore extends beyond simple binary compounds and contains many chemically richer compositions for which formula\-level cues alone are less likely to capture the relevant lattice\-dynamical behavior\.

The unit\-atom and space\-group panels should be read as coverage statistics of the completed MP40 cohort, not as stability trends\. Peaks at 8, 20 and 40 atoms reflect the MP40 source space, conventional\-cell choices and crystallographic representation\. The 40\-atom edge is especially important: it admits ordered superstructures and multicomponent inorganic cells that are more demanding than very small\-cell benchmarks, while avoiding a release dominated by very large phonon supercells\. The ranked space\-group plot shows broad but uneven crystallographic coverage, with high\-symmetry prototypes and lower\-symmetry distorted or ordered structures both present\. Reporting these distributions is part of the audit: users can see the chemical and structural support over which the stability labels and spectra are defined\.

![Refer to caption](https://arxiv.org/html/2607.22573v1/x3.png)Figure 3:Chemical and structural coverage of the completed cohort\.All panels are computed over the 46,899 completed label\+YAML records\.a,Distribution of recorded unit\-cell atom counts within the MP40 scope, with prominent peaks at 8, 20 and 40 atoms\.b,Element occurrence, counted once per completed record when an element appears in the parsed formula\.c,Ranked space\-group occurrence, showing broad but uneven crystallographic coverage\.d,Chemical complexity measured by the number of distinct elements in the parsed formula, with ternary and quaternary compounds forming the largest groups\.
## 6\.Technical validation

PhononBench\-MP40 is validated as an auditable workflow dataset rather than as a claim of model\-independent absolute dynamical stability\. The first validation layer is cohort accounting: task records, recovered local labels and recovered YAML outputs are intersected explicitly, giving 46,899 completed label\+YAML records and identifying 1,067 relaxation failures outside the completed phonon denominator\. This accounting prevents failed relaxations from being silently merged with completed imaginary\-mode spectra\.

The second validation layer is label traceability\. For each completed record, the released label is derived from the sampled high\-symmetry\-path frequencies in the corresponding phonopy YAML file\. The label rule is deterministic: any sampled band frequency below−10−3\-10^\{\-3\}THz gives the completed\-phonon unstable label, while all other completed spectra are labeled Stable\. Because the YAML file is released, users can re\-extractωmin\\omega\_\{\\min\}, reproduce the binary label and test alternative thresholds without relying on a disconnected label table\.

The third validation layer is distributional and physical sanity checking\. The release reports the completed\-cohort element distribution, space\-group distribution, unit\-atom distribution and chemical\-complexity distribution using the same 46,899\-record denominator\. Representative spectra in Fig\.[2](https://arxiv.org/html/2607.22573#S4.F2)further check that the label rule distinguishes stable spectra, shallow soft\-mode spectra and deep\-imaginary spectra in a physically interpretable way\. Additional validation and recommended integrity checks are described in the Supplementary Information, Secs\. S6 and S10\.

## 7\.Benchmark tasks and reuse

PhononBench\-MP40 is useful as a data release because spectra, labels and accounting are distributed together\. It can also be used as a benchmark target because the completed cohort and label rule are fixed\. For users who wish to compare models under a shared setting, we provide a deterministic formula\-group reference split rather than the only valid way to use the release\. The split assigns 37,633 training records, 4,724 validation records and 4,542 test records so that every one of the 35,348 formula groups appears in exactly one partition\. This prevents exact\-formula leakage and gives future studies a common starting point\. If external materials databases are used for pretraining or feature generation, overlap with test structures, formulas or Materials Project identifiers should be reported because it changes how model scores should be interpreted\. Split construction and reporting suggestions are detailed in the Supplementary Information, Secs\. S5 and S7\.

The most direct benchmark task is PhononBench\-defined stability classification on the completed label\+YAML cohort\. Since the completed cohort is imbalanced, with 30,216 completed\-phonon unstable records and 16,683 Stable records, accuracy alone is not a sufficient metric\. Balanced accuracy, macro\-F1, AUROC, AUPRC and class\-specific recall are more informative for future comparisons\. The released spectra also supportωmin\\omega\_\{\\min\}prediction and threshold\-dependent relabeling\. Separately, the 1,067 relaxation failures define a failure\-aware triage problem: predicting whether a workflow produces an auditable spectrum is useful, but it should not be confused with predicting a completed phonon instability\.

For reproducible baseline reporting, we recommend that future model reports include at least one simple composition\-only baseline, one structure\-aware baseline and, when computationally feasible, a graph\-neural\-network baseline such as CGCNN, MEGNet or ALIGNN\[[34](https://arxiv.org/html/2607.22573#bib.bib34),[35](https://arxiv.org/html/2607.22573#bib.bib35),[36](https://arxiv.org/html/2607.22573#bib.bib36)\]\. The purpose of these baselines is not to rank all model families in the present data paper, but to make the benchmark entry point clear: the split, labels, frequency target and failure\-aware task are fixed, while methods can be compared under transparent reporting rules\. The Supplementary Information, Sec\. S7, lists the recommended baseline protocol and metrics\.

This benchmark use should be interpreted with the right boundary\. PhononBench\-MP40 does not claim model\-independent absolute dynamical stability\. It fixes a documented PhononBench/MatterSim–phonopy reference workflow\. Other potentials, denser sampling, larger supercells, anharmonic treatments or first\-principles phonon calculations may change borderline spectra\. That limitation is also why the YAML object is released: users can inspect imaginary branches, measure threshold sensitivity and validate high\-confidence candidates with the workflow or higher\-fidelity calculations when needed\. Additional interpretation boundaries are summarized in the Supplementary Information, Sec\. S8\.

## 8\.Relation to machine\-learning targets

Materials machine learning increasingly relies on fixed tasks, transparent splits and reproducible metrics, as illustrated by Matbench, Matbench Discovery and OC20\[[37](https://arxiv.org/html/2607.22573#bib.bib37),[38](https://arxiv.org/html/2607.22573#bib.bib38),[39](https://arxiv.org/html/2607.22573#bib.bib39)\]\. PhononBench\-MP40 brings the same logic to workflow\-defined phonon stability while retaining the spectral evidence behind each target\. Descriptor\-based materials learning, matminer workflows, crystal graph neural networks, SchNet, MEGNet and ALIGNN are natural model families for future structure\-to\-label and structure\-to\-frequency studies\[[40](https://arxiv.org/html/2607.22573#bib.bib40),[41](https://arxiv.org/html/2607.22573#bib.bib41),[34](https://arxiv.org/html/2607.22573#bib.bib34),[42](https://arxiv.org/html/2607.22573#bib.bib42),[35](https://arxiv.org/html/2607.22573#bib.bib35),[36](https://arxiv.org/html/2607.22573#bib.bib36)\]\. Universal interatomic potentials such as M3GNet, CHGNet, NequIP, MACE and MatterSim have made large atomistic workflows increasingly practical\[[43](https://arxiv.org/html/2607.22573#bib.bib43),[44](https://arxiv.org/html/2607.22573#bib.bib44),[45](https://arxiv.org/html/2607.22573#bib.bib45),[46](https://arxiv.org/html/2607.22573#bib.bib46),[47](https://arxiv.org/html/2607.22573#bib.bib47),[30](https://arxiv.org/html/2607.22573#bib.bib30)\]\. Recent work also shows that phonons remain a demanding and informative test for such models\[[48](https://arxiv.org/html/2607.22573#bib.bib48)\]\. PhononBench\-MP40 gives this ecosystem a concrete reference target for phonon\-stability screening: not an informal pass/fail label, but a spectrum\-backed workflow result\. The present release does not rank model families; it supplies the shared target, split files and spectral evidence needed for future comparisons\.

## 9\.Usage notes and limitations

The release is intended for three common entry points\. First, users can read the completed table, follow theyaml\_relative\_pathfield to a phonopy YAML file, reconstruct the released dispersion and reproduce the label rule\. Second, benchmark users can join the completed table with the reference split table and evaluate models on the fixed formula\-group split\. Third, workflow users can combine the completed table and the failure table to study whether a structure is likely to produce an auditable phonon spectrum under the reference route\.

Several limitations should be retained in downstream use\. PhononBench\-MP40 is not a universal statement about thermodynamic stability, finite\-temperature stability or DFT\-level phonon stability\. It records the result of a specified MP\-derived PhononBench/MatterSim–phonopy workflow, a specified supercell choice, a specified path sampling procedure and a specified numerical threshold\. Borderline structures, especially those with shallow soft modes, should be treated as candidates for follow\-up calculation rather than final physical classifications\. The released spectra and failure accounting are meant to make these limitations measurable instead of hidden\.

## 10\.Conclusion

PhononBench\-MP40 converts a large MP\-derived PhononBench calculation into a reusable spectrum\-resolved benchmark dataset for workflow\-defined phonon stability\. The main release contains 46,899 completed label\+YAML records, including 16,683 Stable records and 30,216 completed\-phonon unstable records, with 1,067 relaxation failures reported separately\. Its central contribution is the release object: labels, spectra, splits and failure accounting are kept together, so each completed stability decision can be traced back to phonon evidence\. This makes PhononBench\-MP40 useful for stability classification, minimum\-frequency analysis, threshold\-dependent relabeling and failure\-aware screening, while making the reference workflow and its boundaries explicit\.

## Data availability statement

The dataset is being archived in Science Data Bank at[https://www\.scidb\.cn/detail?dataSetId=10\.57760/sciencedb\.38735](https://www.scidb.cn/detail?dataSetId=10.57760/sciencedb.38735); the DOI[https://doi\.org/10\.57760/sciencedb\.38735](https://doi.org/10.57760/sciencedb.38735)should be cited once registration is active\. The Science Data Bank record is intended to be the archival source for the dataset\. The companion GitHub repository at[https://github\.com/DreamLufei/phononbench\_mp40](https://github.com/DreamLufei/phononbench_mp40)provides the calculation workflow, scripts and lightweight access utilities, but does not replace the archived dataset record\. Users should report the data record used for the dataset and the Git commit used for code\-level reuse\. Release\-file definitions, recommended integrity checks and example access steps are provided in the Supplementary Information\.

## Acknowledgements

The work is supported by the National Natural Science Foundation of China \(Nos\. 62476278 and 11934020\), Beijing Natural Science Foundation \(No\. Z250005\) and the National Key R&D Program of China \(Grant No\. 2024YFA1408601\)\. Computational resources were provided by the Physical Laboratory of High Performance Computing at Renmin University of China\.

## References

- \[1\]Tian Xie, Xiang Fu, Octavian\-Eugen Ganea, Regina Barzilay, and Tommi Jaakkola\.Crystal diffusion variational autoencoder for periodic material generation, 2022\.
- \[2\]Rui Jiao, Wenbing Huang, Peijia Lin, Jiaqi Han, Pin Chen, Yutong Lu, and Yang Liu\.Crystal structure prediction by joint equivariant diffusion, 2023\.
- \[3\]Amil Merchant, Simon Batzner, Samuel S\. Schoenholz, Muratahan Aykol, Gowoon Cheon, and Ekin D\. Cubuk\.Scaling deep learning for materials discovery\.Nature, 624:80–85, 2023\.
- \[4\]Claudio Zeni, Robert Pinsler, Daniel Zugner, Andrew Fowler, Matthew Horton, Xiang Fu, Zilong Wang, Aliaksandra Shysheya, Jonathan Crabbe, Shoko Ueda, Roberto Sordillo, Lixin Sun, Jake Smith, Bichlien Nguyen, Hannes Schulz, Sarah Lewis, Chin\-Wei Huang, Ziheng Lu, Yichi Zhou, Han Yang, Hongxia Hao, Jielan Li, Chu Yang, et al\.A generative model for inorganic materials design\.Nature, 639:624–632, 2025\.
- \[5\]Anubhav Jain, Shyue Ping Ong, Geoffroy Hautier, Wei Chen, William Davidson Richards, Stephen Dacek, Shreyas Cholia, Dan Gunter, David Skinner, Gerbrand Ceder, and Kristin A\. Persson\.Commentary: The materials project: A materials genome approach to accelerating materials innovation\.APL Materials, 1\(1\):011002, 2013\.
- \[6\]Stefano Curtarolo, Wahyu Setyawan, Gus L\. W\. Hart, Michal Jahnatek, Roman V\. Chepulskii, Richard H\. Taylor, Shidong Wang, Junkai Xue, Kesong Yang, Ohad Levy, Michael J\. Mehl, Harold T\. Stokes, Denis O\. Demchenko, and Dane Morgan\.Aflow: An automatic framework for high\-throughput materials discovery\.Computational Materials Science, 58:218–226, 2012\.
- \[7\]James E\. Saal, Scott Kirklin, Muratahan Aykol, Bryce Meredig, and C\. Wolverton\.Materials design and discovery with high\-throughput density functional theory: The open quantum materials database \(oqmd\)\.JOM, 65\(11\):1501–1509, 2013\.
- \[8\]Scott Kirklin, James E\. Saal, Bryce Meredig, Alex Thompson, Jeff W\. Doak, Muratahan Aykol, Stephan Ruhl, and C\. Wolverton\.The open quantum materials database \(oqmd\): assessing the accuracy of dft formation energies\.npj Computational Materials, 1:15010, 2015\.
- \[9\]Kamal Choudhary, Kevin F\. Garrity, Andrew C\. E\. Reid, Brian DeCost, Adam J\. Biacchi, Angela R\. Hight Walker, Zachary Trautt, Jason Hattrick\-Simpers, A\. Gilad Kusne, Andrea Centrone, Albert Davydov, Jie Jiang, Ruth Pachter, Gowoon Cheon, Evan Reed, Ankit Agrawal, Xiaofeng Qian, Vishu Sharma, Houlong L\. Zhuang, Irina Kalish, Dan Vogel, Jesus Carrete, Qimin Yan, Chao Cao, Ruoqian Yu, Jeff Doak, Jeffrey B\. Neaton, Geoffroy Hautier, and Francesca Tavazza\.The joint automated repository for various integrated simulations \(jarvis\) for data\-driven materials design\.npj Computational Materials, 6:173, 2020\.
- \[10\]Claudia Draxl and Matthias Scheffler\.Nomad: The fair concept for big data\-driven materials science\.MRS Bulletin, 43\(9\):676–682, 2018\.
- \[11\]Leopold Talirz, Snehal Kumbhar, Elsa Passaro, Aliaksandr V\. Yakutovich, Valeria Granata, Fernando Gargiulo, Marco Borelli, Martin Uhrin, Sebastiaan P\. Huber, Spyros Zoupanos, Carl S\. Adorf, Casper W\. Andersen, Ole Schutt, Carlo A\. Pignedoli, Daniele Passerone, Joost VandeVondele, Thomas C\. Schulthess, Berend Smit, Giovanni Pizzi, and Nicola Marzari\.Materials cloud, a platform for open computational science\.Scientific Data, 7:299, 2020\.
- \[12\]Shyue Ping Ong, William Davidson Richards, Anubhav Jain, Geoffroy Hautier, Michael Kocher, Shreyas Cholia, Dan Gunter, Vincent L\. Chevrier, Kristin A\. Persson, and Gerbrand Ceder\.Python materials genomics \(pymatgen\): A robust, open\-source python library for materials analysis\.Computational Materials Science, 68:314–319, 2013\.
- \[13\]Ask Hjorth Larsen, Jens Jorgen Mortensen, Jakob Blomqvist, Ivano E\. Castelli, Rune Christensen, Marcin Dulak, Jesper Friis, Michael N\. Groves, Bjork Hammer, Cory Hargus, Eric D\. Hermes, Paul C\. Jennings, Peter Bjerre Jensen, James Kermode, John R\. Kitchin, Esben Leonhard Kolsbjerg, Joseph Kubal, Kristen Kaasbjerg, Steen Lysgaard, Jon Bergmann Maronsson, Tristan Maxson, Thomas Olsen, Lars Pastewka, Andrew Peterson, Carsten Rostgaard, Jakob Schiotz, Ole Schutt, Mikkel Strange, Kristian S\. Thygesen, Tejs Vegge, Lasse Vilhelmsen, Michael Walter, Zhenhua Zeng, and Karsten W\. Jacobsen\.The atomic simulation environment–a python library for working with atoms\.Journal of Physics: Condensed Matter, 29:273002, 2017\.
- \[14\]Anubhav Jain, Shyue Ping Ong, Wei Chen, Bharat Medasani, Xiaohui Qu, Michael Kocher, Miriam Brafman, Guido Petretto, Gian\-Marco Rignanese, Geoffroy Hautier, Dan Gunter, and Kristin A\. Persson\.Fireworks: a dynamic workflow system designed for high\-throughput applications\.Concurrency and Computation: Practice and Experience, 27\(17\):5037–5059, 2015\.
- \[15\]Kiran Mathew, Joseph H\. Montoya, Alireza Faghaninia, Shyam Dwaraknath, Muratahan Aykol, Hanmei Tang, Iek\-Heng Chu, Tess Smidt, Brandon Bocklund, Matthew Horton, John Dagdelen, Brandon Wood, Zi\-Kui Liu, Jeffrey B\. Neaton, Shyue Ping Ong, Kristin A\. Persson, and Anubhav Jain\.Atomate: A high\-level interface to generate, execute, and analyze computational materials science workflows\.Computational Materials Science, 139:140–152, 2017\.
- \[16\]Giovanni Pizzi, Andrea Cepellotti, Riccardo Sabatini, Nicola Marzari, and Boris Kozinsky\.Aiida: automated interactive infrastructure and database for computational science\.Computational Materials Science, 111:218–230, 2016\.
- \[17\]Guido Petretto, Shyam Dwaraknath, Henrique P\. C\. Miranda, Donald Winston, Matteo Giantomassi, Michiel J\. van Setten, Xavier Gonze, Kristin A\. Persson, Geoffroy Hautier, and Gian\-Marco Rignanese\.High\-throughput density\-functional perturbation theory phonons for inorganic materials\.Scientific Data, 5:180065, 2018\.
- \[18\]Jose J\. Plata, Pinku Nath, Demet Usanmaz, Jesus Carrete, Cormac Toher, Maarten de Jong, Mark Asta, Marco Fornari, Marco Buongiorno Nardelli, and Stefano Curtarolo\.An efficient and accurate framework for calculating lattice thermal conductivity of solids: Aflow–aapl automatic anharmonic phonon library\.npj Computational Materials, 3:45, 2017\.
- \[19\]Jesus Carrete, Wu Li, Natalio Mingo, Shidong Wang, and Stefano Curtarolo\.Finding unprecedentedly low\-thermal\-conductivity half\-heusler semiconductors via high\-throughput materials modeling\.Physical Review X, 4:011019, 2014\.
- \[20\]Nicolas Mounet, Marco Gibertini, Pedro Schwaller, Andrius Merkys, Ivano E\. Castelli, Andrea Cepellotti, Giovanni Pizzi, and Nicola Marzari\.Two\-dimensional materials from high\-throughput computational exfoliation of experimentally known compounds\.Nature Nanotechnology, 13:246–252, 2018\.
- \[21\]Wu Li, Jesus Carrete, Nebil A\. Katcho, and Natalio Mingo\.Shengbte: A solver of the boltzmann transport equation for phonons\.Computer Physics Communications, 185\(6\):1747–1758, 2014\.
- \[22\]Terumasa Tadano, Yoshihiro Gohda, and Shinji Tsuneyuki\.Anharmonic force constants extracted from first\-principles molecular dynamics: applications to heat transfer simulations\.Journal of the Physical Society of Japan, 83:074702, 2014\.
- \[23\]Olle Hellman, Igor A\. Abrikosov, and Sergei I\. Simak\.Temperature dependent effective potential method for accurate free energy calculations of solids\.Physical Review B, 84:180301, 2011\.
- \[24\]Petros Souvatzis, Olle Eriksson, Mikhail I\. Katsnelson, and Sven P\. Rudin\.The self\-consistent ab initio lattice dynamical method\.Computational Materials Science, 44\(3\):888–894, 2009\.
- \[25\]Masato Ohnishi, Tianqi Deng, Pol Torres, Zhihao Xu, Terumasa Tadano, Haoming Zhang, Wei Nong, Masatoshi Hanai, Zeyu Wang, Michimasa Morita, Zhiting Tian, Ming Hu, Xiulin Ruan, Ryo Yoshida, Toyotaro Suzumura, Lucas Lindsay, Alan J\. H\. McGaughey, Tengfei Luo, Kedar Hippalgaonkar, and Junichiro Shiomi\.Database and deep\-learning scalability of anharmonic phonon properties by automated brute\-force first\-principles calculations\.npj Computational Materials, 12:150, 2026\.
- \[26\]Hang Xiao, Rong Li, Xiaoyang Shi, Yan Chen, Liangliang Zhu, Xi Chen, and Lei Wang\.An invertible, invariant crystal representation for inverse design of solid\-state materials using generative deep learning\.Nature Communications, 14:7027, 2023\.
- \[27\]Hyunsoo Park, Anthony Onwuli, and Aron Walsh\.Exploration of crystal chemical space using text\-guided generative artificial intelligence\.Nature Communications, 16:4379, 2025\.
- \[28\]Atsushi Togo and Isao Tanaka\.First principles phonon calculations in materials science\.Scripta Materialia, 108:1–5, 2015\.
- \[29\]Atsushi Togo\.First\-principles phonon calculations with phonopy and phono3py\.Journal of the Physical Society of Japan, 92:012001, 2023\.
- \[30\]Han Yang, Chenxi Hu, Yichi Zhou, Xixian Liu, Yu Shi, Jielan Li, Guanzhi Li, Zekun Chen, Shuizhou Chen, Claudio Zeni, Matthew Horton, Robert Pinsler, Andrew Fowler, Daniel Zugner, Tian Xie, Jake Smith, Lixin Sun, Qian Wang, Lingyu Kong, Chang Liu, Hongxia Hao, and Ziheng Lu\.Mattersim: A deep learning atomistic model across elements, temperatures and pressures, 2024\.
- \[31\]Xiao\-Qi Han, Peng\-Jie Guo, Ze\-Feng Gao, and Zhong\-Yi Lu\.Phononbench: A large\-scale phonon\-based benchmark for dynamical stability in crystal generation, 2025\.GitHub repository and manuscript resource\.
- \[32\]Yoyo Hinuma, Giovanni Pizzi, Yu Kumagai, Fumiyasu Oba, and Isao Tanaka\.Band structure diagram paths based on crystallography\.Computational Materials Science, 128:140–184, 2017\.
- \[33\]Atsushi Togo, Kohei Shinohara, and Isao Tanaka\.Spglib: a software library for crystal symmetry search\.Science and Technology of Advanced Materials: Methods, 4\(1\):2384822, 2024\.
- \[34\]Tian Xie and Jeffrey C\. Grossman\.Crystal graph convolutional neural networks for an accurate and interpretable prediction of material properties\.Physical Review Letters, 120:145301, 2018\.
- \[35\]Chi Chen, Weike Ye, Yunxing Zuo, Chen Zheng, and Shyue Ping Ong\.Graph networks as a universal machine learning framework for molecules and crystals\.Chemistry of Materials, 31\(9\):3564–3572, 2019\.
- \[36\]Kamal Choudhary and Brian DeCost\.Atomistic line graph neural network for improved materials property predictions\.npj Computational Materials, 7:185, 2021\.
- \[37\]Alexander Dunn, Qi Wang, Alex Ganose, Daniel Dopp, and Anubhav Jain\.Benchmarking materials property prediction methods: the matbench test set and automatminer reference algorithm\.npj Computational Materials, 6:138, 2020\.
- \[38\]Janosh Riebesell, Rhys E\. A\. Goodall, Philipp Benner, Yuan Chiang, Bowen Deng, Gerbrand Ceder, Mark Asta, Alpha A\. Lee, Anubhav Jain, and Kristin A\. Persson\.A framework to evaluate machine learning crystal stability predictions\.Nature Machine Intelligence, 7:836–847, 2025\.
- \[39\]Lowik Chanussot, Abhishek Das, Siddharth Goyal, Thibaut Lavril, Muhammed Shuaibi, Morgane Riviere, Kevin Tran, Javier Heras\-Domingo, Caleb Ho, Weihua Hu, Aini Palizhati, Anuroop Sriram, Brandon Wood, Junwoong Yoon, Devi Parikh, C\. Lawrence Zitnick, and Zachary Ulissi\.The open catalyst 2020 \(oc20\) dataset and community challenges\.ACS Catalysis, 11\(10\):6059–6072, 2021\.
- \[40\]Logan Ward, Ankit Agrawal, Alok Choudhary, and Christopher Wolverton\.A general\-purpose machine learning framework for predicting properties of inorganic materials\.npj Computational Materials, 2:16028, 2016\.
- \[41\]Logan Ward, Alexander Dunn, Alireza Faghaninia, Nils E\. R\. Zimmermann, Saurabh Bajaj, Qi Wang, Joseph Montoya, Jiming Chen, Kyle Bystrom, Maxwell Dylla, Kyle Chard, Mark Asta, Kristin Persson, G\. Jeffrey Snyder, and Ian Foster\.Matminer: An open source toolkit for materials data mining\.Computational Materials Science, 152:60–69, 2018\.
- \[42\]Kristof T\. Schutt, Huziel E\. Sauceda, Pieter\-Jan Kindermans, Alexandre Tkatchenko, and Klaus\-Robert Muller\.Schnet: A deep learning architecture for molecules and materials\.The Journal of Chemical Physics, 148:241722, 2018\.
- \[43\]Chi Chen and Shyue Ping Ong\.A universal graph deep learning interatomic potential for the periodic table\.Nature Computational Science, 2:718–728, 2022\.
- \[44\]Bowen Deng, Peichen Zhong, KyuJung Jun, Janosh Riebesell, Kevin Han, Christopher J\. Bartel, and Gerbrand Ceder\.Chgnet as a pretrained universal neural network potential for charge\-informed atomistic modelling\.Nature Machine Intelligence, 5:1031–1041, 2023\.
- \[45\]Simon Batzner, Albert Musaelian, Lixin Sun, Mario Geiger, Jonathan P\. Mailoa, Mordechai Kornbluth, Nicola Molinari, Tess E\. Smidt, and Boris Kozinsky\.E\(3\)\-equivariant graph neural networks for data\-efficient and accurate interatomic potentials\.Nature Communications, 13:2453, 2022\.
- \[46\]Ilyes Batatia, David Peter Kovacs, Gregor N\. C\. Simm, Christoph Ortner, and Gabor Csanyi\.Mace: Higher order equivariant message passing neural networks for fast and accurate force fields, 2022\.
- \[47\]Ilyes Batatia, Philipp Benner, Yuan Chiang, Alin M\. Elena, David P\. Kovacs, Janosh Riebesell, Xavier R\. Advincula, Mark Asta, William J\. Baldwin, Fabian Berger, et al\.A foundation model for atomistic materials chemistry\.The Journal of Chemical Physics, 163:172501, 2025\.
- \[48\]Antoine Loew, Dewen Sun, Hai\-Chen Wang, Silvana Botti, and Miguel A\. L\. Marques\.Universal machine learning interatomic potentials are ready for phonons\.npj Computational Materials, 11:178, 2025\.

Supplementary Information PhononBench\-MP40: a spectrum\-resolved benchmark dataset for phonon stability Wen\-Kao Li1,†, Ze\-Feng Gao1,†,∗, Zhong\-Yi Lu1,∗

1School of Physics and Key Laboratory of Quantum State Construction and Manipulation \(Ministry of Education\), Renmin University of China, Beijing 100872, China †These authors contributed equally\. ∗e\-mail: zfgao@ruc\.edu\.cn; zlu@ruc\.edu\.cn

## S1\. Data schema and release files

High\-level cohort accounting is summarized in Table 1 of the main text\. This section defines the release files used to produce those counts and the manuscript figures\. The release is organized around four tabular data objects, release metadata and the corresponding phonopy YAML outputs\. The master table records the audited MP40 task scope\. The completed table contains only entries with both a local workflow label and a matching local YAML output\. The failure table records relaxation failures without matching YAML outputs\. The split table gives the reference train, validation and test assignment for completed records\. The metadata and checksum files, when provided, define the release version and support file\-integrity checks after download\.

The MP40 scope denotes recorded cells withNatoms≤40N\_\{\\mathrm\{atoms\}\}\\leq 40\. This range was used to balance chemical and structural coverage against the cost of systematic phonon calculations\. Under the2×2×22\\times 2\\times 2supercell used in the workflow, a 40\-atom recorded cell corresponds to an up\-to\-320\-atom force\-evaluation cell before symmetry reduction\.

Table S1\. Release files used for the manuscript and release definition\.

## S2\. Field definitions

The common fields are listed in Table S2\.

Table S2\. Common field definitions\.

## S3\. Workflow and label rule

The labels follow the local PhononBench workflow used for this release\. Structures are relaxed before finite\-displacement phonon calculations\. The phonon calculation uses a2×2×22\\times 2\\times 2supercell matrix, a displacement distance of 0\.01 Å, drift correction on displaced\-supercell forces and symmetrized force constants\. Band paths are generated through the seekpath/spglib convention with 101 points per segment\.

For reproducibility, downstream users should report the data\-release version, the Git commit used for parsing utilities, and the versions of the main phonon\-analysis packages used in any reprocessing\. The manuscript labels in the release are tied to the released YAML spectra and the threshold below; if a user regenerates phonons with a different potential, supercell, band path or threshold, the resulting labels should be reported as a new derived target rather than as the original PhononBench\-MP40 label\.

The completed stability label is derived from the high\-symmetry\-path frequencies\. A structure is labeled unstable when any band frequency satisfies

ω<−10−3​THz\.\\omega<\-10^\{\-3\}\\ \{\\rm THz\}\.Otherwise it is labeled Stable\. The release preserves the YAML spectrum so that users can reconstruct the band structure, extractωmin\\omega\_\{\\min\}, and apply stricter or looser thresholds\.

## S4\. Representative spectra

Three representative examples were selected to show distinct spectral regimes behind the binary labels\.

Table S3\. Representative spectra shown in Fig\. 2 of the main text\.

## S5\. Reference split

The reference split is deterministic at the formula\-group level\. The group key is the parsed formula token\. A SHA\-256 hash of this key assigns each formula group to train, validation or test using the rule train<0\.80<0\.80, validation<0\.90<0\.90and test otherwise\. This split prevents exact\-formula leakage across train, validation and test sets\.

Table S4\. Reference split counts by completed label\.

The completed cohort contains 35,348 formula groups\. Under the reference split, every formula group appears in exactly one split, giving zero exact\-formula leakage\.

## S6\. Coverage statistics

The completed cohort contains 87 observed elements and 182 observed space groups\. The most frequent elements are O, Mg, F, Li, Cu, Fe, Mn, S, Co, V, Ni, Sr, K, Ti, Ba and Al\. The corresponding occurrence percentages over completed records are 37\.2%, 16\.7%, 9\.4%, 9\.1%, 7\.9%, 7\.8%, 6\.7%, 6\.5%, 6\.2%, 5\.8%, 5\.7%, 5\.4%, 5\.2%, 5\.1%, 5\.0% and 4\.8%\.

The leading space groups are 225, 1, 71, 12, 194, 221, 139, 216, 123, 2, 166, 8, 38, 62, 187 and 160\. Their occurrence percentages are 16\.1%, 5\.9%, 4\.6%, 4\.4%, 3\.9%, 3\.8%, 3\.7%, 3\.6%, 3\.3%, 3\.3%, 3\.0%, 2\.9%, 2\.3%, 1\.9%, 1\.6% and 1\.5%\.

Chemical complexity is measured by the number of distinct elements in the parsed formula\. The completed cohort contains 0\.7% unary, 17\.8% binary, 50\.5% ternary, 24\.0% quaternary and 7\.0% five\-or\-more\-element records\.

## S7\. Benchmark uses, tasks and metrics

The release supports several benchmark\-style uses without prescribing a single field\-wide protocol\. For users who wish to compare models under a shared setting, the deterministic formula\-group split provides a reference split rather than the only valid way to use the release\. When another split is used, exact\-formula leakage and overlap with external pretraining data should be reported\.

Stability classification\.The input is a crystal structure in the completed label\+YAML cohort\. The target is the workflow\-defined Stable/unstable label\. Recommended metrics include balanced accuracy, macro\-F1, AUROC, AUPRC and class\-specific recall\. Accuracy alone is not sufficient because the completed cohort is label\-imbalanced\. Scores should be reported for the whole test set and for chemically meaningful subsets, including atom\-count bins, chemical\-complexity classes, oxygen\-containing versus non\-oxygen\-containing records and common space\-group families\.

Minimum\-frequency regression\.The input is the same completed cohort\. The target isωmin\\omega\_\{\\min\}extracted from the YAML spectrum\. Recommended metrics include MAE and RMSE in THz, Spearman rank correlation and error stratified by near\-threshold, soft\-mode and deep\-instability cases\.

Failure\-aware workflow triage\.The input includes completed records and the 1,067 relaxation failures\. The target is whether a record produced an auditable local YAML spectrum\. Recommended metrics include precision, recall and F1 for the failure class\. This task should be reported separately from physical phonon\-stability prediction\.

Candidate validation\.Model predictions on new structures should be treated as screening outputs\. A predicted stable candidate should be validated by rerunning the PhononBench workflow, and by a higher\-fidelity phonon calculation when the downstream use requires stricter confirmation\. This separates model scoring on the release from final materials validation\.

Reporting checklist\.A transparent model report should state the structure representation, model family, training data, external pretraining data, split used for evaluation, subgroup scores when relevant, classification threshold,ωmin\\omega\_\{\\min\}regression metrics when used and the validation route for any newly proposed stable candidates\.

Recommended baseline protocol\.A minimal benchmark report should include a majority\-class baseline, a composition\-only baseline and one structure\-aware baseline\. The composition\-only baseline can use stoichiometric and elemental statistics features, while the structure\-aware baseline can use graph, local\-environment or symmetry\-aware descriptors\. For graph neural networks, CGCNN, MEGNet and ALIGNN are natural starting points\. Baseline scores should be reported on the fixed test split and should not be tuned on the test set\. If a baseline uses pretrained representations or external materials data, exact overlap with PhononBench\-MP40 test formulas, source identifiers or structures should be reported\.

## S8\. Boundaries of interpretation

PhononBench\-MP40 is a release of a defined computational workflow\. The labels inherit the workflow settings, including relaxation protocol, interatomic potential, finite\-displacement setup, band path and numerical threshold\. Borderline structures may change under denser sampling, larger supercells, alternative potentials, anharmonic treatment or higher\-fidelity follow\-up calculations\. The released spectra are intended to make these boundaries explicit: users can inspect the imaginary branch, measure the margin to the threshold and choose application\-specific labels when needed\.

## S9\. Usage notes

The recommended first\-use workflow is:

1. 1\.Download the archived release from Science Data Bank and verify the release tables against the accompanying checksum file, when provided\.
2. 2\.Read the completed table and select one completed record\.
3. 3\.Useyaml\_relative\_pathto locate the corresponding phonopy YAML file\.
4. 4\.Parse the YAML frequencies, compute the minimum sampled frequencyωmin\\omega\_\{\\min\}and apply the thresholdω<−10−3\\omega<\-10^\{\-3\}THz to reproduce the completed label\.
5. 5\.Join the completed table with the split table before model training so that train, validation and test records follow the reference formula\-group split\.

For label\-threshold studies, users should keep the original completed label unchanged and store any threshold\-dependent relabeling in a new derived column\. This makes it possible to compare alternative label conventions without losing the original workflow\-defined target\.

For failure\-aware studies, users should combine the completed table with the failure table and define a separate binary target for whether a task produced an auditable YAML spectrum\. This target should not be mixed with the completed\-phonon Stable/unstable label\.

## S10\. Technical validation checklist

The following checks define the validation logic used for the release and are recommended for downstream mirrors of the data\.

Table S5\. Validation and integrity checks\.

Similar Articles

StabilityBench: Benchmarking Instability in LLMs

arXiv cs.LG

StabilityBench is a benchmark operator that transforms single-turn LLM evaluations into multi-turn interactions with user simulations and baiting modules, revealing significant performance instability in current models and motivating more realistic evaluation protocols.

Introducing BetterBench - more accurate PP and TPS measurement

Reddit r/LocalLLaMA

BetterBench is a new benchmarking tool that provides more accurate PP and TPS measurements by ensuring content consistency within 1% and testing across different content types, addressing variability in LLM performance benchmarks.

PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research

Hugging Face Daily Papers

PRL-Bench is a comprehensive benchmark for evaluating LLMs' capabilities in frontier physics research, constructed from 100 curated Physical Review Letters papers across five physics subfields. The benchmark reveals significant gaps in current LLM performance (best scores below 50%), designed to test end-to-end research workflows, complex reasoning, and autonomous exploration.