NVIDIA, Google DeepMind, and other organizations have released predicted 3D structures for over 2,800 viral proteins using AlphaFold2 and NVIDIA BioNeMo tools to help researchers prepare for future pandemics, making the data openly available.
<div id="bsf_rt_marker"></div><p><span style="font-weight: 400;">When COVID-19 emerged, scientists had a crucial advantage: Decades of prior research on coronaviruses meant they understood the virus’ key proteins well enough to design vaccines in record time. The next pandemic may not offer the same head start. </span></p>
<p><span style="font-weight: 400;">To help improve the odds, NVIDIA has joined a coalition of global research organizations, including Google DeepMind and the European Molecular Biology Laboratory’s European Bioinformatics Institute (EMBL-EBI), to release predicted 3D structures for the protein complexes of more than 2,800 viruses — openly available to any scientist, anywhere, through the AlphaFold Database.</span></p>
<p><span style="font-weight: 400;">The structures in the newly released dataset were inferred using AlphaFold2 — Google DeepMind’s AI model for predicting how proteins fold into 3D shapes — with optimization from </span><a target="_blank" href="https://docs.nvidia.com/bionemo/inference-runtime/overview"><span style="font-weight: 400;">NVIDIA BioNeMo Inference Runtime</span></a><span style="font-weight: 400;">. This allowed the team to scale inference to thousands of viral proteomes, predicting the complexes, or groups of interacting proteins, encoded within each virus. </span></p>
<p><span style="font-weight: 400;">“Our ambition with the AlphaFold Database has always been to democratize access to foundational biology at scale,” said Risha Patel, life sciences partnerships manager at Google DeepMind. “This collaboration to bring thousands of viral complexes into the database will equip scientists around the world with insights they need to help prepare for future outbreaks.”</span></p>
<p><span style="font-weight: 400;">NVIDIA is also openly releasing the </span><a target="_blank" href="https://github.com/NVIDIA-BioNeMo/BioNeMo-Structure-Prediction-Pipeline"><span style="font-weight: 400;">BioNeMo Structure Prediction Pipeline</span></a><span style="font-weight: 400;">, the GPU-accelerated workflow used to generate the dataset, so researchers can go from protein sequence to predicted 3D structure for their own targets.</span></p>
<p><span style="font-weight: 400;">Preparation for the next pandemic must begin now. An analysis by the Center for Global Development estimates a </span><a target="_blank" href="https://www.cgdev.org/blog/the-next-pandemic-could-come-soon-and-be-deadlier"><span style="font-weight: 400;">roughly 50% chance</span></a><span style="font-weight: 400;"> of the world facing a pandemic as severe as COVID-19 by 2050.</span></p>
<p><span style="font-weight: 400;">“When the next pandemic happens, there may be something that comes out of the blue, and we’ll be lacking the knowledge we had for COVID,” said Joe Grove, professor of molecular virology at the Medical Research Council-University of Glasgow Centre for Virus Research and a collaborator on the project. “What we’re trying to do is stockpile some of that knowledge ahead of time.”</span></p>
<p><span style="font-weight: 400;">About 30% of the protein interactions being added to the database are completely new to science, showing interaction shapes that have never been documented in the Protein Data Bank, the main repository of experimentally determined protein structures. This translates to new insights for the biological community to explore and harness to generate new knowledge.</span></p>
<p><span style="font-weight: 400;">“This database is an engine for hypothesis generation,” said Chris Dallago, applied research science team lead in digital biology at NVIDIA. “We’re enabling biologists and the AI community to investigate protein interactions, not just as single molecules but as complexes, so the whole field can move forward.”</span></p>
<h2><b>Predicting Complex Protein Structures</b></h2>
<p><span style="font-weight: 400;">Most proteins don’t work alone — they come together in complexes of multiple molecules to perform sophisticated functions. Those structures are often what a vaccine or drug must target to disrupt viral function. </span></p>
<p><span style="font-weight: 400;">Understanding the 3D structure of the COVID-19 virus’ spike protein, for example, proved foundational to vaccine design. For thousands of other viruses, no such structural knowledge exists today. This dataset begins to fill that gap.</span></p>
<p><span style="font-weight: 400;">Traditional methods for determining protein structures — crystallizing proteins and shooting X-rays at them — can take years and cost thousands of dollars per structure. AlphaFold2, which was optimized with </span><a target="_blank" href="https://github.com/NVIDIA-BioNeMo"><span style="font-weight: 400;">NVIDIA BioNeMo</span></a><span style="font-weight: 400;"> to efficiently run on NVIDIA GPUs, predicts a structure in minutes and can be run in bulk. Scientists can then verify high-confidence predictions through experimental methods. </span></p>
<p><span style="font-weight: 400;">For this project, the team systematically worked through the protein structures of viral families known to infect humans, from common-cold viruses to emerging threats like Mpox. </span></p>
<h2><b>A Global Collaboration With Global Access</b></h2>
<p><span style="font-weight: 400;">The collaboration spans the Coalition for Epidemic Preparedness Innovations, EMBL-EBI, Google DeepMind, NVIDIA, Seoul National University, Sungkyunkwan University, the Swiss Institute of Bioinformatics and the University of Glasgow. </span></p>
<p><span style="font-weight: 400;">The dataset release — coinciding with a United Nations General Assembly meeting convened by the World Economic Forum on pandemic prevention, preparedness and response taking place this week in New York City — contributes to the AlphaFold Database, which now holds more than 260 million protein and protein complex predictions covering nearly every cataloged protein known to science.</span></p>
<p><span style="font-weight: 400;">“Making this data open is critical for understanding viral diagnostics and developing treatments and vaccines,” said Jo McEntyre, interim director of EMBL-EBI. “The dataset also covers lesser-studied viruses and lowers the barriers for scientists in low-resource settings who are confronting outbreaks firsthand.”</span></p>
<p><span style="font-weight: 400;">Predictions in the open dataset are labeled by confidence. The structures show what viral complexes may look like and how individual proteins might interact within a viral proteome. </span></p>
<p><span style="font-weight: 400;">Overall, the new data represents a major contribution to the information available for scientists across digital biology and disease research.</span></p>
<p><span style="font-weight: 400;">“When I did my Ph.D., there were no structures for any of the proteins we were investigating. It was like working in the dark — we had to guess what was going on,” said Grove. “This dataset is a powerful tool for all the researchers doing their Ph.D.s now, giving them high-quality structural data that’s going to accelerate fundamental science.”</span></p>
<p><i><span style="font-weight: 400;">Explore the viral protein complex dataset on the </span></i><a target="_blank" href="https://alphafold.ebi.ac.uk/"><i><span style="font-weight: 400;">AlphaFold Database Pandemic Preparedness Portal</span></i></a><i><span style="font-weight: 400;">, predict structures for protein targets with the </span></i><a target="_blank" href="https://github.com/NVIDIA-BioNeMo/BioNeMo-Structure-Prediction-Pipeline"><i><span style="font-weight: 400;">BioNeMo Structure Prediction Pipeline</span></i></a><i><span style="font-weight: 400;">, and learn more about </span></i><a target="_blank" href="https://github.com/NVIDIA-BioNeMo"><i><span style="font-weight: 400;">NVIDIA BioNeMo</span></i></a><i><span style="font-weight: 400;">.</span></i></p>
# How Open Science Can Help Researchers Prepare for the Next Pandemic
Source: [https://blogs.nvidia.com/blog/open-protein-dataset/](https://blogs.nvidia.com/blog/open-protein-dataset/)
When COVID\-19 emerged, scientists had a crucial advantage: Decades of prior research on coronaviruses meant they understood the virus’ key proteins well enough to design vaccines in record time\. The next pandemic may not offer the same head start\.
To help improve the odds, NVIDIA has joined a coalition of global research organizations, including Google DeepMind and the European Molecular Biology Laboratory’s European Bioinformatics Institute \(EMBL\-EBI\), to release predicted 3D structures for the protein complexes of more than 2,800 viruses — openly available to any scientist, anywhere, through the AlphaFold Database\.
The structures in the newly released dataset were inferred using AlphaFold2 — Google DeepMind’s AI model for predicting how proteins fold into 3D shapes — with optimization from[NVIDIA BioNeMo Inference Runtime](https://docs.nvidia.com/bionemo/inference-runtime/overview)\. This allowed the team to scale inference to thousands of viral proteomes, predicting the complexes, or groups of interacting proteins, encoded within each virus\.
“Our ambition with the AlphaFold Database has always been to democratize access to foundational biology at scale,” said Risha Patel, life sciences partnerships manager at Google DeepMind\. “This collaboration to bring thousands of viral complexes into the database will equip scientists around the world with insights they need to help prepare for future outbreaks\.”
NVIDIA is also openly releasing the[BioNeMo Structure Prediction Pipeline](https://github.com/NVIDIA-BioNeMo/BioNeMo-Structure-Prediction-Pipeline), the GPU\-accelerated workflow used to generate the dataset, so researchers can go from protein sequence to predicted 3D structure for their own targets\.
Preparation for the next pandemic must begin now\. An analysis by the Center for Global Development estimates a[roughly 50% chance](https://www.cgdev.org/blog/the-next-pandemic-could-come-soon-and-be-deadlier)of the world facing a pandemic as severe as COVID\-19 by 2050\.
“When the next pandemic happens, there may be something that comes out of the blue, and we’ll be lacking the knowledge we had for COVID,” said Joe Grove, professor of molecular virology at the Medical Research Council\-University of Glasgow Centre for Virus Research and a collaborator on the project\. “What we’re trying to do is stockpile some of that knowledge ahead of time\.”
About 30% of the protein interactions being added to the database are completely new to science, showing interaction shapes that have never been documented in the Protein Data Bank, the main repository of experimentally determined protein structures\. This translates to new insights for the biological community to explore and harness to generate new knowledge\.
“This database is an engine for hypothesis generation,” said Chris Dallago, applied research science team lead in digital biology at NVIDIA\. “We’re enabling biologists and the AI community to investigate protein interactions, not just as single molecules but as complexes, so the whole field can move forward\.”
## **Predicting Complex Protein Structures**
Most proteins don’t work alone — they come together in complexes of multiple molecules to perform sophisticated functions\. Those structures are often what a vaccine or drug must target to disrupt viral function\.
Understanding the 3D structure of the COVID\-19 virus’ spike protein, for example, proved foundational to vaccine design\. For thousands of other viruses, no such structural knowledge exists today\. This dataset begins to fill that gap\.
Traditional methods for determining protein structures — crystallizing proteins and shooting X\-rays at them — can take years and cost thousands of dollars per structure\. AlphaFold2, which was optimized with[NVIDIA BioNeMo](https://github.com/NVIDIA-BioNeMo)to efficiently run on NVIDIA GPUs, predicts a structure in minutes and can be run in bulk\. Scientists can then verify high\-confidence predictions through experimental methods\.
For this project, the team systematically worked through the protein structures of viral families known to infect humans, from common\-cold viruses to emerging threats like Mpox\.
## **A Global Collaboration With Global Access**
The collaboration spans the Coalition for Epidemic Preparedness Innovations, EMBL\-EBI, Google DeepMind, NVIDIA, Seoul National University, Sungkyunkwan University, the Swiss Institute of Bioinformatics and the University of Glasgow\.
The dataset release — coinciding with a United Nations General Assembly meeting convened by the World Economic Forum on pandemic prevention, preparedness and response taking place this week in New York City — contributes to the AlphaFold Database, which now holds more than 260 million protein and protein complex predictions covering nearly every cataloged protein known to science\.
“Making this data open is critical for understanding viral diagnostics and developing treatments and vaccines,” said Jo McEntyre, interim director of EMBL\-EBI\. “The dataset also covers lesser\-studied viruses and lowers the barriers for scientists in low\-resource settings who are confronting outbreaks firsthand\.”
Predictions in the open dataset are labeled by confidence\. The structures show what viral complexes may look like and how individual proteins might interact within a viral proteome\.
Overall, the new data represents a major contribution to the information available for scientists across digital biology and disease research\.
“When I did my Ph\.D\., there were no structures for any of the proteins we were investigating\. It was like working in the dark — we had to guess what was going on,” said Grove\. “This dataset is a powerful tool for all the researchers doing their Ph\.D\.s now, giving them high\-quality structural data that’s going to accelerate fundamental science\.”
*Explore the viral protein complex dataset on the*[*AlphaFold Database Pandemic Preparedness Portal*](https://alphafold.ebi.ac.uk/)*, predict structures for protein targets with the*[*BioNeMo Structure Prediction Pipeline*](https://github.com/NVIDIA-BioNeMo/BioNeMo-Structure-Prediction-Pipeline)*, and learn more about*[*NVIDIA BioNeMo*](https://github.com/NVIDIA-BioNeMo)*\.*
Google DeepMind and Isomorphic Labs outline their joint approach to bioresilience, focusing on preventing misuse of AI models and using AI to improve prevention, detection, and response to infectious disease outbreaks through partnerships and tools like AlphaFold, Gemini, and SynthID.
A roundup of 12 AI co-scientist systems in 2026, including DeepMind's Co-Scientist finding a fibrosis drug candidate and OpenAI's reasoning model solving an 80-year-old geometry problem, highlighting open-source tools for biology, fluid simulation, and automated research.
Google DeepMind introduces a computational discovery prototype that uses AlphaEvolve and Empirical Research Assistance to develop and score thousands of code variations in parallel, enabling faster testing of modelling approaches for epidemiology.
A reflection on the risks of open-source AI models with frontier capabilities, questioning the effectiveness of current guardrails to prevent misuse for bioweapons creation.