Tag
Alexander Kalian argues that biology lacks sufficient high-quality data for AI, while César de la Fuente counters that data exists but needs better organization, highlighting the AllTheBacteria preprint as an open platform for bacterial genomes.
This paper systematically benchmarks classical machine learning models (Random Forest, XGBoost, etc.) for ER status prediction using multi-omics data from TCGA-BRCA, finding that RNA expression provides the strongest predictive signal and that Random Forest achieves 90.3% balanced accuracy in the integrated multi-omic setting.
Google DeepMind and Isomorphic Labs outline their joint approach to bioresilience, focusing on preventing misuse of AI models and using AI to improve prevention, detection, and response to infectious disease outbreaks through partnerships and tools like AlphaFold, Gemini, and SynthID.
A practical guide on sequencing your own genome at home using the Oxford Nanopore MinION, covering hardware, protocol steps, and analysis tools.
An introductory guide to genomics written for engineers and computer scientists, covering cells, chromosomes, DNA, and genes using a bakery analogy.
OpenAI introduces GeneBench-Pro, a research-level benchmark to test AI agents' ability to navigate messy biological data, choose analysis paths, and make judgment calls in computational biology.
GeneBench-Pro is a comprehensive benchmark from OpenAI designed to evaluate AI models on complex genomics tasks, including somatic oncology, functional genomics, and clinical carrier screening.
OpenAI introduces GeneBench-Pro, a research-level benchmark designed to test AI agents' ability to perform judgment-heavy analyses in computational biology, covering genomics, quantitative biology, and translational medicine.
This paper introduces GRAFT, a curated multimodal dataset linking gene expression profiles and phenotypic traits in Arabidopsis thaliana, along with graph and hypergraph benchmarks for phenotype prediction. It aims to advance genome-to-phenome mapping in plant biology.
A Twitter user shares recommendations from others on the best accounts to follow for AI in biology, covering protein design, genomics, biosecurity, and policy.
Colossal and the US Fish and Wildlife Service will sequence the genomes of all endangered species in the US, storing biological samples in Colossal's BioVault and making genomic data openly accessible to aid conservation efforts.
A new Nature paper from the Pakistan Genomic Resource (PGR) analyzes 173,303 Pakistanis from consanguineous communities, identifying human knockouts for nearly one-third of protein-coding genes, overturning biological assumptions like PRDM9 essentiality for fertility.
OpenAI highlights how o3 Deep Research can aid rare disease diagnosis by integrating clinical features, inheritance patterns, variant evidence, and scientific literature into actionable hypotheses for specialists.
A new study reveals that cockroach genomes contain thousands of pieces of bacterial DNA acquired through horizontal gene transfer, challenging the assumption that such transfers are rare in complex animals.
This paper investigates how post-training stages such as continued pre-training, supervised fine-tuning, and reinforcement learning affect generalization in biological reasoning models, finding that these stages have distinct impacts on in-domain and out-of-domain performance.
A new study reveals that the genomes of early complex cells (eukaryotes) were built through multiple waves of gene transfers from various bacteria and archaea, complicating the simple fusion model.
The author recounts the tragic loss of his son Owen to a rare lung disease and introduces Gamow Labs, a platform built with AI-assisted coding aimed at helping families with similar genetic diagnoses.
LDARNet is a 120M-parameter hierarchical genomic foundation model that introduces learnable adaptive tokenization (inspired by H-Net's dynamic chunking) for masked language modeling on DNA sequences. It achieves state-of-the-art results on 5 histone modification tasks and outperforms models up to 20× larger on several genomic benchmarks, with learned token boundaries aligning with biological features like promoter motifs and splice junctions.
GENEB is a large-scale diagnostic benchmark that evaluates 40 genomic foundation models across 100 tasks in 13 functional categories under a unified probing protocol, exposing that aggregate leaderboards are unstable and that architectural alignment often outweighs model scale. The work addresses the fragmented evaluation landscape in genomic machine learning, analogous to what MTEB did for NLP.
OpenAI introduces an updated GPT-Rosalind model purpose-built for life sciences research, with improved performance in medicinal chemistry, genomics, and drug-discovery workflows, and new benchmarks like LifeSciBench and MedChemBench.