Tag
RIBOSPAN is a large bidirectional RNA foundation model pretrained on up to 10,240 nucleotides, enabling high-resolution full-transcript modeling and mRNA generation through discrete diffusion.
SNAIL is a hybrid named entity recognition framework that automatically identifies bioinformatics software and database names from scientific literature, outperforming existing methods and enabling large-scale tool usage analysis.
GenEx is a novel graph-based pipeline that converts SARS-CoV-2 gene sequences into codon co-occurrence graphs to detect variants, using techniques like MSCG and LAPCG, and demonstrates effectiveness with benchmarked ML models.
ProteinDPO is a method that uses LLM preference learning techniques to improve the stability of protein models, developed by researchers at Arc Institute.
MIT Technology Review profiles Deanne Taylor, a bioinformatics director who pushed for pediatric representation in the Human Cell Atlas and helped launch the dGTEx project to map healthy gene expression in children, addressing the lack of baseline data for pediatric medicine.
Presents EGRL, a graph neural network framework for RNA-protein interaction prediction that improves cold-start generalization via edge generation and multi-relational attention, achieving competitive results on benchmark datasets.
CellWorld introduces a latent-space predictive pretraining approach for spatial transcriptomics foundation models, predicting latent representations of masked cells instead of reconstructing gene measurements. Across held-out datasets, even small variants outperform existing baselines on all benchmarks, showing that scaling and broad biological diversity improve transferability.
A new 1.1B-parameter DNA foundation model, MarinDNA v0.5 scaling ladder, was released on Hugging Face; it reads and generates DNA sequences and reportedly rivals Evo 2 40B on variant effect prediction.
BioM-JEPA introduces a joint-embedding predictive architecture that learns single-cell representations by predicting graph-connected gene blocks instead of individual genes, showing improved efficiency and downstream performance in perturbation-response tasks.
This arXiv paper introduces CohortHijack, a robustness audit that removes non-target cells from single-cell query cohorts to test how annotation tools can be manipulated without altering the target cell's expression profile. It shows that structured removal and search strategies can change refined labels in popular pipelines while preserving the target, identifying query cohort composition as a vulnerability surface.
The paper presents TourSynbio-Search, an LLM-driven agent framework for unified protein engineering search across literature and biological databases, powered by the TourSynbio-7B multimodal model with dual PaperSearch and ProteinSearch components.
Introduces AutoProteinEngine (AutoPE), an LLM-driven agent framework that enables biologists without deep learning expertise to perform multimodal AutoML for protein engineering via natural language, showing improvements over zero-shot and manual fine-tuning approaches.
This paper studies a self-supervised task for generating single-cell gene expression vectors using an autoregressive transformer with a quantized VAE tokenizer. It reports scaling laws and a compute-optimal frontier for single-cell foundation models, with potential fine-tuning for perturbation prediction.
This paper introduces CoCoS, a contrastive pretraining framework that learns whole-cell representations from complementary transcriptomic views, addressing limitations of masked gene reconstruction in single-cell foundation models. Experiments on cell-type annotation and gene regulatory network inference show competitive transfer performance.
This paper introduces LLantia, an LLM-based automated method for inferring neural circuit function from connectome data, demonstrated on the adult fruit fly brain.
This paper analyzes the Virtual Cell Challenge benchmark for held-out CRISPRi perturbation prediction, finding that simple magnitude-based scalar features outperform deep MLP encoders, and that magnitude-only predictors transfer better across cell types.
SynBio Studio es un IDE para biología sintética que permite a los científicos programar circuitos genéticos, mapear plásmidos en tiempo real y simular su cinética antes de realizar experimentos en laboratorio.
This bioRxiv preprint introduces Diff-Switch, a framework that uses diffusion-based ensemble sampling to generate conformational states for de novo protein switch design, improving the success rate of finding switch-compatible sequences.
An experience report from BOSC 2026 on using generative AI to pre-review open-source software submissions, with human reviewers making final decisions. Most reviewers found the AI-assisted pre-review useful but preferred to verify AI conclusions independently.
Arcee AI, Loka, AWS, and Prime Intellect post-trained an open model using reinforcement learning to improve scientific tool use and biological reasoning, achieving notable gains on drug tool and Gene Ontology benchmarks.