GSK’s new $110M AI deal shows why quality biological data > bigger models

Reddit r/ArtificialInteligence News

Summary

GSK expands its $110M AI drug discovery partnership with Relation Therapeutics, emphasizing that clean, disease-specific biological data from lab experiments is more valuable than larger models. The move highlights an industry-wide pivot toward proprietary data as a competitive moat in AI-driven pharma.

The $110M Deal: GSK and Relation Therapeutics are expanding their partnership. Relation will conduct lab experiments to produce large-scale cellular datasets to train AI models (like their MORGAN platform) for drug target discovery. Public Biological Data Hits a Wall: Unlike LLMs that improve as you feed them more web text, single-cell biological models plateau quickly on public databases (e.g., CZ CELLxGENE, Human Cell Atlas) due to lab noise, differing protocols, and data overlap. The "Lab-in-the-Loop" Model: Relation uses physical perturbation experiments alongside single-cell and spatial transcriptomics to measure exact cellular responses to genetic changes and drugs, feeding clean data directly back into ML models. Proven Concept: Relation has already used this strategy to build Osteomics, a proprietary single-cell bone atlas for studying osteoporosis. Industry-Wide Pivot: Big Pharma (including similar moves by AstraZeneca and Pathos AI) is realizing that clean, specialized, disease-specific data is becoming the ultimate moat in AI drug discovery.
Original Article

Similar Articles

Closing the data loop in AI-driven drug discovery

MIT Technology Review

AI accelerates drug discovery by predicting candidates and reducing costs, but success depends on high-quality data and integration with lab systems to close the data loop and validate predictions.