@nkwang24: 1/ Curious how protein language models such as @biohub's ESM-C represent structural information from sequence alone? Co…
Summary
Goodfire's Silico research platform is now publicly available, and a demo shows how protein language models like ESM-C represent structural information from sequence alone, enabling frontier-scale model interpretation and training.
View Cached Full Text
Cached at: 08/05/26, 06:20 AM
1/ Curious how protein language models such as @biohub’s ESM-C represent structural information from sequence alone? Come explore this interactive demo we built with @GoodfireAI’s new Silico research platform and see for yourself!
2/ We extend the categorical Jacobian approach in https://doi.org/10.1073/pnas.2406285121… to extract a layerwise readout of ESM-C’s predicted contacts, i.e. which sequence positions are likely to be close physically. Remarkable concordance with known structures despite not being trained on any.
3/ Looking at proteins with structural repeats such as beta propellers, we find that repeated elements overlap in a 3D PCA projection despite each structural repeat differing in sequence.
4/ Contact precision emerges late in the model — at a similar relative depth across model sizes — and continues to rise through the final layers.
5/ Residue-level structural properties also peak towards the final layers but are more legible throughout.
6/ Follow @GoodfireAI for more and try extending this analysis yourself today!
Similar Articles
Biohub/esm
Biohub releases ESMC, ESMFold2, and ESM Atlas — a world model for protein biology enabling state-of-the-art prediction, design, and discovery across scales, including a billion-structure atlas.
@SylvainGariel: Took me a while to figure out what all the ESMFold2 rage was about. At first, the benchmarking data didn't look super r…
ESMFold2 is an open-source AI model for protein structure prediction that achieves state-of-the-art performance on protein interactions and antibodies, with a massive structure database (ESM Atlas).
The Prose of Proteins - A Lesson in Taste and Vision through the Work of Brian Hie
This article profiles researcher Brian Hie, highlighting how his unique background in literature and computer science informed the development of ESM, a BERT-like model for protein sequences.
ProtSent: Protein Sentence Transformers
This article introduces ProtSent, a contrastive fine-tuning framework for protein language models that improves embedding quality for downstream tasks like remote homology detection and structural retrieval.
Structural Interpretations of Protein Language Model Representations via Differentiable Graph Partitioning
This paper proposes SoftBlobGIN, a framework that enhances the interpretability of protein language model representations by projecting them onto contact graphs for structure-aware message passing. It demonstrates improved performance on enzyme classification and binding-site detection while providing auditable structural explanations.