PairSAE: Mechanistic Interpretability from Pair Representations in Protein Co-Folding

arXiv cs.LG Papers

Summary

Introduces PairSAE, a method that adapts sparse autoencoders to interpret pairwise representations in protein co-folding models, enabling the discovery of interpretable features that align with biological annotations and predict binding affinities.

arXiv:2606.27440v1 Announce Type: new Abstract: Foundation models for structural biology have achieved remarkable performance in predicting biomolecular structure and show promise for the design of proteins and small molecules. Yet understanding which internal features drive their outputs remains challenging. Standard sparse autoencoders (SAEs), effective on transformer-style sequence embeddings, do not transfer cleanly to pairformer-like architectures: naively operating on pairwise representations yields a quadratic blow-up of features and obscures concepts distributed jointly across sequence and pair representations. We introduce PairSAE, which summarizes pairwise tensors via an N-mode SVD into token-wise interaction roles, then uses a sparse autoencoder to learn a shared set of token-level features that decode into both sequence and pair representations. Evaluated on Boltz-2 activations for PLINDER protein-ligand complexes, PairSAE yields interpretable features that align with UniProt annotations and predict Boltz-2 affinity values. These results indicate that PairSAE links the latent space of foundation models for structural biology to interpretable structural concepts, clarifying what the model "knows" while avoiding pairformer-induced pitfalls that limit conventional SAEs.
Original Article
View Cached Full Text

Cached at: 06/29/26, 05:22 AM

# PairSAE: Mechanistic Interpretability from Pair Representations in Protein Co-Folding
Source: [https://arxiv.org/html/2606.27440](https://arxiv.org/html/2606.27440)
Giosue Migliorini University of California, Irvine &Aristofanis Rontogiannis Flagship Pioneering &Grigori Guitchounts Flagship Pioneering &Nicholas Franklin Flagship Pioneering &Axel Elaldi Flagship Pioneering &Olivia Viessmann Flagship Pioneering

###### Abstract

Foundation models for structural biology have achieved remarkable performance in predicting biomolecular structure and show promise for the design of proteins and small molecules\. Yet understanding*which*internal features drive their outputs remains challenging\. Standard sparse autoencoders \(SAEs\), effective on transformer\-style*sequence embeddings*, do not transfer cleanly to pairformer\-like architectures: naïvely operating on pairwise representations yields a quadratic blow\-up of features and obscures concepts distributed jointly across sequence and pair representations\. We introducePairSAE, which summarizes pairwise tensors via anNN\-mode SVD into token\-wise interaction roles, then uses a sparse autoencoder to learn a*shared*set of token\-level features that decode into both sequence and pair representations\. Evaluated on Boltz\-2 activations for PLINDER protein–ligand complexes, PairSAE yields interpretable features that align with UniProt annotations and predict Boltz\-2 affinity values\. These results indicate that PairSAE links the latent space of foundation models for structural biology to interpretable structural concepts, clarifying what the model “knows” while avoiding pairformer\-induced pitfalls that limit conventional SAEs\.

## 1Introduction

Recent advances in protein structure prediction, most notably AlphaFold and the Boltz models\(Jumperet al\.,[2021](https://arxiv.org/html/2606.27440#bib.bib18); Abramsonet al\.,[2024](https://arxiv.org/html/2606.27440#bib.bib19); Wohlwendet al\.,[2025](https://arxiv.org/html/2606.27440#bib.bib20); Passaroet al\.,[2025](https://arxiv.org/html/2606.27440#bib.bib21)\), have transformed structural biology and now extend to DNA, RNA, small molecules, and their complexes\. These models achieve high accuracy without explicitly encoding physical or chemical laws, instead exploiting statistical regularities in large datasets of experimentally determined structures\. This raises a central question: to what extent do they implicitly capture the physics and chemistry that govern folding? Answering this requires interpretability – explanations that let researchers assess prediction reliability and biophysical plausibility – yet the highly nonlinear architectures of deep models make such analysis challenging\. Work in mechanistic interpretability shows that apparent “polysemantic” neurons can arise when networks represent more sparse features than available dimensions, forcing features into superposition rather than cleanly localizing them\(Elhageet al\.,[2022](https://arxiv.org/html/2606.27440#bib.bib29); Scherliset al\.,[2022](https://arxiv.org/html/2606.27440#bib.bib31)\)\. Sparse autoencoders \(SAEs\) address this by disentangling superposed features into more monosemantic components via sparse, overcomplete dictionaries\(Olshausen and Field,[1997](https://arxiv.org/html/2606.27440#bib.bib24); Ng and others,[2011](https://arxiv.org/html/2606.27440#bib.bib23); Elhageet al\.,[2022](https://arxiv.org/html/2606.27440#bib.bib29)\)\. Results in large language models\(Cunninghamet al\.,[2024](https://arxiv.org/html/2606.27440#bib.bib30); Brickenet al\.,[2023](https://arxiv.org/html/2606.27440#bib.bib28)\)and protein language models \(pLMs\)\(Simon and Zou,[2024](https://arxiv.org/html/2606.27440#bib.bib36); Adamset al\.,[2025](https://arxiv.org/html/2606.27440#bib.bib3)\)suggest that SAEs can recover interpretable latent features, making them a suitable candidate for probing the representations learned by protein structure prediction models

However, applying SAEs naively to structure prediction models is nontrivial\. Unlike most transformer\-based LLMs and pLMs, these models use pairformer blocks with pairwise representations alongside sequence embeddings\(Abramsonet al\.,[2024](https://arxiv.org/html/2606.27440#bib.bib19)\)\. Learning separate dictionaries per pair scales quadratically and obscures analysis, and repeated pair–sequence interactions suggest features may be superposed across both spaces\. To address this, we introducePairSAE, which reconstructs both sequence\-level and pairwise embeddings from a shared feature set\. Applied to Boltz\-2, PairSAE links latent representations to interpretable structural concepts, clarifying what the model “knows” about folding and enabling more principled evaluation and discovery in structural biology\.

![Refer to caption](https://arxiv.org/html/2606.27440v1/figures/transmembrane_struct.png)

![Refer to caption](https://arxiv.org/html/2606.27440v1/figures/dis_bond_struct.png)

![Refer to caption](https://arxiv.org/html/2606.27440v1/figures/protease.png)

![Refer to caption](https://arxiv.org/html/2606.27440v1/figures/transmembrane.png)

![Refer to caption](https://arxiv.org/html/2606.27440v1/figures/dis_bonds.png)

![Refer to caption](https://arxiv.org/html/2606.27440v1/figures/FT__Active_site__For_protease_activity_shared_with_dimeric_partner.png)

Figure 1:Interpretable features recovered by PairSAEs trained at layers 33 and 64 \(third recycling step\) of Boltz\-2\.Top left:feature 889 onE\. coliComplex II \(PDB 1NEK\); activates ontransmembraneproteins \(tokenF1=0\.58F\_\{1\}=0\.58, complexF1=0\.65F\_\{1\}=0\.65\)\.Top center:feature 55 on influenza N2 neuraminidase chain B \(PDB 4K1I\); activates ondisulfide bonds\(tokenF1=0\.79F\_\{1\}=0\.79, complexF1=0\.81F\_\{1\}=0\.81\)\.Top right:feature 885 on an HIV\-1 protease inhibitor \(PDB 1HPS\); activates onprotease active sitesshared with a dimeric partner \(tokenF1=0\.93F\_\{1\}=0\.93, complexF1=0\.93F\_\{1\}=0\.93\)\.Bottom:per\-residue activations with UniProt ground\-truth \(blue\)\.
## 2Background

### 2\.1Sparse autoencoders

Sparse dictionary learning\(Olshausen and Field,[1997](https://arxiv.org/html/2606.27440#bib.bib24); Leeet al\.,[2006](https://arxiv.org/html/2606.27440#bib.bib16); Ng and others,[2011](https://arxiv.org/html/2606.27440#bib.bib23)\)is a long\-standing representation learning technique that aims at finding a dictionary𝐃=\{𝐝j\}j=1D\\mathbf\{D\}=\\\{\\mathbf\{d\}\_\{j\}\\\}\_\{j=1\}^\{D\}approximating input vectors𝐱∈ℝn\\mathbf\{x\}\\in\\mathbb\{R\}^\{n\}via a linear combination of its columns with a set of sparse latent features𝐡∈ℝD\\mathbf\{h\}\\in\\mathbb\{R\}^\{D\}, extracted from𝐱\\mathbf\{x\}itself\. It has recently been shown that this formulation can yield interpretable features in transformed\-based models such as large language models\(Elhageet al\.,[2022](https://arxiv.org/html/2606.27440#bib.bib29); Templetonet al\.,[2024](https://arxiv.org/html/2606.27440#bib.bib35); Cunninghamet al\.,[2024](https://arxiv.org/html/2606.27440#bib.bib30)\), protein language models like ESM2\(Simon and Zou,[2024](https://arxiv.org/html/2606.27440#bib.bib36); Garcia and Ansuini,[2025](https://arxiv.org/html/2606.27440#bib.bib2); Adamset al\.,[2025](https://arxiv.org/html/2606.27440#bib.bib3); Gujralet al\.,[2025](https://arxiv.org/html/2606.27440#bib.bib4); Parsanet al\.,[2025](https://arxiv.org/html/2606.27440#bib.bib17)\), and foundation models for genomics like Evo 2\(Brixiet al\.,[2025](https://arxiv.org/html/2606.27440#bib.bib41)\)\.

Such features can be obtained by training a sparse autoencoder using a reconstruction loss with a sparsity constraint\. In mech interp the autoencoder is typically linear, with both a sparsity and positivity constraint on the latent features\. A common construction is

𝐡=σ​\(𝐄​𝐱\+𝐛enc\),𝐱^=𝐃​𝐡\+𝐛dec,\\mathbf\{h\}=\\sigma\\left\(\\mathbf\{E\}\\,\\mathbf\{x\}\+\\mathbf\{b\}\_\{\\text\{enc\}\}\\right\),\\quad\\hat\{\\mathbf\{x\}\}\\ =\\ \\mathbf\{D\}\\,\\mathbf\{h\}\\ \+\\ \\mathbf\{b\}\_\{\\text\{dec\}\},\(1\)where𝐄∈ℝD×n,𝐛enc∈ℝD\\mathbf\{E\}\\in\\mathbb\{R\}^\{D\\times n\},\\mathbf\{b\}\_\{\\text\{enc\}\}\\in\\mathbb\{R\}^\{D\}are the weights and bias terms for the encoder,𝐃∈ℝn×D,𝐛dec∈ℝn\\mathbf\{D\}\\in\\mathbb\{R\}^\{n\\times D\},\\mathbf\{b\}\_\{\\text\{dec\}\}\\in\\mathbb\{R\}^\{n\}are the weights and bias terms for the decoder, andσ\\sigmais a sparsity\-inducing nonlinearity such as ReLU\(Cunninghamet al\.,[2024](https://arxiv.org/html/2606.27440#bib.bib30); Brickenet al\.,[2023](https://arxiv.org/html/2606.27440#bib.bib28)\), TopK\(Gaoet al\.,[2025](https://arxiv.org/html/2606.27440#bib.bib27)\), BatchTopK\(Bussmannet al\.,[2024](https://arxiv.org/html/2606.27440#bib.bib37)\), and others\.

### 2\.2Pairformer representations

In a pairformer module\(Abramsonet al\.,[2024](https://arxiv.org/html/2606.27440#bib.bib19)\), an input sequence𝐱\\mathbf\{x\}of lengthNtokN\_\{\\text\{tok\}\}tokens is encoded into a sequence embedding matrix𝐒∈ℝNtok×ns\\mathbf\{S\}\\in\\mathbb\{R\}^\{N\_\{\\text\{tok\}\}\\times n\_\{s\}\}, whoseii\-th row is the token embedding vector𝐬i∈ℝns\\mathbf\{s\}\_\{i\}\\in\\mathbb\{R\}^\{n\_\{s\}\}, and a three\-dimensional pair representation tensor𝒵∈ℝNtok×Ntok×nz\\mathcal\{Z\}\\in\\mathbb\{R\}^\{N\_\{\\text\{tok\}\}\\times N\_\{\\text\{tok\}\}\\times n\_\{z\}\}, whose\(i,j\)\(i,j\)\-th slice𝒵i,j∈ℝnz\\mathcal\{Z\}\_\{i,j\}\\in\\mathbb\{R\}^\{n\_\{z\}\}represents the embedding of token pair\(i,j\)\(i,j\)\. Here,nsn\_\{s\}andnzn\_\{z\}are the sequence\-level and pairwise embedding dimensions, respectively\. Pair representations control the information flow in the updates of the sequence\-level embeddings at each layer of the pairformer, by biasing the attention logitsAbramsonet al\.\([2024](https://arxiv.org/html/2606.27440#bib.bib19)\)\. Pair representations were first introduced with the Evoformer, part of the architecture of AlphaFold 2Jumperet al\.\([2021](https://arxiv.org/html/2606.27440#bib.bib18)\)\.

## 3PairSAE

#### Motivation\.

Our goal is to obtain token\-level sparse features that jointly reconstruct sequence and pair representations\. Pairwise embeddings impose no structure, so we first compress them into a token\-wise summary that preserves row/column interaction roles\. A shared feature set then decodes into both representation types\.

#### Summarizing pairwise interactions\.

We propose to perform anN−N\-mode singular value decomposition \(SVD\) of the tensor𝒵\\mathcal\{Z\}, and use it to construct a sequence representation that is specifically designed to incorporate information about how each token interacts in the sequence\. We include a brief introduction toN−N\-mode SVD in Appendix[A](https://arxiv.org/html/2606.27440#A1)\. Information about the role that each token plays*as a row and as a column*in𝒵\\mathcal\{Z\}can be obtained by inspecting the mode\-1 and mode\-2 matrices of left singular vectors𝐔\(1\),𝐔\(2\)∈ℝNtok×Ntok\\mathbf\{U\}^\{\(1\)\},\\mathbf\{U\}^\{\(2\)\}\\in\\mathbb\{R\}^\{N\_\{\\text\{tok\}\}\\times N\_\{\\text\{tok\}\}\}\(De Lathauweret al\.,[2000a](https://arxiv.org/html/2606.27440#bib.bib33); Vasilescu and Terzopoulos,[2002](https://arxiv.org/html/2606.27440#bib.bib32)\)\. We consider a truncation to the firstrrcolumns𝐔:,1:r\(1\)\\mathbf\{U\}^\{\(1\)\}\_\{:,1:r\}and𝐔:,1:r\(2\)\\mathbf\{U\}^\{\(2\)\}\_\{:,1:r\}, ordered by their singular value\. We then concatenate them column\-wise into a new sequence level embedding

𝐦i=\[𝐔i,1:r\(1\)∥𝐔i,1:r\(2\)\],i=1,…,Ntok\.\\mathbf\{m\}\_\{i\}=\\big\[\\mathbf\{U\}^\{\(1\)\}\_\{i,1:r\}\\;\\\|\\;\\mathbf\{U\}^\{\(2\)\}\_\{i,1:r\}\\big\],\\quad i=1,\\dots,N\_\{\\text\{tok\}\}\.\(2\)

#### Autoencoder\.

PairSAE latent features are obtained by encoding a concatenation of this new embedding matrix with the original sequence\-level representations from the pairformer model:

𝐡i=σ\(𝐄\[𝐬i∥𝐦i\]\+𝐛enc\),,i=1,…,Ntok,\\mathbf\{h\}\_\{i\}=\\sigma\\left\(\\mathbf\{E\}\\,\[\\mathbf\{s\}\_\{i\}\\;\\\|\\;\\mathbf\{m\}\_\{i\}\]\+\\mathbf\{b\}\_\{\\text\{enc\}\}\\right\),,\\quad i=1,\\dots,N\_\{\\text\{tok\}\},\(3\)where𝐄∈ℝD×\(ns\+2​r\),𝐛enc∈ℝD,\\mathbf\{E\}\\in\\mathbb\{R\}^\{D\\times\(n\_\{s\}\+2r\)\},\\,\\mathbf\{b\}\_\{\\text\{enc\}\}\\in\\mathbb\{R\}^\{D\},andσ\\sigmais a composition of BatchTopK and ReLU\(Bussmannet al\.,[2024](https://arxiv.org/html/2606.27440#bib.bib37)\)\. We decode with shared features into both spaces:

𝐬^i=𝐃s​𝐡i\+𝐛decs,𝐳^i,j=𝐃zrow​𝐡i\+𝐃zcol​𝐡j\+𝐛decz,i,j∈\{1,…,Ntok\},\\hat\{\\mathbf\{s\}\}\_\{i\}=\\mathbf\{D\}^\{s\}\\,\\mathbf\{h\}\_\{i\}\+\\mathbf\{b\}^\{s\}\_\{\\text\{dec\}\},\\quad\\hat\{\\mathbf\{z\}\}\_\{i,j\}=\\mathbf\{D\}^\{z\_\{\\text\{row\}\}\}\\mathbf\{h\}\_\{i\}\+\\mathbf\{D\}^\{z\_\{\\text\{col\}\}\}\\mathbf\{h\}\_\{j\}\+\\mathbf\{b\}^\{z\}\_\{\\text\{dec\}\},\\quad i,j\\in\\\{1,\\dots,N\_\{\\text\{tok\}\}\\\},\(4\)where𝐃s∈ℝns×D,𝐛decs∈ℝns,𝐃zrow,𝐃zcol∈ℝnz×D,\\mathbf\{D\}^\{s\}\\in\\mathbb\{R\}^\{n\_\{s\}\\times D\}\\,,\\mathbf\{b\}^\{s\}\_\{\\text\{dec\}\}\\in\\mathbb\{R\}^\{n\_\{s\}\},\\,\\mathbf\{D\}^\{z\_\{\\text\{row\}\}\},\\mathbf\{D\}^\{z\_\{\\text\{col\}\}\}\\in\\mathbb\{R\}^\{n\_\{z\}\\times D\},and𝐛decz∈ℝnz\\mathbf\{b\}^\{z\}\_\{\\text\{dec\}\}\\in\\mathbb\{R\}^\{n\_\{z\}\}\.

#### Loss function\.

Choosing the dictionary sizeDDcan have major consequences on the types of features that are learned, and a larger dictionary does not necessarily translate to better performance on downstream tasks\(Templetonet al\.,[2024](https://arxiv.org/html/2606.27440#bib.bib35); Gaoet al\.,[2025](https://arxiv.org/html/2606.27440#bib.bib27)\)\. It has recently been shown that a Matryoshka SAE loss\(Bussmannet al\.,[2025](https://arxiv.org/html/2606.27440#bib.bib38)\)exhibits robust performance across dictionary sizes\. This is attained by slicing the decoder matrix and the feature vector across several nested levels \(we use three nested widthsc1<c2<c3=Dc\_\{1\}<c\_\{2\}<c\_\{3\}=D\), and summing them to compute the objectives

ℒ​\(𝐬i\)≔∑k=13‖𝐬i−𝐃\[:,1:ck\]s​𝐡i,\[1:ck\]−𝐛decs‖22,\\displaystyle\\mathcal\{L\}\(\\mathbf\{s\}\_\{i\}\)\\coloneq\\sum\_\{k=1\}^\{3\}\\left\\\|\\mathbf\{s\}\_\{i\}\-\\mathbf\{D\}^\{s\}\_\{\[:,\\,1:c\_\{k\}\]\}\\mathbf\{h\}\_\{i,\[1:c\_\{k\}\]\}\-\\mathbf\{b\}^\{s\}\_\{\\text\{dec\}\}\\right\\\|^\{2\}\_\{2\},\(5\)ℒ​\(𝐳i​j\)≔∑k=13‖𝐳i​j−𝐃\[:,1:ck\]zrow​𝐡i,\[1:ck\]−𝐃\[:,1:ck\]zcol​𝐡j,\[1:ck\]−𝐛decz‖22\\displaystyle\\mathcal\{L\}\(\\mathbf\{z\}\_\{ij\}\)\\coloneq\\sum\_\{k=1\}^\{3\}\\left\\\|\\mathbf\{z\}\_\{ij\}\-\\mathbf\{D\}^\{z\_\{\\text\{row\}\}\}\_\{\[:,\\,1:c\_\{k\}\]\}\\mathbf\{h\}\_\{i,\[1:c\_\{k\}\]\}\-\\mathbf\{D\}^\{z\_\{\\text\{col\}\}\}\_\{\[:,\\,1:c\_\{k\}\]\}\\mathbf\{h\}\_\{j,\[1:c\_\{k\}\]\}\-\\mathbf\{b\}^\{z\}\_\{\\text\{dec\}\}\\right\\\|^\{2\}\_\{2\}\(6\)Following\(Gaoet al\.,[2025](https://arxiv.org/html/2606.27440#bib.bib27)\), we include an auxiliary loss functionℒaux\\mathcal\{L\}\_\{\\text\{aux\}\}that can "revitalize" dead features, resulting in the loss

ℒ​\(𝐒,𝒵\)=∑i=1Ntokℒ​\(𝐬i\)\+λpair​∑j=1Ntokℒ​\(𝐳i​j\)\+λaux​ℒaux\.\\mathcal\{L\}\(\\mathbf\{S\},\\mathcal\{Z\}\)=\\sum\_\{i=1\}^\{N\_\{\\text\{tok\}\}\}\\mathcal\{L\}\(\\mathbf\{s\}\_\{i\}\)\+\\lambda\_\{\\text\{pair\}\}\\sum\_\{j=1\}^\{N\_\{\\text\{tok\}\}\}\\mathcal\{L\}\(\\mathbf\{z\}\_\{ij\}\)\+\\lambda\_\{\\text\{aux\}\}\\mathcal\{L\}\_\{\\text\{aux\}\}\.\\\(7\)We speed up training by we computing a Monte Carlo approximation of the second summation and only consider a singlejjfor eachii\. We setλpair=Ntok−1\\lambda\_\{\\text\{pair\}\}=N\_\{\\text\{tok\}\}^\{\-1\}to have both representation be equally weighted\.

![Refer to caption](https://arxiv.org/html/2606.27440v1/figures/probe_domain.png)

![Refer to caption](https://arxiv.org/html/2606.27440v1/figures/probe_token.png)

Figure 2:Count of concepts withF1≥0\.5F\_\{1\}\\geq 0\.5using complex\-level recall \(left\) and token level \(right\), grouped by UniProt annotation category \(test\-set counts in parentheses\)\.Table 1:Percentage of concepts in the test set withF1F\_\{1\}≥\\geq0\.5\.

## 4Results

We evaluate PairSAE on Boltz\-2 representations for protein–ligand complexes from PLINDER\(Durairajet al\.,[2024](https://arxiv.org/html/2606.27440#bib.bib26)\)\. We train two models at the third recycling step, at layers 33 and 64 \(R3\-L33, R3\-L64\)\. Interpretability can be assessed via linear probes\(Gaoet al\.,[2025](https://arxiv.org/html/2606.27440#bib.bib27); Simon and Zou,[2024](https://arxiv.org/html/2606.27440#bib.bib36)\), and we test whether features predict UniProt residue annotations\(UniProt,[2025](https://arxiv.org/html/2606.27440#bib.bib6)\)and PLINDER system annotations\.

Boltz\-2 demonstrated a strong improvement on the speed/accuracy Pareto frontier for binding affinity prediction\(Passaroet al\.,[2025](https://arxiv.org/html/2606.27440#bib.bib21)\), and has the potential of further accelerating virtual screenings\. We aim to identify features that exhibit strong association with the predicted affinity, and to do so we build on the hypothesis generation techniques laid out inMovvaet al\.\([2025](https://arxiv.org/html/2606.27440#bib.bib1)\)\. Details of each experiment can be found in Appendix[B](https://arxiv.org/html/2606.27440#A2)\.

### 4\.1Linear probing

Following\(Simon and Zou,[2024](https://arxiv.org/html/2606.27440#bib.bib36)\), we normalize each feature to\[0,1\]\[0,1\], fit a single\-threshold classifier per concept–feature pair via a grid search with candidate thresholds spaced at increments of 0\.1 across the unit interval, and select the best feature by validationF1F\_\{1\}\. As noted inSimon and Zou \([2024](https://arxiv.org/html/2606.27440#bib.bib36)\), it is often the case that latent features at a single token might capture properties of the entire system, hence we followSimon and Zou \([2024](https://arxiv.org/html/2606.27440#bib.bib36)\)and compute theF1F\_\{1\}score by considering the recall at the complex level as well as the token level, resulting in two separate metrics\. For each metric, we display a count of how many features we could predict with a score above 0\.5 on a test set in Figure[2](https://arxiv.org/html/2606.27440#S3.F2)\. We find that PairSAE features are highly predictive of UniProt annotations, as shown in Table[1](https://arxiv.org/html/2606.27440#S3.T1), and outperform classifiers based on neurons from the last layer of ESM\-2\-650M, considered powerful embeddings for many downstream applications\(Linet al\.,[2023](https://arxiv.org/html/2606.27440#bib.bib13)\)\.

### 4\.2Hypothesis generation: explaining binding affinity predictions

Because affinity is defined at the complex level, we form a complex embedding by max\-pooling features over tokens:h^k=maxi≤Ntok⁡hi,k\\hat\{h\}\_\{k\}=\\max\_\{i\\leq N\_\{\\text\{tok\}\}\}h\_\{i,k\},k=1,…,Dk=1,\\dots,D\. We then predictA​VAVvia LASSO\(Tibshirani,[1996](https://arxiv.org/html/2606.27440#bib.bib15); Movvaet al\.,[2025](https://arxiv.org/html/2606.27440#bib.bib1)\), minimizing on a training set𝒟\\mathcal\{D\}of 980 PLINDER complexes:

∑j∈𝒟‖A​Vj−𝜷​𝐡^j‖22\+λ​\|𝜷\|1\.\\sum\_\{j\\in\\mathcal\{D\}\}\\left\\\|AV\_\{j\}\-\\boldsymbol\{\\beta\}\\,\\mathbf\{\\hat\{h\}\}\_\{j\}\\right\\\|^\{2\}\_\{2\}\+\\lambda\\left\|\\boldsymbol\{\\beta\}\\right\|\_\{1\}\.\(8\)We test our results on Posebusters\(Buttenschoenet al\.,[2024](https://arxiv.org/html/2606.27440#bib.bib12)\)\.

Our first finding is on the predictive power of the PairSAE features on affinity values: by choosingλ\\lambdausing cross\-validation, we can predictA​VAVwith anR2R^\{2\}of 0\.528 from R3\-L64 using 291 features out of 16,384\. Errors increase at higher predicted affinity, partly due to label shift \(fewer high\-affinity examples in training than in PoseBusters; see Fig\.[3](https://arxiv.org/html/2606.27440#S4.F3)\)\.

Moreover, by lettingλ\\lambdaincrease we can highlight features that carry most of the predictive power\. By looking at the top 10 most influential features at both layers we find most of them to be activated on ligands, for which we have very limited annotations\. In Figure[3](https://arxiv.org/html/2606.27440#S4.F3)we highlight a feature from R3\-L64, selected by settingλ\\lambdain \([8](https://arxiv.org/html/2606.27440#S4.E8)\) such that only one feature has nonzero coefficient\. This feature activates on complexes with higher binding affinity, with a strong group difference when compared to inactive complexes\. Additional results are reported in Appendix[C](https://arxiv.org/html/2606.27440#A3)\.

![Refer to caption](https://arxiv.org/html/2606.27440v1/figures/lasso_out.png)

![Refer to caption](https://arxiv.org/html/2606.27440v1/figures/group_diff_2299_test.png)

![Refer to caption](https://arxiv.org/html/2606.27440v1/figures/affinity_feat_2299.png)

Figure 3:Left:negative affinity values from Boltz\-2 \(higher means more affine\)\(Passaroet al\.,[2025](https://arxiv.org/html/2606.27440#bib.bib21)\), and predicted values from the LASSO regression in \([8](https://arxiv.org/html/2606.27440#S4.E8)\)\.Center:feature 2299 from R3\-L64, displaying strong group difference in affinity values measured by Welch t\-test\.Right:values of feature 2299 overlaid on a ligand where it is highly activated\.

## 5Conclusion

We introduced PairSAE, a sparse dictionary learning method that discovers a shared feature basis jointly explaining sequence and pair representations in pairformer\-style models\. We showed that these features encode biophysically meaningful concepts and can surface hypotheses about a model’s inner workings for downstream tasks such as binding\-affinity prediction\. Future work includes scaling up the study to more layers and recycling steps, integrating PairSAE features into automated interpretability pipelines to accelerate concept discovery, and developing steering methods for interpretable protein design via feature interventions\. A key limitation is that we did not map ligand\-activating sparse features to specific concepts, due to the lack of ligand annotations\.

## References

- J\. Abramson, J\. Adler, J\. Dunger, R\. Evans, T\. Green, A\. Pritzel, O\. Ronneberger, L\. Willmore, A\. J\. Ballard, J\. Bambrick,et al\.\(2024\)Accurate structure prediction of biomolecular interactions with alphafold 3\.Nature630\(8016\),pp\. 493–500\.Cited by:[§1](https://arxiv.org/html/2606.27440#S1.p1.1),[§1](https://arxiv.org/html/2606.27440#S1.p2.1),[§2\.2](https://arxiv.org/html/2606.27440#S2.SS2.p1.11)\.
- E\. Adams, L\. Bai, M\. Lee, Y\. Yu, and M\. AlQuraishi \(2025\)From mechanistic interpretability to mechanistic biology: training, evaluating, and interpreting sparse autoencoders on protein language models\.Forty\-second International Conference on Machine Learning\.Cited by:[§1](https://arxiv.org/html/2606.27440#S1.p1.1),[§2\.1](https://arxiv.org/html/2606.27440#S2.SS1.p1.4)\.
- T\. Bricken, A\. Templeton, J\. Batson, B\. Chen, A\. Jermyn, T\. Conerly, N\. Turner, C\. Anil, C\. Denison, A\. Askell, R\. Lasenby, Y\. Wu, S\. Kravec, N\. Schiefer, T\. Maxwell, N\. Joseph, Z\. Hatfield\-Dodds, A\. Tamkin, K\. Nguyen, B\. McLean, J\. E\. Burke, T\. Hume, S\. Carter, T\. Henighan, and C\. Olah \(2023\)Towards monosemanticity: decomposing language models with dictionary learning\.Note:[https://transformer\-circuits\.pub/2023/monosemantic\-features/index\.html](https://transformer-circuits.pub/2023/monosemantic-features/index.html)Transformer Circuits ThreadCited by:[§1](https://arxiv.org/html/2606.27440#S1.p1.1),[§2\.1](https://arxiv.org/html/2606.27440#S2.SS1.p2.3)\.
- G\. Brixi, M\. G\. Durrant, J\. Ku, M\. Poli, G\. Brockman, D\. Chang, G\. A\. Gonzalez, S\. H\. King, D\. B\. Li, A\. T\. Merchant,et al\.\(2025\)Genome modeling and design across all domains of life with evo 2\.BioRxiv,pp\. 2025–02\.Cited by:[§2\.1](https://arxiv.org/html/2606.27440#S2.SS1.p1.4)\.
- B\. Bussmann, P\. Leask, and N\. Nanda \(2024\)Batchtopk sparse autoencoders\.arXiv preprint arXiv:2412\.06410\.Cited by:[§2\.1](https://arxiv.org/html/2606.27440#S2.SS1.p2.3),[§3](https://arxiv.org/html/2606.27440#S3.SS0.SSS0.Px3.p1.2)\.
- B\. Bussmann, N\. Nabeshima, A\. Karvonen, and N\. Nanda \(2025\)Learning multi\-level features with matryoshka sparse autoencoders\.InForty\-second International Conference on Machine Learning,Cited by:[§3](https://arxiv.org/html/2606.27440#S3.SS0.SSS0.Px4.p1.2)\.
- M\. Buttenschoen, G\. M\. Morris, and C\. M\. Deane \(2024\)PoseBusters: ai\-based docking methods fail to generate physically valid poses or generalise to novel sequences\.Chemical Science15\(9\),pp\. 3130–3139\.Cited by:[§B\.3](https://arxiv.org/html/2606.27440#A2.SS3.p1.1),[§4\.2](https://arxiv.org/html/2606.27440#S4.SS2.p1.5)\.
- H\. Cunningham, A\. Ewart, L\. Riggs, R\. Huben, and L\. Sharkey \(2024\)Sparse autoencoders find highly interpretable features in language models\.ICLR\.Cited by:[§1](https://arxiv.org/html/2606.27440#S1.p1.1),[§2\.1](https://arxiv.org/html/2606.27440#S2.SS1.p1.4),[§2\.1](https://arxiv.org/html/2606.27440#S2.SS1.p2.3)\.
- L\. De Lathauwer, B\. De Moor, and J\. Vandewalle \(2000a\)A multilinear singular value decomposition\.SIAM journal on Matrix Analysis and Applications21\(4\),pp\. 1253–1278\.Cited by:[Appendix A](https://arxiv.org/html/2606.27440#A1.p1.4),[§3](https://arxiv.org/html/2606.27440#S3.SS0.SSS0.Px2.p1.8)\.
- L\. De Lathauwer, B\. De Moor, and J\. Vandewalle \(2000b\)On the best rank\-1 and rank\-\(R1,R2,…,RnR\_\{1\},R\_\{2\},\.\.\.,R\_\{n\}\) approximation of higher\-order tensors\.SIAM journal on Matrix Analysis and Applications21\(4\),pp\. 1324–1342\.Cited by:[Appendix A](https://arxiv.org/html/2606.27440#A1.p2.3)\.
- J\. Durairaj, Y\. Adeshina, Z\. Cao, X\. Zhang, V\. Oleinikovas, T\. Duignan, Z\. McClure, X\. Robin, G\. Studer, D\. Kovtun,et al\.\(2024\)PLINDER: the protein\-ligand interactions dataset and evaluation resource\.bioRxiv,pp\. 2024–07\.Cited by:[§B\.1](https://arxiv.org/html/2606.27440#A2.SS1.p2.2),[§4](https://arxiv.org/html/2606.27440#S4.p1.1)\.
- N\. Elhage, T\. Hume, C\. Olsson, N\. Schiefer, T\. Henighan, S\. Kravec, Z\. Hatfield\-Dodds, R\. Lasenby, D\. Drain, C\. Chen,et al\.\(2022\)Toy models of superposition\.arXiv preprint arXiv:2209\.10652\.Cited by:[§1](https://arxiv.org/html/2606.27440#S1.p1.1),[§2\.1](https://arxiv.org/html/2606.27440#S2.SS1.p1.4)\.
- L\. Gao, T\. D\. la Tour, H\. Tillman, G\. Goh, R\. Troll, A\. Radford, I\. Sutskever, J\. Leike, and J\. Wu \(2025\)Scaling and evaluating sparse autoencoders\.InThe Thirteenth International Conference on Learning Representations,Cited by:[§2\.1](https://arxiv.org/html/2606.27440#S2.SS1.p2.3),[§3](https://arxiv.org/html/2606.27440#S3.SS0.SSS0.Px4.p1.2),[§3](https://arxiv.org/html/2606.27440#S3.SS0.SSS0.Px4.p1.3),[§4](https://arxiv.org/html/2606.27440#S4.p1.1)\.
- E\. N\. V\. Garcia and A\. Ansuini \(2025\)Interpreting and steering protein language models through sparse autoencoders\.arXiv preprint arXiv:2502\.09135\.Cited by:[§2\.1](https://arxiv.org/html/2606.27440#S2.SS1.p1.4)\.
- O\. Gujral, M\. Bafna, E\. Alm, and B\. Berger \(2025\)Sparse autoencoders uncover biologically interpretable features in protein language model representations\.Proceedings of the National Academy of Sciences122\(34\),pp\. e2506316122\.Cited by:[§2\.1](https://arxiv.org/html/2606.27440#S2.SS1.p1.4)\.
- M\. Ishteva, P\. Absil, S\. Van Huffel, and L\. De Lathauwer \(2011\)Tucker compression and local optima\.Chemometrics and Intelligent Laboratory Systems106\(1\),pp\. 57–64\.Cited by:[Appendix A](https://arxiv.org/html/2606.27440#A1.p2.3)\.
- J\. Jumper, R\. Evans, A\. Pritzel, T\. Green, M\. Figurnov, O\. Ronneberger, K\. Tunyasuvunakool, R\. Bates, A\. Žídek, A\. Potapenko,et al\.\(2021\)Highly accurate protein structure prediction with alphafold\.nature596\(7873\),pp\. 583–589\.Cited by:[§1](https://arxiv.org/html/2606.27440#S1.p1.1),[§2\.2](https://arxiv.org/html/2606.27440#S2.SS2.p1.11)\.
- D\. P\. Kingma and J\. Ba \(2014\)Adam: a method for stochastic optimization\.arXiv preprint arXiv:1412\.6980\.Cited by:[§B\.1](https://arxiv.org/html/2606.27440#A2.SS1.p2.2)\.
- H\. Lee, A\. Battle, R\. Raina, and A\. Ng \(2006\)Efficient sparse coding algorithms\.Advances in neural information processing systems19\.Cited by:[§2\.1](https://arxiv.org/html/2606.27440#S2.SS1.p1.4)\.
- Z\. Lin, H\. Akin, R\. Rao, B\. Hie, Z\. Zhu, W\. Lu, N\. Smetanin, R\. Verkuil, O\. Kabeli, Y\. Shmueli,et al\.\(2023\)Evolutionary\-scale prediction of atomic\-level protein structure with a language model\.Science379\(6637\),pp\. 1123–1130\.Cited by:[§4\.1](https://arxiv.org/html/2606.27440#S4.SS1.p1.3)\.
- R\. Movva, K\. Peng, N\. Garg, J\. Kleinberg, and E\. Pierson \(2025\)Sparse autoencoders for hypothesis generation\.InForty\-second International Conference on Machine Learning,Cited by:[§4\.2](https://arxiv.org/html/2606.27440#S4.SS2.p1.4),[§4](https://arxiv.org/html/2606.27440#S4.p2.1)\.
- A\. Nget al\.\(2011\)Sparse autoencoder\.CS294A Lecture notes72\(2011\),pp\. 1–19\.Cited by:[§1](https://arxiv.org/html/2606.27440#S1.p1.1),[§2\.1](https://arxiv.org/html/2606.27440#S2.SS1.p1.4)\.
- B\. A\. Olshausen and D\. J\. Field \(1997\)Sparse coding with an overcomplete basis set: a strategy employed by v1?\.Vision research37\(23\),pp\. 3311–3325\.Cited by:[§1](https://arxiv.org/html/2606.27440#S1.p1.1),[§2\.1](https://arxiv.org/html/2606.27440#S2.SS1.p1.4)\.
- N\. Parsan, D\. J\. Yang, and J\. J\. Yang \(2025\)Towards interpretable protein structure prediction with sparse autoencoders\.InLearning Meaningful Representations of Life \(LMRL\) Workshop at ICLR 2025,Cited by:[§2\.1](https://arxiv.org/html/2606.27440#S2.SS1.p1.4)\.
- S\. Passaro, G\. Corso, J\. Wohlwend, M\. Reveiz, S\. Thaler, V\. R\. Somnath, N\. Getz, T\. Portnoi, J\. Roy, H\. Stark,et al\.\(2025\)Boltz\-2: towards accurate and efficient binding affinity prediction\.BioRxiv,pp\. 2025–06\.Cited by:[§1](https://arxiv.org/html/2606.27440#S1.p1.1),[Figure 3](https://arxiv.org/html/2606.27440#S4.F3),[§4](https://arxiv.org/html/2606.27440#S4.p2.1)\.
- F\. Pedregosa, G\. Varoquaux, A\. Gramfort, V\. Michel, B\. Thirion, O\. Grisel, M\. Blondel, P\. Prettenhofer, R\. Weiss, V\. Dubourg, J\. Vanderplas, A\. Passos, D\. Cournapeau, M\. Brucher, M\. Perrot, and E\. Duchesnay \(2011\)Scikit\-learn: machine learning in Python\.Journal of Machine Learning Research12,pp\. 2825–2830\.Cited by:[§B\.3](https://arxiv.org/html/2606.27440#A2.SS3.p2.1)\.
- S\. Salentin, S\. Schreiber, V\. J\. Haupt, M\. F\. Adasme, and M\. Schroeder \(2015\)PLIP: fully automated protein–ligand interaction profiler\.Nucleic acids research43\(W1\),pp\. W443–W447\.Cited by:[§B\.2](https://arxiv.org/html/2606.27440#A2.SS2.p2.4)\.
- A\. Scherlis, K\. Sachan, A\. S\. Jermyn, J\. Benton, and B\. Shlegeris \(2022\)Polysemanticity and capacity in neural networks\.arXiv preprint arXiv:2210\.01892\.Cited by:[§1](https://arxiv.org/html/2606.27440#S1.p1.1)\.
- E\. Simon and J\. Zou \(2024\)Interplm: discovering interpretable features in protein language models via sparse autoencoders\.bioRxiv,pp\. 2024–11\.Cited by:[§1](https://arxiv.org/html/2606.27440#S1.p1.1),[§2\.1](https://arxiv.org/html/2606.27440#S2.SS1.p1.4),[§4\.1](https://arxiv.org/html/2606.27440#S4.SS1.p1.3),[§4](https://arxiv.org/html/2606.27440#S4.p1.1)\.
- A\. Templeton, T\. Conerly, J\. Marcus, J\. Lindsey, T\. Bricken, B\. Chen, A\. Pearce, C\. Citro, E\. Ameisen, A\. Jones, H\. Cunningham, N\. L\. Turner, C\. McDougall, M\. MacDiarmid, C\. D\. Freeman, T\. R\. Sumers, E\. Rees, J\. Batson, A\. Jermyn, S\. Carter, C\. Olah, and T\. Henighan \(2024\)Scaling monosemanticity: extracting interpretable features from claude 3 sonnet\.Transformer Circuits Thread\.External Links:[Link](https://transformer-circuits.pub/2024/scaling-monosemanticity/index.html)Cited by:[§2\.1](https://arxiv.org/html/2606.27440#S2.SS1.p1.4),[§3](https://arxiv.org/html/2606.27440#S3.SS0.SSS0.Px4.p1.2)\.
- R\. Tibshirani \(1996\)Regression shrinkage and selection via the lasso\.Journal of the Royal Statistical Society Series B: Statistical Methodology58\(1\),pp\. 267–288\.Cited by:[§4\.2](https://arxiv.org/html/2606.27440#S4.SS2.p1.4)\.
- UniProt \(2025\)UniProt: the universal protein knowledgebase in 2025\.Nucleic acids research53\(D1\),pp\. D609–D617\.Cited by:[§B\.2](https://arxiv.org/html/2606.27440#A2.SS2.p1.1),[§4](https://arxiv.org/html/2606.27440#S4.p1.1)\.
- N\. Vannieuwenhoven, R\. Vandebril, and K\. Meerbergen \(2012\)A new truncation strategy for the higher\-order singular value decomposition\.SIAM Journal on Scientific Computing34\(2\),pp\. A1027–A1052\.Cited by:[Appendix A](https://arxiv.org/html/2606.27440#A1.p2.3)\.
- M\. A\. O\. Vasilescu and D\. Terzopoulos \(2002\)Multilinear analysis of image ensembles: tensorfaces\.InEuropean conference on computer vision,pp\. 447–460\.Cited by:[Appendix A](https://arxiv.org/html/2606.27440#A1.p1.13),[§3](https://arxiv.org/html/2606.27440#S3.SS0.SSS0.Px2.p1.8)\.
- S\. Velankar, J\. M\. Dana, J\. Jacobsen, G\. Van Ginkel, P\. J\. Gane, J\. Luo, T\. J\. Oldfield, C\. O’Donovan, M\. Martin, and G\. J\. Kleywegt \(2012\)SIFTS: structure integration with function, taxonomy and sequences resource\.Nucleic acids research41\(D1\),pp\. D483–D489\.Cited by:[§B\.2](https://arxiv.org/html/2606.27440#A2.SS2.p1.1)\.
- J\. Wohlwend, G\. Corso, S\. Passaro, N\. Getz, M\. Reveiz, K\. Leidal, W\. Swiderski, L\. Atkinson, T\. Portnoi, I\. Chinn,et al\.\(2025\)Boltz\-1 democratizing biomolecular interaction modeling\.BioRxiv,pp\. 2024–11\.Cited by:[§1](https://arxiv.org/html/2606.27440#S1.p1.1)\.

## Appendix ANN\-mode Singular Value Decomposition

The mode\-kkunfolding of a tensor𝒯∈ℝn1×n2×⋯×nN\\mathcal\{T\}\\in\\mathbb\{R\}^\{n\_\{1\}\\times n\_\{2\}\\times\\dots\\times n\_\{N\}\}is obtained by flattening all but thekthk^\{\\text\{th\}\}dimension into a matrix𝒯\(k\)∈ℝnk×∏i≠kni\\mathcal\{T\}^\{\(k\)\}\\in\\mathbb\{R\}^\{n\_\{k\}\\times\\prod\_\{i\\neq k\}n\_\{i\}\}\. FollowingDe Lathauweret al\.\[[2000a](https://arxiv.org/html/2606.27440#bib.bib33)\], every tensor admits the higher\-order singular value decomposition

𝒯=𝒞×1𝐔\(1\)×2𝐔\(2\)​⋯×N𝐔\(N\),\\mathcal\{T\}=\\mathcal\{C\}\\times\_\{1\}\\mathbf\{U\}^\{\(1\)\}\\times\_\{2\}\\mathbf\{U\}^\{\(2\)\}\\dots\\times\_\{N\}\\mathbf\{U\}^\{\(N\)\},where𝒞∈ℝn1×n2×⋯×nN\\mathcal\{C\}\\in\\mathbb\{R\}^\{n\_\{1\}\\times n\_\{2\}\\times\\dots\\times n\_\{N\}\}is a core tensor,×k\\times\_\{k\}denotes the mode\-kktensor–matrix product, and each𝐔\(k\)∈ℝnk×nk\\mathbf\{U\}^\{\(k\)\}\\in\\mathbb\{R\}^\{n\_\{k\}\\times n\_\{k\}\}is an orthogonal matrix obtained from the left\-singular vectors of the mode\-kkunfolding𝒯\(k\)\\mathcal\{T\}^\{\(k\)\}\[Vasilescu and Terzopoulos,[2002](https://arxiv.org/html/2606.27440#bib.bib32)\]\. We refer to𝐔\(k\)\\mathbf\{U\}^\{\(k\)\}as the mode\-kkleft singular vectors of𝒯\\mathcal\{T\}\.

Unlike matrices, truncating the ordered left\-singular vectors to theirrk<nkr\_\{k\}<n\_\{k\}columns and fitting a new, smaller core tensor does not yield an optimal \(r1,r2,…,rNr\_\{1\},r\_\{2\},\\dots,r\_\{N\}\)\-rank approximation of𝒯\\mathcal\{T\}in the Frobenius norm\[De Lathauweret al\.,[2000b](https://arxiv.org/html/2606.27440#bib.bib34)\]\. Multiple algorithms have been proposed to improve upon a naive truncation\[De Lathauweret al\.,[2000b](https://arxiv.org/html/2606.27440#bib.bib34), Vannieuwenhovenet al\.,[2012](https://arxiv.org/html/2606.27440#bib.bib40)\], but convergence guarantees are limited to local optima\[Ishtevaet al\.,[2011](https://arxiv.org/html/2606.27440#bib.bib39)\]\.

## Appendix BExperimental Details

### B\.1PairSAE

Following the standard implementation of Boltz\-2, our sequence representation has dimensionns=384n\_\{s\}=384, while the pair representation has dimensionnz=128n\_\{z\}=128\. We compute theN−N\-mode SVD by simply flattening𝒵\\mathcal\{Z\}to its mode\-1 and mode\-2 unfolding, and obtain the SVD of these matrices usingnumpy\.linalg\.svd\. We truncate to the firstr=64r=64columns, and ifNtok<rN\_\{\\text\{tok\}\}<rwe fill the remaining entries with zeroes\. After concatenating sequence embeddings𝐬\\mathbf\{s\}and the SVD\-derived embedding𝐦\\mathbf\{m\}into a 512\-dimensional vector, we perform a layer normalization\.

We let PairSAE have a dictionary size ofD=16,384D=16,384, corresponding to an expansion factor of32×32\\times\. We train on minibatches of 2,048 tokens using the Adam optimizer\[Kingma and Ba,[2014](https://arxiv.org/html/2606.27440#bib.bib7)\], with learning rate 0\.0002\. We train for 250,000 steps on 40,000 systems sampled from the PLINDER training set\[Durairajet al\.,[2024](https://arxiv.org/html/2606.27440#bib.bib26)\]with at most 512 residues, and at each step we sample tokens at random\. Due to compute constraints, we do not use the MSA when training and evaluating the PairSAE\.

### B\.2Linear probing

We consider annotations from UniProt\[UniProt,[2025](https://arxiv.org/html/2606.27440#bib.bib6)\], and binarize them by one\-hot encoding each annotation\. We map each chain in a PLINDER system to its UniProt annotation using the SIFTS database\[Velankaret al\.,[2012](https://arxiv.org/html/2606.27440#bib.bib9)\], and match sequences with a minimum of 95% correspondance between the two databases, ignoring sequences that can’t be matched\.

We also consider a subset of the system\-level annotations found in PLINDER, as well as the PLIP interaction fingerprints\[Salentinet al\.,[2015](https://arxiv.org/html/2606.27440#bib.bib10)\]that are also found in PLINDER\. Iterating over featureshih\_\{i\}fori=1,…,Di=1,\\dots,Dand concept annotationsyjy\_\{j\}forj=1,…,Nyj=1,\\dots,N\_\{y\}, we consider a simple classifier

y^j=𝟏​\{hi\>τi​j\}\.\\hat\{y\}\_\{j\}=\\boldsymbol\{1\}\\\{h\_\{i\}\>\\tau\_\{ij\}\\\}\.\(9\)We pick the best thresholdτi​j\\tau\_\{ij\}on a training set of 7,680 complexes from the PLINDER training set, we select the best featureiifor each conceptjjbased onF1F\_\{1\}scores on a validation set of 2,560 complexes, and finally reportF1F\_\{1\}scores on a test set comprised of 5,120 randomly selected complexes\. We restrict our analysis to concepts that appear on at least five different complexes\.

### B\.3Hypothesis generation

We build a training set for LASSO regression by obtaining Boltz\-2 predictions for affinity values in 980 systems from PLINDER, held out from the PairSAE training set\. The test set is Posebusters\[Buttenschoenet al\.,[2024](https://arxiv.org/html/2606.27440#bib.bib12)\]\. In figure[B\.4](https://arxiv.org/html/2606.27440#A2.F4)we highlight how these differ in terms of the response variable, and note the lack of high affinity complexes in our training set\.

LASSO regression is fit usingsklearn\.linear\_model\.Lasso\[Pedregosaet al\.,[2011](https://arxiv.org/html/2606.27440#bib.bib8)\]\. Detailed results for both models are presented in Table[B\.2](https://arxiv.org/html/2606.27440#A2.T2)\. See Figure[B\.5](https://arxiv.org/html/2606.27440#A2.F5)for true vs predicted Boltz\-2 affinity values based on the PairSAE at R3\-L33\.

As mentioned in Section[B](https://arxiv.org/html/2606.27440#A2), we did not use the MSA for our analysis\. In Figure[B\.6](https://arxiv.org/html/2606.27440#A2.F6)we compare the Boltz\-2 output with and without MSA, and note a substantial difference in predicted affinity values and affinity probability\. As MSA\-based predictions are conditioned on more information, we expect these to be more accurate\. In future work, we aim at replicating our analysis using MSA inputs\.

Table B\.2:LASSO regression on max\-pooled PairSAE features\.ρ\\rhodenotes Spearman rank correlation\.![Refer to caption](https://arxiv.org/html/2606.27440v1/figures/lasso_train_test.png)Figure B\.4:Boltz\-2 affinity values for the training and test set used in the hypothesis generation task\.![Refer to caption](https://arxiv.org/html/2606.27440v1/figures/lasso_33.png)Figure B\.5:Negative affinity values from Boltz\-2 \(higher means more affine\), and predicted values from the LASSO regression in \([8](https://arxiv.org/html/2606.27440#S4.E8)\) using the PairSAE at R3\-L33\.![Refer to caption](https://arxiv.org/html/2606.27440v1/figures/msa_vs_no_msa_pb.png)Figure B\.6:Comparison of having MSA on and off in the Boltz\-2 affinity values and affinity probabilities, in the Posebusters dataset\.

## Appendix CAdditional results

In Figure[C\.7](https://arxiv.org/html/2606.27440#A3.F7)we display how lasso coefficients evolve for varying levels of regularization, and highlight features having strong association with the predicted affinity\. Among these, we highlight how feature 1744 and 1770 from R3\-L33 activate in two ligands in Figure[C\.8](https://arxiv.org/html/2606.27440#A3.F8)\.

In Section[4\.2](https://arxiv.org/html/2606.27440#S4.SS2)we highlighted feature 2299 from R3\-L64, that exhibits a strong group difference by activating on complexes with higher predicted affinity\. In R3\-L33 we found a feature representing the opposite, and activating only samples that on average have a lower predicted affinity\. Group differences for these two features are presented in Figure[C\.9](https://arxiv.org/html/2606.27440#A3.F9)\.

In Figure[C\.10](https://arxiv.org/html/2606.27440#A3.F10)we display values of feature 3888 overlaid on ligands where it is activated\. In these examples, the predicted complexes do not appear to present a plausible docked pose: the small molecule is displaced from the protein interface and does not form stable contacts\.

![Refer to caption](https://arxiv.org/html/2606.27440v1/figures/lasso_paths_32.png)

![Refer to caption](https://arxiv.org/html/2606.27440v1/figures/lasso_paths_64.png)

Figure C\.7:LASSO regression coefficients for varying values of theλ\\lambdaregularization parameter in \([8](https://arxiv.org/html/2606.27440#S4.E8)\), fit on the PairSAE at R3\-L33 \(left\) and R3\-L64 \(right\)\.![Refer to caption](https://arxiv.org/html/2606.27440v1/figures/1770.png)

![Refer to caption](https://arxiv.org/html/2606.27440v1/figures/1744.png)

Figure C\.8:Ligands activating on feature 1770 \(left\) and feature 1774 \(right\)\. These are the top 2 positive features that predict the Boltz\-2 affinity in R3\-L33, as highlighted in Figure[C\.7](https://arxiv.org/html/2606.27440#A3.F7)\.![Refer to caption](https://arxiv.org/html/2606.27440v1/figures/group_diff_2299.png)

![Refer to caption](https://arxiv.org/html/2606.27440v1/figures/nav_group_diff.png)

![Refer to caption](https://arxiv.org/html/2606.27440v1/figures/nav_group_diff_test.png)

Figure C\.9:Group difference in Boltz\-2 affinity values when splitting the dataset based on whether a feature activates or not\.Left:feature 2299 from R3\-L64, training set\.Right:feature 3888 from R3\-L33, training set\.Bottom:feature 3888 from R3\-L33, test set\.![Refer to caption](https://arxiv.org/html/2606.27440v1/figures/3888_1.png)

![Refer to caption](https://arxiv.org/html/2606.27440v1/figures/3888_2.png)

![Refer to caption](https://arxiv.org/html/2606.27440v1/figures/3888_4.png)

![Refer to caption](https://arxiv.org/html/2606.27440v1/figures/3888_3.png)

Figure C\.10:Examples where feature 3888 is activated\.

Similar Articles

Co-folding model guided by structural proteomics

arXiv cs.LG

Introduces AIMS-Fold, an inference-time guided-diffusion framework that integrates cross-linking mass spectrometry (XL-MS) and hydrogen-deuterium exchange (HDX-MS) data to improve protein co-folding predictions for induced proximity drug targets.

From Sparse Features to Trustworthy Proxies: Certifying SAE-Based Interpretability

arXiv cs.LG

This paper proposes a post-hoc certification framework for sparse autoencoder (SAE) based interpretability, deriving an upper bound on the frozen language model's risk using measurable quantities. The framework is validated on GPT-2 Small, Gemma-2B, and Llama-3-8B, showing non-vacuous bounds and revealing depth-dependent behavior.

Turn-Averaged SAEs for Feature Discovery and Long-Context Attribution

arXiv cs.CL

This paper introduces turn-averaged sparse autoencoders (SAEs) that operate on average activations across conversational turns, enabling efficient feature discovery and attribution graphs for long contexts. It also proposes a nested architecture for joint training with per-token features.

Sparse Autoencoders for Interpretable Out-of-Distribution Detection

arXiv cs.LG

This paper introduces a novel method using sparse autoencoders (SAEs) to learn interpretable features from intermediate network activations for out-of-distribution (OOD) detection, achieving state-of-the-art performance and providing insights into how distribution shifts affect learned representations.