Preference-based Antibody Expression Ranking: Scaling with Large-scale Weak Supervision
Summary
This paper proposes a preference-based learning framework for antibody expression ranking, integrating scarce quantitative data with large-scale weak positive supervision from immunization sequences. The method adapts Direct Preference Optimization to protein language models using a union-masked log-likelihood approximation and IMGT-based alignment, achieving improved ranking performance on a diverse internal dataset.
View Cached Full Text
Cached at: 07/21/26, 06:48 AM
# Scaling with Large-scale Weak Supervision
Source: [https://arxiv.org/html/2607.16263](https://arxiv.org/html/2607.16263)
## Preference\-based Antibody Expression Ranking: Scaling with Large\-scale Weak Supervision
###### Abstract
Antibody expression ranking is a critical task in antibody design, yet its modelling is severely hindered by the scarcity of labeled experimental data\. To address this, we propose a unified preference\-based learning framework that integrates scarce quantitative expression data with large\-scale weak positive supervision from immunization data\. We adapt Direct Preference Optimization \(DPO\) to protein language models by introducing a union\-masked log\-likelihood approximation and IMGT\-based alignment, enabling efficient training on variable\-length sequences\. Evaluating on a diverse internal dataset of 1254 labeled sequences and 4 million unlabeled camelid\-derived antibodies, we show that our method consistently outperforms baselines on most metrics\. Our results demonstrate that preference learning can effectively learn from weak supervision, providing a scalable solution for antibody expressibility optimization in data\-constrained settings\. Project page: https://kisoji\-biotechnology\-inc\.github\.io/Preference\-Expression\-Ranking/\.
AI for science, AI for drug discovery, protein language model\.
## 1Introduction
Antibody expression ranking is a critical bottleneck in antibody development, yet the biophysical principles governing why some sequences achieve high yields while others fail remain poorly understood\. This unpredictability motivates the use of computational models to guide candidate selection\. In practical antibody campaigns, enrichment and AI\-based generation routinely produce hundreds to thousands of candidates, so the relevant objective is not precise yield prediction for individual sequences, but a ranked prioritization of candidates with the highest likelihood of expressibility\. Realizing this objective, however, is challenging due to the extreme scarcity of labeled experimental data; public benchmarks typically contain only hundreds of sequences with limited diversity in antibody length and yield distributions, failing to reflect the complexity of real design \(Table[1](https://arxiv.org/html/2607.16263#S1.T1)\)\.
Simultaneously, industrial antibody discovery workflows generate additional sources of supervision that are largely absent from existing benchmarks\. Standard immunization\-based discovery pipelines routinely produce millions of unique antibody sequences derived from immune responses\(Hankeet al\.,[2020](https://arxiv.org/html/2607.16263#bib.bib14); Tsurutaet al\.,[2024](https://arxiv.org/html/2607.16263#bib.bib6)\)\. Although these sequences are not associated with quantitative expression measurements, empirical evidence from internal experimental studies suggests that more than 90% of camelid\-derived antibodies are expressible, whereas the expression rate drops to around 60% for antibodies derived from transgenic mice or generated by AI models\. This provides a form of weak positive supervision, an informative but underutilized signal that lies between supervised and unsupervised learning\.
In this work, we investigate the integration of scarce, quantitative expression measurements with large\-scale weak positive supervision for antibody optimization\. We utilize Exp\-1K, a proprietary dataset of 1254 sequences that, unlike existing benchmarks \(Table[1](https://arxiv.org/html/2607.16263#S1.T1)\), includes both expressible and non\-expressible examples with substantial length variability\. To complement this dataset, we leverage another private dataset, Camel\-4M, a collection of four million camelid\-derived antibody sequences obtained from immunization, which we treat as a source of weak positive supervision\.
DatasetNNYield Range \(mg/L\)Non\-expAvg\. Len\.Len\. Std\.SourceGarbinski et al\. \(2023\)94\[68\.3, 373\.3\]×\\times122\.62\.2\(Chungyounet al\.,[2024](https://arxiv.org/html/2607.16263#bib.bib1)\),originally unpublishedJain et al\. \(2017\)136\[6\.6, 277\.2\]×\\times119\.53\.2\(Jainet al\.,[2017](https://arxiv.org/html/2607.16263#bib.bib2)\)Szkodny et al\. \(2024\)178\[1\.5, 22\.8\]×\\times120\.00\.0\(Szkodny and Lee,[2024](https://arxiv.org/html/2607.16263#bib.bib5)\)PROPHET\-Ab246\[34\.3, 781\.9\]×\\times149\.00\.0\(Arsiwalaet al\.,[2025](https://arxiv.org/html/2607.16263#bib.bib3)\)Exp\-1K \(Ours\)1254\[0\.0, 435\.5\]✓\\checkmark123\.312\.8InternalCamel\-4M \(Ours\)4\.2M\-\-119\.74\.6InternalTable 1:Comparison of antibody expression benchmarks\. Our Exp\-1K dataset significantly exceeds existing benchmarks in both scale and coverage, notably including non\-expressible sequences and exhibiting higher sequence length variability\. Camel\-4M provides an additional large\-scale source of weak positive supervision derived from immunization\.The central challenge is to design a training framework that can naturally absorb these heterogeneous supervision signals\. We address this challenge through Direct Preference Optimization \(DPO\)\(Rafailovet al\.,[2023](https://arxiv.org/html/2607.16263#bib.bib8)\)\. DPO provides a unified abstraction for antibody expression ranking: quantitative expression yields can be converted into relative preferences between sequences, while weak positive supervision induces implicit preferences over the sequence space without requiring explicit labels\. Under this formulation, a single model can jointly learn to rank sequences by relative yield score and discriminate expressible from non\-expressible antibodies\.
Our framework begins with a domain\-alignment phase, where the model undergoes continual pretraining with masked language modelling on immunization\-derived Camel\-4M to adapt existing protein language models \(pLMs\) to specific protein families\(Alleyet al\.,[2019](https://arxiv.org/html/2607.16263#bib.bib12); Biswaset al\.,[2021](https://arxiv.org/html/2607.16263#bib.bib13)\)\. This establishes a specialized representation of the antibody space, providing a robust foundation for preference learning\. Building on this initialization, we introduce a modified DPO objective to pLMs\. Applying DPO to pLMs is non\-trivial, as pLMs are mostly masked language models and the sequence likelihood is defined differently from causal language models\. Specifically, pLMs rely on pseudo\-log\-likelihood\(Salazaret al\.,[2020](https://arxiv.org/html/2607.16263#bib.bib11); Meieret al\.,[2021](https://arxiv.org/html/2607.16263#bib.bib9)\), which is computationally expensive when applied independently at each sequence position\. We address this limitation by adopting a union\-masked log\-likelihood approximation\(Ferraguet al\.,[2025](https://arxiv.org/html/2607.16263#bib.bib7)\), where differing positions between two sequences are jointly masked and evaluated in a single forward pass\. To enable consistent construction of such preference pairs for sequences with variable lengths, we align antibody sequences with IMGT numbering\(Lefrancet al\.,[2003a](https://arxiv.org/html/2607.16263#bib.bib37),[b](https://arxiv.org/html/2607.16263#bib.bib38)\)using ANARCI software\(Dunbar and Deane,[2016](https://arxiv.org/html/2607.16263#bib.bib39)\)\. This design allows efficient and scalable preference optimization for antibody sequences under realistic length variability\.
Empirically, we evaluate our framework on the Exp\-1K dataset, demonstrating that preference\-based learning effectively captures antibody expressibility\. Our results show that our model consistently outperforms standard regression and classification baselines in both yield ranking and binary prediction\. Specifically, we observe that the combined effect of domain\-aligned MLM and preference optimization leads to substantial performance gains compared to training on labeled data alone\. Furthermore, we show that our framework scales effectively with increasing amounts of weak supervision, confirming that it can successfully extract useful biophysical signals from large\-scale, unlabeled antibody sequences\. These findings suggest that preference\-based optimization is a practical and robust approach for antibody design in data\-constrained scenarios\.
We summarize our primary contributions as follows: \(a\) Unified Framework: We develop a preference\-based learning framework that integrates scarce quantitative yields and large\-scale weak positive supervision within a single training objective\. \(b\) Efficient DPO for pLMs: We adapt a union\-masked log\-likelihood approximation and the IMGT\-based alignment, enabling efficient training on variable\-length sequences\. \(c\) Empirical Validation: We demonstrate that our approach consistently outperforms baselines with multiple protein language backbones\. \(d\) Scaling Analysis: We show that our framework effectively scales with increasing volumes of weak supervision, proving its ability to leverage unlabeled industrial data to mitigate labeled data scarcity\.
Conflict of Interest Disclosure
The author Josh Qixuan Sun conducted this work during an internship at Kisoji Biotech and received financial support from Kisoji Biotech\. This project was also supported in part by Mitacs\. The authors declare no other financial conflicts of interest that could reasonably be perceived to influence the work\.
## 2Related Work
##### Preference learning for protein language models\.
For general protein design, ProteinDPO\(Xuet al\.,[2025](https://arxiv.org/html/2607.16263#bib.bib20)\)integrates structural feedback from folding simulators \(e\.g\., AlphaFold\) to optimize inverse folding models, significantly improving TM\-scores on CATH benchmarks\. ResiDPO\(Xueet al\.,[2025](https://arxiv.org/html/2607.16263#bib.bib21)\)further refines this by introducing residue\-level preference signals based on local confidence scores \(pLDDT\), enhancing the success rate of complex enzyme design\. In the domain of antibody engineering, AbDPO\(Zhouet al\.,[2024](https://arxiv.org/html/2607.16263#bib.bib22)\)and AlignAb\(Wenet al\.,[2024](https://arxiv.org/html/2607.16263#bib.bib23)\)utilize physics\-based energy functions to align diffusion models toward Pareto\-optimal binders, balancing binding affinity with structural stability\. To address the conservatism of KL\-regularized DPO, SimBinder\-IF\(Zhaoet al\.,[2025](https://arxiv.org/html/2607.16263#bib.bib24)\)employs SimPO, a reference\-free objective with a calibrated margin, achieving state\-of\-the\-art Spearman correlation in zero\-shot antibody\-antigen affinity prediction\. Beyond proteins, the paradigm has extended to RNA design through RiboPO\(Sunet al\.,[2025](https://arxiv.org/html/2607.16263#bib.bib25)\), which aligns sequences with thermodynamic stability and geometric fidelity\. To address the computational bottlenecks, g\-DPO\(Ferraguet al\.,[2025](https://arxiv.org/html/2607.16263#bib.bib7)\)implements a scalable framework that prunes redundant sequence pairs, achieving significant training acceleration across protein fitness and stability tasks\. Collectively, these works demonstrate that preference\-based fine\-tuning effectively bridges the gap between evolutionary probability distributions and specific biophysical engineering goals\.
##### Protein developability prediction\.
Computational approaches for developability have evolved from heuristic\-based servers like Protein\-Sol\(Hebditchet al\.,[2017](https://arxiv.org/html/2607.16263#bib.bib26)\)to sophisticated deep learning architectures\. Geometric approaches, such as GVP\-GNN\(Kimet al\.,[2023](https://arxiv.org/html/2607.16263#bib.bib27)\)and D\-GNN\(Khadeet al\.,[2023](https://arxiv.org/html/2607.16263#bib.bib28)\), utilize equivariant graph neural networks to model the 3D spatial relationships and surface patches of residues, which are critical for predicting thermostability \(TmT\_\{m\}\) and solubility\. However, the current state\-of\-the\-art is increasingly dominated by finetuning pLMs\. Large\-scale models such as ProtBERT\(Brandeset al\.,[2022](https://arxiv.org/html/2607.16263#bib.bib32)\)leverage masked language modelling on hundreds of millions of sequences to capture structural and functional constraints without explicit supervision\. Recent research highlights the efficacy of task\-specific adaptation: ESM\-1v\(Meieret al\.,[2021](https://arxiv.org/html/2607.16263#bib.bib9)\)demonstrates high accuracy in predicting antibody non\-specificity through logistic regression on its embeddings\. NetSolP\(Thumuluriet al\.,[2022](https://arxiv.org/html/2607.16263#bib.bib35)\)finetunes embeddings using a Transformer architecture to predict protein solubility and purification suitability\. DeepSTABp\(Junget al\.,[2023](https://arxiv.org/html/2607.16263#bib.bib33)\)integrates PLM representations with experimental metadata to regress continuousTmT\_\{m\}values\. RP3Net\(Tankhilevichet al\.,[2026](https://arxiv.org/html/2607.16263#bib.bib42)\)takes the presentation from protein language model and trains a linear layer to predict whether the given antibody sequence is expressible or not\. Overall, most existing pLM\-based developability predictors follow a paradigm, where a pretrained pLM is frozen and its final\-layer embeddings are used as features for downstream regression or classification heads\. While effective, this strategy does not explicitly modify the sequence\-level likelihood of the pLM\. In contrast, preference learning methods, as discussed above, directly finetune the generative model itself using pairwise or ranked supervision, enabling the model to align its sequence probabilities with developability\-related objectives\.
## 3Method
### 3\.1Problem Setup
We consider antibody expression ranking as a sequence\-level scoring problem\. An antibody sequence is represented asy=\(s1,s2,…,sL\)y=\(s\_\{1\},s\_\{2\},\\ldots,s\_\{L\}\), where each amino acid tokensis\_\{i\}belongs to the standard set of 20 residues\. We focus on variable\-length VHH antibody sequences and make no assumption of fixed length\.
Given a sequenceyy, the objective is to assign a scalar score that reflects its expression behavior, such that non\-expressible sequences are ranked lower than expressible ones, and among expressible antibodies, sequences with higher expression yields are ranked higher\. This formulation naturally supports both ranking\-based evaluation and binary expressibility prediction without requiring separate task\-specific models\.
### 3\.2Sequence Scoring via Pseudo\-Log\-Likelihood
We derive sequence\-level scores from protein language models \(pLMs\) using pseudo\-log\-likelihood \(PLL\)\(Salazaret al\.,[2020](https://arxiv.org/html/2607.16263#bib.bib11); Meieret al\.,[2021](https://arxiv.org/html/2607.16263#bib.bib9)\)\. For a sequencey=\(s1,…,sL\)y=\(s\_\{1\},\\ldots,s\_\{L\}\), the PLL score is defined as
PLL\(y\)=1L∑i=1Llogpθ\(si∣y∖i\),\\mathrm\{PLL\}\(y\)=\\frac\{1\}\{L\}\\sum\_\{i=1\}^\{L\}\\log p\_\{\\theta\}\(s\_\{i\}\\mid y\_\{\\setminus i\}\),wherey∖iy\_\{\\setminus i\}denotes the sequence with theii\-th position masked\. We use the average PLL as a length\-normalized scalar score for ranking antibody sequences\.
This scoring function induces a total ordering over sequences and can be directly applied in a zero\-shot manner\. To obtain a binary expressibility prediction from PLL scores, we adopt a percentile\-based thresholding scheme\. Specifically, given a training set with a known fraction of expressible sequences, we compute PLL scores for all training samples and select the corresponding percentile as a global decision threshold\. This threshold is then applied consistently to validation and test sets, yielding a binary classifier without introducing additional trainable parameters\.
### 3\.3Preference Optimization with Protein Language Models
While pseudo\-log\-likelihood provides a meaningful sequence\-level score, it does not directly incorporate available supervision\. We therefore adopt a preference\-based learning framework to adapt protein language models using heterogeneous supervision signals\.
#### 3\.3\.1Direct Preference Optimization
Preference learning formulates antibody expression ranking through pairwise comparisons, and we adapt the Direct Preference Optimization \(DPO\)\(Rafailovet al\.,[2023](https://arxiv.org/html/2607.16263#bib.bib8)\)framework\. DPO optimizes a policy modelπθ\\pi\_\{\\theta\}to better satisfy preferences compared to a reference modelπref\\pi\_\{\\text\{ref \}\}\(typically the initial pretrained model\), using a dataset𝒟\\mathcal\{D\}of preference pairs \(yw,yly\_\{w\},y\_\{l\}\) for a given promptxx, whereywy\_\{w\}is preferred overyly\_\{l\}\. The DPO objective maximizes the likelihood of preferred responses while regularizing against large deviations from the reference model via a KL divergence penalty:
ℒDPO\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{DPO\}\}\(πθ;πref\)=−𝔼\(x,yw,yl\)∼𝒟\[\\displaystyle\\left\(\\pi\_\{\\theta\};\\pi\_\{\\mathrm\{ref\}\}\\right\)=\-\\mathbb\{E\}\_\{\(x,y\_\{w\},y\_\{l\}\)\\sim\\mathcal\{D\}\}\\Bigl\[logσ\(βlogπθ\(yw∣x\)πref\(yw∣x\)−βlogπθ\(yl∣x\)πref\(yl∣x\)\)\],\\displaystyle\\log\\sigma\\Bigl\(\\beta\\log\\frac\{\\pi\_\{\\theta\}\(y\_\{w\}\\mid x\)\}\{\\pi\_\{\\mathrm\{ref\}\}\(y\_\{w\}\\mid x\)\}\{\}\-\\beta\\log\\frac\{\\pi\_\{\\theta\}\(y\_\{l\}\\mid x\)\}\{\\pi\_\{\\mathrm\{ref\}\}\(y\_\{l\}\\mid x\)\}\\Bigr\)\\Bigr\],
whereσ\\sigmais the sigmoid function, andβ\\betais a hyperparameter controlling the strength of the preference relative to the regularization\. In our context,xxis none,ywy\_\{w\}andyly\_\{l\}are candidate antibody sequences,πref\\pi\_\{ref\}is a pretrained pLM, andπθ\\pi\_\{\\theta\}is the model being finetuned\. This formulation allows both strong and weak supervision to be expressed uniformly as relative preferences, without introducing explicit reward models or task\-specific heads\.
#### 3\.3\.2Likelihood Approximation for Protein Language Models
The DPO objective relies on evaluating log\-probabilities of sequences under both the policy modelπθ\\pi\_\{\\theta\}and the reference modelπref\\pi\_\{\\mathrm\{ref\}\}\. However, protein language models are typically trained with masked language modelling objectives and do not define tractable autoregressive likelihoods of the formlogπ\(y∣x\)\\log\\pi\(y\\mid x\)\. To address this, we calculate sequence log\-likelihood induced by the policyπ\\piusing pseudo\-log\-likelihood \(PLL\)\(Salazaret al\.,[2020](https://arxiv.org/html/2607.16263#bib.bib11); Meieret al\.,[2021](https://arxiv.org/html/2607.16263#bib.bib9)\):
logπ\(y∣x\)=PLLπ\(y\)=1L∑i=1Llogπ\(si∣y∖i\),\\log\\pi\(y\\mid x\)=\\mathrm\{PLL\}\_\{\\pi\}\(y\)=\\frac\{1\}\{L\}\\sum\_\{i=1\}^\{L\}\\log\\pi\(s\_\{i\}\\mid y\_\{\\setminus i\}\),wherey∖iy\_\{\\setminus i\}denotes the sequence with theii\-th position masked\.
Under this calculation, the log\-probability ratios in the DPO objective are replaced by differences in PLL scores\. However, computingPLLπ\(y\)\\mathrm\{PLL\}\_\{\\pi\}\(y\)exactly requires one forward pass per sequence position, which is computationally expensive during preference optimization\. Crucially, DPO only depends on relative scores between preferred and less preferred sequences\. For a preference pair\(yw,yl\)\(y\_\{w\},y\_\{l\}\), we define the union mask:
ℳ\(yw,yl\)=\{i∣\(yw\)i≠\(yl\)i\}\.\\mathcal\{M\}\(y\_\{w\},y\_\{l\}\)=\\\{\\,i\\mid\(y\_\{w\}\)\_\{i\}\\neq\(y\_\{l\}\)\_\{i\}\\,\\\}\.To make the definition ofℳ\(yw,yl\)\\mathcal\{M\}\(y\_\{w\},y\_\{l\}\)well\-defined for antibody sequences with variable lengths, we first align sequences using a standardized numbering scheme\. Specifically, we map each antibody sequence to IMGT numbering, which assigns homologous positions across variable\-length variable regions\. All sequence comparisons and masking operations are then performed in the aligned IMGT coordinate space\.
We perform a single forward pass on the partially masked sequence\. The log\-score ofywy\_\{w\}under policyπ\\piis then approximated as\(Zhaoet al\.,[2024](https://arxiv.org/html/2607.16263#bib.bib15);[Hawkins\-Hookeret al\.,](https://arxiv.org/html/2607.16263#bib.bib16); Ferraguet al\.,[2025](https://arxiv.org/html/2607.16263#bib.bib7)\):
PLLπ\(yw\)≈1\|ℳ\|∑i∈ℳlogπ\(\(yw\)i∣y∖ℳ\),\\mathrm\{PLL\}\_\{\\pi\}\(y\_\{w\}\)\\;\\approx\\;\\frac\{1\}\{\|\\mathcal\{M\}\|\}\\sum\_\{i\\in\\mathcal\{M\}\}\\log\\pi\\bigl\(\(y\_\{w\}\)\_\{i\}\\mid y\_\{\\setminus\\mathcal\{M\}\}\\bigr\),with an analogous expression foryly\_\{l\}\. Positions outsideℳ\\mathcal\{M\}need not be evaluated, as their contributions cancel in the score difference used by DPO\.
Figure 1:Construction of preference pairs\. Strong supervision \(top\) uses yield differences with marginδ\\delta; weak supervision \(bottom\) pairs immunization\-derived sequences with zero\-yield antibodies\. Both contribute to the unified preference dataset𝒟pair\\mathcal\{D\}\_\{\\text\{pair\}\}\.
#### 3\.3\.3Preference Pair Construction
We construct preference pairs from two supervision sources through a unified pipeline, illustrated in Fig\.[1](https://arxiv.org/html/2607.16263#S3.F1)\. Both strong and weak supervision are converted into pairwise comparisons and merged into a single preference dataset\.
##### Strong supervision\.
Let𝒟sup=\{\(yi,ri\)\}i=1N\\mathcal\{D\}\_\{\\text\{sup\}\}=\\\{\(y\_\{i\},r\_\{i\}\)\\\}\_\{i=1\}^\{N\}denote the set of antibody sequences with measured expression yieldsri∈ℝ≥0r\_\{i\}\\in\\mathbb\{R\}\_\{\\geq 0\}\. From this set, we construct preference pairs by comparing yield values\. Specifically, for any pair\(yi,yj\)\(y\_\{i\},y\_\{j\}\), we define a preference wheneverri−rj≥δr\_\{i\}\-r\_\{j\}\\geq\\deltawhereδ\>0\\delta\>0is a predefined yield margin\. In this case,yiy\_\{i\}is treated as the preferred sequenceywy\_\{w\}andyjy\_\{j\}as the less preferred sequenceyly\_\{l\}\. Pairs that do not satisfy this margin are discarded to reduce ambiguity due to experimental noise\. In this work, we setδ=30\\delta=30\.
##### Weak supervision\.
Weak supervision is provided by a large collection of immunization\-derived antibody sequences, denoted as𝒟weak=\{yk\}k=1M\\mathcal\{D\}\_\{\\text\{weak\}\}=\\\{y\_\{k\}\\\}\_\{k=1\}^\{M\}\. These sequences lack quantitative expression measurements and are treated as weak positive supervision\.
To induce preferences without assigning explicit negative labels, we pair weakly supervised sequences against strongly supervised sequences with zero measured expression yield\. Let𝒟0=\{yi∣\(yi,ri\)∈𝒟sup,ri=0\}\\mathcal\{D\}\_\{0\}=\\\{\\,y\_\{i\}\\mid\(y\_\{i\},r\_\{i\}\)\\in\\mathcal\{D\}\_\{\\text\{sup\}\},\\ r\_\{i\}=0\\,\\\}denote the subset of non\-expressible sequences\. For eachyk∈𝒟weaky\_\{k\}\\in\\mathcal\{D\}\_\{\\text\{weak\}\}and a sampledy0∈𝒟0y\_\{0\}\\in\\mathcal\{D\}\_\{0\}, we construct a preference pair\(yw,yl\)=\(yk,y0\)\(y\_\{w\},y\_\{l\}\)=\(y\_\{k\},y\_\{0\}\)\.
##### Unified preference dataset\.
The final preference dataset is defined as𝒟pair=𝒟suppair∪𝒟weakpair\\mathcal\{D\}\_\{\\text\{pair\}\}=\\mathcal\{D\}\_\{\\text\{sup\}\}^\{\\text\{pair\}\}\\;\\cup\\;\\mathcal\{D\}\_\{\\text\{weak\}\}^\{\\text\{pair\}\}, where𝒟suppair\\mathcal\{D\}\_\{\\text\{sup\}\}^\{\\text\{pair\}\}and𝒟weakpair\\mathcal\{D\}\_\{\\text\{weak\}\}^\{\\text\{pair\}\}denote preference pairs constructed from strong and weak supervision, respectively\. Both supervision sources are thus integrated into a single preference learning framework\. In practice, the number of pairs derived from these two sources is highly imbalanced, with weak supervision pairs vastly outnumbering the supervised ones\. To ensure that the model effectively learns from the scarce but high\-quality quantitative data, we implement a sampling rate hyperparameterα\\alphato oversample𝒟suppair\\mathcal\{D\}\_\{\\text\{sup\}\}^\{\\text\{pair\}\}relative to𝒟weakpair\\mathcal\{D\}\_\{\\text\{weak\}\}^\{\\text\{pair\}\}\. This ensures a balanced gradient signal from both heterogeneous sources throughout the optimization process\. The impact of different sampling rates will be further discussed in our ablation study\.
### 3\.4Training Procedure
Training proceeds in two stages: domain\-aligned continual pretraining with masked language modelling, followed by preference optimization\.
##### Stage I: Continual pretraining\.
We first perform continual pretraining of the pretrained protein language model using a standard masked language modelling objective on the weakly supervised dataset𝒟weak\\mathcal\{D\}\_\{\\text\{weak\}\}\. This stage adapts the model to the antibody\-specific sequence distribution and is sometimes referred to as*evo\-tuning*\(Alleyet al\.,[2019](https://arxiv.org/html/2607.16263#bib.bib12); Biswaset al\.,[2021](https://arxiv.org/html/2607.16263#bib.bib13); Hsuet al\.,[2022](https://arxiv.org/html/2607.16263#bib.bib17);[Gordonet al\.,](https://arxiv.org/html/2607.16263#bib.bib18)\)\. No expression\-related supervision is used at this stage\.
##### Stage II: Preference finetuning\.
Starting from the domain\-aligned initialization, we apply DPO using the preference dataset𝒟pair\\mathcal\{D\}\_\{\\text\{pair\}\}\. Model parameters are optimized to satisfy relative preferences between antibody sequences under the DPO objective, integrating both strong quantitative supervision and weak positive supervision\.
## 4Experiments
### 4\.1Datasets
To evaluate our preference\-based learning framework under realistic antibody discovery conditions, we curate two internal datasets:Exp\-1Kfor supervised evaluation andCamel\-4Mfor large\-scale weak supervision\. The statistical comparison with existing public benchmarks is summarized in Table[1](https://arxiv.org/html/2607.16263#S1.T1)\.
Figure 2:Histogram of Exp\-1K yield score\.
Figure 3:Sequence length density across different datasets\.
##### Exp\-1K: Labeled Expression Benchmark\.
This dataset comprises 1254 unique VHH sequences derived from internal antibody designs\. Unlike public datasets that primarily consist of expressible sequences, Exp\-1K explicitly includes approximately 27% \(335 sequences\) of non\-expressible candidates \(no detectable expression band, expressibility Yield \(mg/L\)<1<1mg/L\)\. As illustrated in Figure[3](https://arxiv.org/html/2607.16263#S4.F3), the yield distribution is highly skewed with a distinct zero\-yield peak\. This sparsity and imbalance make Exp\-1K a more challenging and realistic benchmark than existing small\-scale datasets\.
##### Camel\-4M: Large\-scale Weak Supervision\.
To provide a broad representation of the antibody sequence space, we utilize Camel\-4M, a corpus of 4\.2 million non\-redundant VHH sequences\. These sequences were derived from PBMC samples of immunized camelid species, including alpacas and llamas\. Figure[3](https://arxiv.org/html/2607.16263#S4.F3)shows the sequence length density across different datasets\. We treat the sequences from the immune maturation process as weak positive supervision\.
### 4\.2Experimental Setup
BinaryRankingCompositeBackboneTypeMethodStage IAUC↑\\uparrowMCC↑\\uparrowTau↑\\uparrowSCC↑\\uparrowRecall↑\\uparrowAvg↑\\uparrowAntiBERTa2\(Bartonet al\.,[2024](https://arxiv.org/html/2607.16263#bib.bib29)\)aLMZero\-shot×\\times52\.66\.8\-0\.2\-4\.111\.5\-Zero\-shot✓51\.413\.2\-0\.1\-0\.10\.1\-Supervised×\\times74\.529\.137\.642\.242\.333\.1Supervised✓76\.221\.838\.442\.442\.329\.0Ours \(Stage II\)×\\times77\.226\.733\.832\.838\.530\.8Ours \(Stage I\+II\)✓82\.443\.746\.350\.846\.245\.0IgBERT\(Kenlayet al\.,[2024](https://arxiv.org/html/2607.16263#bib.bib30)\)aLMZero\-shot×\\times50\.4\-12\.6\-0\.2\-0\.915\.4\-Zero\-shot✓56\.0\-1\.59\.09\.727\.0\-Supervised×\\times78\.67\.540\.043\.534\.617\.2Supervised✓75\.736\.838\.345\.846\.237\.5Ours \(Stage II\)×\\times73\.017\.134\.341\.546\.224\.2Ours \(Stage I\+II\)✓74\.837\.542\.954\.761\.540\.1ProtBERT\(Brandeset al\.,[2022](https://arxiv.org/html/2607.16263#bib.bib32)\)pLMZero\-shot×\\times43\.6\-3\.7\-8\.8\-11\.011\.5\-Zero\-shot✓54\.117\.44\.02\.826\.98\.4Supervised×\\times77\.816\.737\.340\.434\.624\.9Supervised✓76\.426\.137\.142\.138\.531\.1Ours \(Stage II\)×\\times74\.528\.629\.827\.319\.229\.2Ours \(Stage I\+II\)✓78\.040\.140\.544\.553\.840\.3Table 2:Main results on the Exp\-1K test set across two antibody language model \(aLM\) and one protein language model \(pLM\) backbone\. AUC, MCC, and Recall evaluate binary expressibility prediction, while Kendall’s Tau \(Tau\) and Spearman’s correlation coefficient \(SCC\) evaluate yield ranking\. Avg denotes the geometric mean of MCC and Tau\. Recall is reported at top 20% according to the predicted sequence scores\.Model Backbones and Initialization\. We evaluate our framework across three representative backbones: AntiBERTa2, IgBERT, and ProtBERT\. All models utilize their default architectures without structural modifications\. For our DPO\-based approach, the reference modelπref\\pi\_\{\\text\{ref\}\}is initialized from the weights after Stage I \(continual pretraining\) rather than the vanilla pLMs, ensuring a consistent starting point for preference alignment\.
Training Details\. Our training pipeline consists of two stages\. Stage I: Continual Pretraining \(CPT\)\. We perform CPT on the Camel\-4M dataset using the standard masked language modelling \(MLM\) objective with a 15% mask ratio\. Training is conducted for 1 epoch with a learning rate of1×10−51\\times 10^\{\-5\}, a batch size of 128, and a maximum sequence length of 168\. Stage II: Preference Fine\-tuning\. We use the AdamW optimizer with a learning rate of1×10−41\\times 10^\{\-4\}, 1K warmup steps, and a weight decay of 0\.01\. The KL\-divergence regularization parameterβ\\betais set to 0\.1, and the yield marginδ\\deltafor constructing preference pairs is fixed at 30\. For the main results, we use a sampling rateα=1000\\alpha=1000to balance the weak and strong supervision\.
Baselines and Fair Comparison\. To demonstrate the efficacy of preference\-based modelling, we implement a standard supervised baseline optimized with Mean Squared Error \(MSE\) loss\. It utilizes the same pLM backbones followed by an attention pooling layer and a linear head\. For a rigorous comparison, we report results using vanilla pLM weights and Stage I\-tuned weights\.
Evaluation Metrics\. Following established benchmarks such as ProteinGym and FLAb, we evaluate model performance across two dimensions: sequence ranking and binary classification\. \(1\) Ranking Metrics: To evaluate the correlation between predicted scores and experimental yields, we report Kendall’sτ\\tauand Spearman’s Correlation Coefficient \(SCC\)\. While SCC is a standard metric in protein fitness landscapes, it is sensitive to ”ties” \(identical ranks\)\. Given that approximately 27% of the Exp\-1K dataset consists of non\-expressible sequences \(yield=0=0\), the resulting heavy ties in ground\-truth labels make SCC less robust\. Therefore, we adopt Kendall’sτ\\tauas our primary ranking metric due to its superior handling of tied observations\. To provide a more granular view, SCC is calculated specifically on the subset of expressible sequences \(yield\>0\>0\)\. \(2\) Binary Metrics: We use the Matthews Correlation Coefficient \(MCC\) to assess the model’s ability to distinguish between expressible and non\-expressible antibodies\. \(3\) Composite Metric \(Avg\): We introduce a unified score calculated as the geometric mean of MCC and Kendall’sτ\\tau\. We prioritize the geometric mean over the arithmetic mean to penalize models that achieve high performance in one dimension at the significant expense of the other, thereby ensuring a balanced optimization for real\-world antibody screening\.
Data Split and Implementation\. The Exp\-1K dataset is partitioned into training, validation, and test sets with an 80/10/10 ratio\. Model selection is performed based on validation performance, and final results are reported on the unseen test set\. Experiments are conducted on NVIDIA H100 GPUs using bf16 precision, with batch sizes adjusted according to the backbone scale\.
### 4\.3Main Results
Table[2](https://arxiv.org/html/2607.16263#S4.T2)reports the main results on the Exp\-1K test set across two antibody language model backbones \(aLM\) and one protein language model \(pLM\) backbone\. We evaluate both binary expressibility prediction and yield ranking to reflect the dual nature of the task\. Across all evaluated metrics, our two\-stage method consistently achieves the strongest overall performance\. For all three backbones, the proposed approach outperforms baselines mostly across all metrics\. This demonstrates that preference optimization jointly improves ranking and classification performance\.
Limitations of zero\-shot baselines\.We first observe that zero\-shot baselines perform poorly across all metrics, regardless of whether continual pretraining is applied\. In particular, Kendall’s Tau and MCC are close to zero or even negative\. This suggests that sequence likelihood under a pretrained language model is weak for expressibility\. This behavior is expected, as expressibility differs fundamentally from properties such as thermostability or evolutionary fitness\. Many antibody sequences in the evaluation set are synthetically designed and do not correspond to naturally occurring sequences\. As a result, their probability under the natural sequence distribution modeled by a protein language model is not strongly correlated with downstream expression outcomes\. Continual pretraining alone does not resolve this mismatch, since it optimizes the same likelihood\-based objective\.
Limited gains from Stage I with supervised heads\.Next, we compare supervised learning baselines with and without Stage I continual pretraining\. Applying continual pretraining yields marginal or inconsistent improvements, and in some cases leads to performance degradation\. This trend is visible across multiple metrics\. This result indicates that representations optimized via continual pretraining are not necessarily aligned with the features required by downstream supervised objectives\. Simply initializing a supervised head from a continually pretrained backbone does not guarantee better expressibility prediction, especially when labeled data remain limited\.
Preference learning vs\. Stage I \+ supervised learning\.A key comparison in Table[2](https://arxiv.org/html/2607.16263#S4.T2)is between Stage I CPT followed by supervised learning and our preference\-based post\-training\. This comparison directly addresses the question of whether continual pretraining combined with standard regression is sufficient for expressibility prediction\. Despite using the same continually pretrained backbone, our method substantially outperforms supervised baselines across all reported metrics\. In particular, gains in Kendall’s Tau and MCC indicate improved alignment with both ranking and classification objectives\. This highlights the importance of preference\-based optimization, which directly matches the relative supervision signal available from weakly labeled data and avoids objective mismatch\.
Synergy between Stage I and Stage II\.Comparing ”Ours \(Stage II\)” with ”Ours \(Stage I\+II\)” reveals that Stage I continual pretraining is essential for maximizing performance\. For AntiBERTa2, the full pipeline boosts MCC from 26\.7 to 43\.7 and Tau from 33\.8 to 46\.3\. This demonstrates that Stage I creates a foundational representation space that is uniquely synergistic with Stage II preference optimization, significantly outperforming the use of Stage II alone\.
### 4\.4Scaling with Weak Supervision
Figure 4:Scaling with weak supervision\. Results are shown with and without Stage I continual pretraining \(CPT\), while keeping a consistent exposure ratio between supervised and weak pairs\.To investigate how our framework benefits from increasing scales of weak supervision, we evaluate the model performance by varying the size of𝒟weak\\mathcal\{D\}\_\{\\text\{weak\}\}from 1K to 4\.2M\. We conduct this analysis with AntiBERTa2 backbone under two settings: \(i\) w/ Stage I continual pretraining \(CPT\) on Camel\-4M, and \(ii\) w/o Stage I, where the preference learning starts directly from the vanilla pLM\. To ensure a fair comparison, the sampling rateα\\alphais dynamically adjusted to maintain a consistent exposure ratio between supervised pairs and weak pairs\.
As shown in Figure[4](https://arxiv.org/html/2607.16263#S4.F4), the average metric \(Avg\) generally increases with the amount of weak supervision for both backbones\. For the w/o Stage I setting, the curve exhibits noticeable fluctuations as the data scale grows\. In contrast, when Stage I pretraining is applied, the improvement with additional weak supervision is more stable and sustained, indicating that the proposed preference learning setup can continue to benefit from larger amounts of unlabeled immunization data\.
### 4\.5Sensitivity to Data Composition
The balance between high\-fidelity supervised pairs𝒟sup\\mathcal\{D\}\_\{\\text\{sup\}\}and large\-scale weak preference pairs𝒟weak\\mathcal\{D\}\_\{\\text\{weak\}\}is crucial for effective preference learning\. To manage the vast size disparity between these sets, we introduce a sampling hyperparameterα\\alphato modulate the exposure frequency of supervised data\. Formally, a weak preference pair is selected with probabilitypweak=Nweak/\(Nweak\+α⋅Nsup\)p\_\{\\text\{weak\}\}=N\_\{\\text\{weak\}\}/\(N\_\{\\text\{weak\}\}\+\\alpha\\cdot N\_\{\\text\{sup\}\}\), whereNweakN\_\{\\text\{weak\}\}andNsupN\_\{\\text\{sup\}\}denote the total number of available pairs in each set respectively\. Detailed calculations are in Appendix[A\.5](https://arxiv.org/html/2607.16263#A1.SS5)\.
α\\alphapweakp\_\{\\text\{weak\}\}MCCTauSCC1099\.7%13\.912\.28\.710096\.7%21\.714\.20\.0100074\.6%40\.140\.544\.5500037\.1%46\.938\.037\.41000022\.7%31\.130\.726\.8Table 3:Sensitivity analysis of the sampling hyperparameterα\\alpha\.pweakp\_\{\\text\{weak\}\}represents the theoretical probability of sampling a weak preference pair in each training step\.As shown in Table[3](https://arxiv.org/html/2607.16263#S4.T3), we evaluate the model performance across differentα\\alphavalues\. Contrary to the observation in standard multi\-task learning, the performance exhibits a clear sensitivity to the sampling ratio\. We find that a default value ofα=1000\\alpha=1000\(corresponding topweak≈74%p\_\{\\text\{weak\}\}\\approx 74\\%\) provides the optimal trade\-off\. Lower values ofα\\alphalead to the ”swamping” of the high\-quality supervised signal by the massive weak data, while excessively high values cause the model to overfit the small\-scale Exp\-1K dataset, sacrificing the general structural priors offered by the immunization data\.
### 4\.6Regularization and Stability
A fundamental question arises: does weak preference learning act as a regularizer? We investigate this by examining the optimization dynamics and generalization gaps across different learning paradigms\.
Generalization Gap\.As summarized in Table[4](https://arxiv.org/html/2607.16263#S4.T4), we report the discrepancyΔ=\|Avgval−Avgtest\|\\Delta=\|\\text\{Avg\}\_\{\\text\{val\}\}\-\\text\{Avg\}\_\{\\text\{test\}\}\|for the average metric\. Our full framework exhibits the tightest gap, whereas removing weak data or switching to a standard supervised regression objective significantly widens this gap\. This suggests that the large\-scale𝒟weak\\mathcal\{D\}\_\{\\text\{weak\}\}anchors the model within a biologically viable sequence space, preventing it from fitting experimental noise specific to the 1K samples\.
MethodΔ\\DeltaAvgOurs \(Full Framework\)5\.5w/o Weak Data14\.1CPT \+ Supervised Regression16\.3Table 4:Generalization gapΔ=\|Avgval−Avgtest\|\\Delta=\|\\text\{Avg\}\_\{\\text\{val\}\}\-\\text\{Avg\}\_\{\\text\{test\}\}\|across different settings\. LowerΔ\\Deltaindicates better generalization\.Figure 5:Validation performance trajectories over training for different methods\.Optimization Stability\.Figure[5](https://arxiv.org/html/2607.16263#S4.F5)illustrates the validation performance trajectories over training\. To enable a fair comparison between the step\-based DPO \(evaluated every 250 steps\) and the epoch\-based regression \(evaluated every epoch\), we align their training progress on a relative scale\. We observe that models trained without weak supervision exhibit erratic performance oscillations and ”spiky” validation curves\. In contrast, our unified preference objective yields a remarkably smooth convergence trajectory, suggesting that the massive preference pairs provide a flatter and more robust optimization landscape\.
### 4\.7Analysis of Preference Construction
To investigate the intrinsic behavior of preference\-based learning in antibody expression modelling, we first evaluate whether DPO remains effective under a strictly supervised setting\. Table[5](https://arxiv.org/html/2607.16263#S4.T5)compares the performance of the supervised baseline against DPO trained solely on the 1\.2k labeled sequences from Exp\-1K\. Interestingly, we observe that without the inclusion of large\-scale weak supervision \(Camel\-4M\), the pure DPO approach yields slightly inferior results compared to direct regression\. We attribute this to the fact that preference learning provides a sparser gradient signal than point\-wise regression when the labeled dataset is extremely small\. This strongly emphasizes the necessity of incorporating massive weakly labeled data within the preference learning framework\.
MethodMCC↑\\uparrowTau↑\\uparrowSCC↑\\uparrowSupervised29\.137\.642\.2w/o Camel\-4M32\.835\.136\.2Ours43\.746\.350\.8Table 5:Comparison between supervised regression and DPO on the Exp\-1K labeled set \(w/o Camel\-4M\)\.ThresholdMCC↑\\uparrowTau↑\\uparrowSCC↑\\uparrowNo threshold43\.746\.350\.80\.3135\.737\.638\.50\.596\.418\.319\.4Table 6:Impact of similarity\-based pair filtering on DPO performance\. Thresholds are chosen based on the local minima of the dataset’s similarity distribution\.Building on this, we examine the impact of sequence similarity on preference construction\. We selected two specific thresholds to partition the dataset, which correspond to the natural local minima \(Figure[6](https://arxiv.org/html/2607.16263#A1.F6)\) in the pairwise similarity distribution of Exp\-1K\. This allows us to compare the effectiveness of learning from closely related mutants versus more divergent antibody scaffolds\. Our results \(Table[6](https://arxiv.org/html/2607.16263#S4.T6)\) demonstrate that utilizing the full distribution consistently yields superior performance\. This justifies our choice to utilize the entire sequence space without heuristic filtering\.
## 5Conclusion
In this work, we present a unified preference\-based framework for antibody expression modelling, addressing the scarcity of labeled experimental data by integrating large\-scale weak supervision\. By adapting Direct Preference Optimization \(DPO\) to protein language models with a union\-masked log\-likelihood objective, we successfully leverage massive immunization data to learn biological priors\. Our experiments demonstrate that while traditional binary classifiers remain highly effective for simple expressibility categorization, our preference\-based approach offers superior performance in complex ranking tasks and exhibits significantly better generalization and optimization stability\. This discovery suggests that preference learning acts as a robust biophysical regularizer, providing a scalable solution for fine\-grained antibody developability prediction\. Future research will focus on hybrid architectures that combine the discriminative power of binary supervision with the ranking precision of preference\-based alignment\.
## Impact Statement
This paper presents work whose goal is to advance the field of AI for drug discovery\. There are many potential societal consequences of our work, none of which we feel must be specifically highlighted here\.
## References
- E\. C\. Alley, G\. Khimulya, S\. Biswas, M\. AlQuraishi, and G\. M\. Church \(2019\)Unified rational protein engineering with sequence\-based deep representation learning\.Nature methods16\(12\),pp\. 1315–1322\.Cited by:[§1](https://arxiv.org/html/2607.16263#S1.p5.1),[§3\.4](https://arxiv.org/html/2607.16263#S3.SS4.SSS0.Px1.p1.1)\.
- A\. Arsiwala, R\. Bhatt, Y\. Yang, P\. Quintero Cadena, K\. Anderson, X\. Ao, L\. van Niekerk, A\. Rosenbaum, A\. Bhatt, A\. Smith, L\. Grippo, X\. Cao, R\. Cohen, J\. Patel, O\. Allen, A\. Faraj, A\. Nandy, J\. Hocking, B\. Tural, S\. Salvador, J\. Jacobowitz, K\. Schaven, M\. Sherman, S\. Shah, P\. M\. Tessier, and D\. Borhani \(2025\)A high\-throughput platform for biophysical antibody developability assessment to enable ai/ml model training\.bioRxiv\.External Links:[Document](https://dx.doi.org/10.1101/2025.05.01.651684),[Link](https://www.biorxiv.org/content/early/2025/05/02/2025.05.01.651684),https://www\.biorxiv\.org/content/early/2025/05/02/2025\.05\.01\.651684\.full\.pdfCited by:[Table 1](https://arxiv.org/html/2607.16263#S1.T1.5.5.7)\.
- J\. Barton, J\. D\. Galson, and J\. Leem \(2024\)Enhancing antibody language models with structural information\.BioRxiv,pp\. 2023–12\.Cited by:[Table 2](https://arxiv.org/html/2607.16263#S4.T2.7.7.2.1.1.2)\.
- S\. Biswas, G\. Khimulya, E\. C\. Alley, K\. M\. Esvelt, and G\. M\. Church \(2021\)Low\-n protein engineering with data\-efficient deep learning\.Nature methods18\(4\),pp\. 389–396\.Cited by:[§1](https://arxiv.org/html/2607.16263#S1.p5.1),[§3\.4](https://arxiv.org/html/2607.16263#S3.SS4.SSS0.Px1.p1.1)\.
- N\. Brandes, D\. Ofer, Y\. Peleg, N\. Rappoport, and M\. Linial \(2022\)ProteinBERT: a universal deep\-learning model of protein sequence and function\.Bioinformatics38\(8\),pp\. 2102–2110\.Cited by:[§2](https://arxiv.org/html/2607.16263#S2.SS0.SSS0.Px2.p1.2),[Table 2](https://arxiv.org/html/2607.16263#S4.T2.13.13.2.1.1.2)\.
- M\. Chungyoun, J\. Ruffolo, and J\. Gray \(2024\)FLAb: benchmarking deep learning methods for antibody fitness prediction\.BioRxiv,pp\. 2024–01\.Cited by:[Table 1](https://arxiv.org/html/2607.16263#S1.T1.2.2.7.2.1.1.1)\.
- J\. Dunbar and C\. M\. Deane \(2016\)ANARCI: antigen receptor numbering and receptor classification\.Bioinformatics32\(2\),pp\. 298–300\.Cited by:[§A\.1](https://arxiv.org/html/2607.16263#A1.SS1.p2.1),[§1](https://arxiv.org/html/2607.16263#S1.p5.1)\.
- C\. Ferragu, J\. D\. Ziegler, N\. Deutschmann, A\. Lindoulsi, E\. Bixby, and C\. M\. Team \(2025\)G\-dpo: scalable preference optimization for protein language models\.arXiv preprint arXiv:2510\.19474\.Cited by:[§1](https://arxiv.org/html/2607.16263#S1.p5.1),[§2](https://arxiv.org/html/2607.16263#S2.SS0.SSS0.Px1.p1.1),[§3\.3\.2](https://arxiv.org/html/2607.16263#S3.SS3.SSS2.p3.2)\.
- M\. F\. Flajnik, N\. Deschacht, and S\. Muyldermans \(2011\)A case of convergence: why did a simple alternative to canonical antibodies arise in sharks and camels?\.PLoS biology9\(8\),pp\. e1001120\.Cited by:[§C\.3](https://arxiv.org/html/2607.16263#A3.SS3.p1.1)\.
- \[10\]C\. W\. Gordon, A\. X\. Lu, and P\. AbbeelProtein language model fitness is a matter of preference\.InThe Thirteenth International Conference on Learning Representations,Cited by:[§3\.4](https://arxiv.org/html/2607.16263#S3.SS4.SSS0.Px1.p1.1)\.
- L\. Hanke, L\. Vidakovics Perez, D\. J\. Sheward, H\. Das, T\. Schulte, A\. Moliner\-Morro, M\. Corcoran, A\. Achour, G\. B\. Karlsson Hedestam, B\. M\. Hällberg,et al\.\(2020\)An alpaca nanobody neutralizes sars\-cov\-2 by blocking receptor interaction\.Nature communications11\(1\),pp\. 4420\.Cited by:[§1](https://arxiv.org/html/2607.16263#S1.p2.1)\.
- \[12\]A\. Hawkins\-Hooker, J\. Kmec, O\. Bent, and P\. DuckworthLikelihood\-based fine\-tuning of protein language models for few\-shot fitness prediction and design\.InICML’24 Workshop ML for Life and Material Science: From Theory to Industry Applications,Cited by:[§3\.3\.2](https://arxiv.org/html/2607.16263#S3.SS3.SSS2.p3.2)\.
- M\. Hebditch, M\. A\. Carballo\-Amador, S\. Charonis, R\. Curtis, and J\. Warwicker \(2017\)Protein–sol: a web tool for predicting protein solubility from sequence\.Bioinformatics33\(19\),pp\. 3098–3100\.Cited by:[§2](https://arxiv.org/html/2607.16263#S2.SS0.SSS0.Px2.p1.2)\.
- C\. Hsu, H\. Nisonoff, C\. Fannjiang, and J\. Listgarten \(2022\)Learning protein fitness models from evolutionary and assay\-labeled data\.Nature biotechnology40\(7\),pp\. 1114–1122\.Cited by:[§3\.4](https://arxiv.org/html/2607.16263#S3.SS4.SSS0.Px1.p1.1)\.
- T\. Jain, T\. Sun, S\. Durand, A\. Hall, N\. R\. Houston, J\. H\. Nett, B\. Sharkey, B\. Bobrowicz, I\. Caffry, Y\. Yu,et al\.\(2017\)Biophysical properties of the clinical\-stage antibody landscape\.Proceedings of the National Academy of Sciences114\(5\),pp\. 944–949\.Cited by:[Table 1](https://arxiv.org/html/2607.16263#S1.T1.3.3.7)\.
- F\. Jung, K\. Frey, D\. Zimmer, and T\. Mühlhaus \(2023\)DeepSTABp: a deep learning approach for the prediction of thermal protein stability\.International Journal of Molecular Sciences24\(8\),pp\. 7444\.Cited by:[§2](https://arxiv.org/html/2607.16263#S2.SS0.SSS0.Px2.p1.2)\.
- H\. Kenlay, F\. A\. Dreyer, A\. Kovaltsuk, D\. Miketa, D\. Pires, and C\. M\. Deane \(2024\)Large scale paired antibody language models\.PLOS Computational Biology20\(12\),pp\. e1012646\.Cited by:[Table 2](https://arxiv.org/html/2607.16263#S4.T2.10.10.2.1.1.2)\.
- P\. M\. Khade, M\. Maser, V\. Gligorijevic, and A\. Watkins \(2023\)Mixed structure\- and sequence\-based approach for protein graph neural networks with application to antibody developability prediction\.bioRxiv\.External Links:[Document](https://dx.doi.org/10.1101/2023.06.26.546331),[Link](https://www.biorxiv.org/content/early/2023/06/28/2023.06.26.546331),https://www\.biorxiv\.org/content/early/2023/06/28/2023\.06\.26\.546331\.full\.pdfCited by:[§2](https://arxiv.org/html/2607.16263#S2.SS0.SSS0.Px2.p1.2)\.
- H\. Y\. Kim, S\. Kim, W\. Park, and D\. Kim \(2023\)G\-rank: an equivariant graph neural network for the scoring of protein–protein docking models\.Bioinformatics Advances3\(1\),pp\. vbad011\.Cited by:[§2](https://arxiv.org/html/2607.16263#S2.SS0.SSS0.Px2.p1.2)\.
- M\. Lefranc, C\. Pommié, M\. Ruiz, V\. Giudicelli, E\. Foulquier, L\. Truong, V\. Thouvenin\-Contet, and G\. Lefranc \(2003a\)IMGT unique numbering for immunoglobulin and t cell receptor variable domains and ig superfamily v\-like domains\.Developmental & Comparative Immunology27\(1\),pp\. 55–77\.Cited by:[§A\.1](https://arxiv.org/html/2607.16263#A1.SS1.p2.1),[§1](https://arxiv.org/html/2607.16263#S1.p5.1)\.
- M\. Lefranc, C\. Pommié, M\. Ruiz, V\. Giudicelli, E\. Foulquier, L\. Truong, V\. Thouvenin\-Contet, and G\. Lefranc \(2003b\)IMGT unique numbering for immunoglobulin and t cell receptor variable domains and ig superfamily v\-like domains\.Developmental & Comparative Immunology27\(1\),pp\. 55–77\.Cited by:[§A\.1](https://arxiv.org/html/2607.16263#A1.SS1.p2.1),[§1](https://arxiv.org/html/2607.16263#S1.p5.1)\.
- Z\. Liang, T\. Wang, Y\. Sun, W\. Yang, Z\. Liu, J\. Fei, Y\. Guo, Q\. Ma, Q\. Pan, and L\. Ren \(2015\)A comprehensive analysis of immunoglobulin heavy chain genes in the bactrian camel \(camelus bactrianus\)\.Frontiers of Agricultural Science and Engineering2\(3\),pp\. 249–259\.Cited by:[§C\.3](https://arxiv.org/html/2607.16263#A3.SS3.p1.1)\.
- J\. Meier, R\. Rao, R\. Verkuil, J\. Liu, T\. Sercu, and A\. Rives \(2021\)Language models enable zero\-shot prediction of the effects of mutations on protein function\.Advances in neural information processing systems34,pp\. 29287–29303\.Cited by:[§1](https://arxiv.org/html/2607.16263#S1.p5.1),[§2](https://arxiv.org/html/2607.16263#S2.SS0.SSS0.Px2.p1.2),[§3\.2](https://arxiv.org/html/2607.16263#S3.SS2.p1.1),[§3\.3\.2](https://arxiv.org/html/2607.16263#S3.SS3.SSS2.p1.4)\.
- R\. Rafailov, A\. Sharma, E\. Mitchell, C\. D\. Manning, S\. Ermon, and C\. Finn \(2023\)Direct preference optimization: your language model is secretly a reward model\.Advances in neural information processing systems36,pp\. 53728–53741\.Cited by:[§1](https://arxiv.org/html/2607.16263#S1.p4.1),[§3\.3\.1](https://arxiv.org/html/2607.16263#S3.SS3.SSS1.p1.7)\.
- J\. Salazar, D\. Liang, T\. Q\. Nguyen, and K\. Kirchhoff \(2020\)Masked language model scoring\.InProceedings of the 58th annual meeting of the association for computational linguistics,pp\. 2699–2712\.Cited by:[§1](https://arxiv.org/html/2607.16263#S1.p5.1),[§3\.2](https://arxiv.org/html/2607.16263#S3.SS2.p1.1),[§3\.3\.2](https://arxiv.org/html/2607.16263#S3.SS3.SSS2.p1.4)\.
- M\. Sun, H\. Cao, Z\. Zhang, C\. Wei, L\. Wang, T\. Jia, Z\. Liu, T\. Fu, X\. Tang, Y\. Choi,et al\.\(2025\)RiboPO: preference optimization for structure\-and stability\-aware rna design\.arXiv preprint arXiv:2510\.21161\.Cited by:[§2](https://arxiv.org/html/2607.16263#S2.SS0.SSS0.Px1.p1.1)\.
- A\. C\. Szkodny and K\. H\. Lee \(2024\)A systemic approach to identifying sequence frameworks that decrease mab production in a transient chinese hamster ovary cell expression system\.Biotechnology Progress40\(5\),pp\. e3466\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1002/btpr.3466),[Link](https://aiche.onlinelibrary.wiley.com/doi/abs/10.1002/btpr.3466),https://aiche\.onlinelibrary\.wiley\.com/doi/pdf/10\.1002/btpr\.3466Cited by:[Table 1](https://arxiv.org/html/2607.16263#S1.T1.4.4.7)\.
- E\. Tankhilevich, S\. Martinez Cuesta, I\. Barrett, C\. Berg, L\. Holmberg Schiavone, and A\. R\. Leach \(2026\)RP3Net: a deep learning model for predicting recombinant protein production in escherichia coli\.Bioinformatics,pp\. btag003\.Cited by:[§2](https://arxiv.org/html/2607.16263#S2.SS0.SSS0.Px2.p1.2)\.
- V\. Thumuluri, H\. Martiny, J\. J\. Almagro Armenteros, J\. Salomon, H\. Nielsen, and A\. R\. Johansen \(2022\)NetSolP: predicting protein solubility in escherichia coli using language models\.Bioinformatics38\(4\),pp\. 941–946\.Cited by:[§2](https://arxiv.org/html/2607.16263#S2.SS0.SSS0.Px2.p1.2)\.
- H\. Tsuruta, H\. Yamazaki, R\. Maeda, R\. Tamura, and A\. Imura \(2024\)A sars\-cov\-2 interaction dataset and vhh sequence corpus for antibody language models\.Advances in Neural Information Processing Systems37,pp\. 116149–116171\.Cited by:[§1](https://arxiv.org/html/2607.16263#S1.p2.1)\.
- Y\. Wen, C\. Xu, J\. Y\. Hu, and H\. Liu \(2024\)Alignab: pareto\-optimal energy alignment for designing nature\-like antibodies\.arXiv preprint arXiv:2412\.20984\.Cited by:[§2](https://arxiv.org/html/2607.16263#S2.SS0.SSS0.Px1.p1.1)\.
- J\. Xu, Z\. Gao, X\. Zhou, hujie, X\. Cheng, L\. Song, G\. Chen, P\. Heng, and J\. Qiu \(2025\)Protein inverse folding from structure feedback\.InThe Thirty\-ninth Annual Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=ORsrbGTXQB)Cited by:[§2](https://arxiv.org/html/2607.16263#S2.SS0.SSS0.Px1.p1.1)\.
- F\. Xue, A\. Kubaney, Z\. Guo, J\. K\. Min, G\. Liu, Y\. Yang, and D\. Baker \(2025\)Improving protein sequence design through designability preference optimization\.arXiv preprint arXiv:2506\.00297\.Cited by:[§2](https://arxiv.org/html/2607.16263#S2.SS0.SSS0.Px1.p1.1)\.
- T\. Zhang, X\. Cheng, D\. Yu, F\. Lin, N\. Hou, X\. Cheng, S\. Hao, J\. Wei, L\. Ma, Y\. Fu,et al\.\(2018\)Genetic removal of the ch1 exon enables the production of heavy chain\-only igg in mice\.Frontiers in Immunology9,pp\. 2202\.Cited by:[§C\.3](https://arxiv.org/html/2607.16263#A3.SS3.p1.1)\.
- J\. Zhao, C\. Zhang, and Y\. Luo \(2024\)Contrastive fitness learning: reprogramming protein language models for low\-n learning of protein fitness landscape\.InInternational Conference on Research in Computational Molecular Biology,pp\. 470–474\.Cited by:[§3\.3\.2](https://arxiv.org/html/2607.16263#S3.SS3.SSS2.p3.2)\.
- X\. Zhao, Y\. Tang, R\. Monsia, V\. J\. Cantu, A\. K\. Ramesh, X\. Liu, Z\. An, X\. Jiang, and Y\. Kim \(2025\)Structure\-aware antibody design with affinity\-optimized inverse folding\.arXiv preprint arXiv:2512\.17815\.Cited by:[§2](https://arxiv.org/html/2607.16263#S2.SS0.SSS0.Px1.p1.1)\.
- X\. Zhou, D\. Xue, R\. Chen, Z\. Zheng, L\. Wang, and Q\. Gu \(2024\)Antigen\-specific antibody design via direct energy\-based preference optimization\.Advances in Neural Information Processing Systems37,pp\. 120861–120891\.Cited by:[§2](https://arxiv.org/html/2607.16263#S2.SS0.SSS0.Px1.p1.1)\.
## Appendix AExperimental Implementation Details
### A\.1Dataset Statistics and Preprocessing
Exp\-1K Statistics\.The Exp\-1K dataset serves as our primary benchmark for quantitative yield prediction\. The sequence counts for the 80/10/10 split are 1,003 \(Train\), 125 \(Val\), and 126 \(Test\)\.
IMGT Alignment\.We use ANARCI\(Dunbar and Deane,[2016](https://arxiv.org/html/2607.16263#bib.bib39)\)to number all sequences according to the IMGT standard numbering\(Lefrancet al\.,[2003a](https://arxiv.org/html/2607.16263#bib.bib37),[b](https://arxiv.org/html/2607.16263#bib.bib38)\)\. Sequences are mapped to 128 standard positions, with missing residues filled by gap tokens \(−\-\)\. This alignment is crucial for the position\-wise log\-likelihood calculation in our DPO framework\.
Hamming Similarity\.To quantify the similarity between two antibody sequences, we adapt the Hamming similarity based on the IMGT numbering scheme\. Given two antibody sequencess1s\_\{1\}ands2s\_\{2\}, we first map them to a standardized coordinate system of IMGT residue indicesℐ\\mathcal\{I\}\(e\.g\., positions 1 to 128\)\. Letpos\(s,i\)\\text\{pos\}\(s,i\)be the amino acid residue of sequencessat the IMGT indexi∈ℐi\\in\\mathcal\{I\}\. The similarity scoreh\(s1,s2\)h\(s\_\{1\},s\_\{2\}\)is defined as the fraction of identical residues across the intersection of their aligned positions:
h\(s1,s2\)=∑i∈ℐs1∩ℐs2𝕀\(pos\(s1,i\)=pos\(s2,i\)\)\|ℐs1∩ℐs2\|,h\(s\_\{1\},s\_\{2\}\)=\\frac\{\\sum\_\{i\\in\\mathcal\{I\}\_\{s\_\{1\}\}\\cap\\mathcal\{I\}\_\{s\_\{2\}\}\}\\mathbb\{I\}\(\\text\{pos\}\(s\_\{1\},i\)=\\text\{pos\}\(s\_\{2\},i\)\)\}\{\|\\mathcal\{I\}\_\{s\_\{1\}\}\\cap\\mathcal\{I\}\_\{s\_\{2\}\}\|\},where𝕀\(⋅\)\\mathbb\{I\}\(\\cdot\)is the indicator function and\|ℐs1∩ℐs2\|\|\\mathcal\{I\}\_\{s\_\{1\}\}\\cap\\mathcal\{I\}\_\{s\_\{2\}\}\|represents the number of shared IMGT positions\.
Figure 6:Pairwise similarity distribution of the Exp\-1K dataset\. The similarity is calculated as the IMGT\-aligned Hamming identity across all sequence pairs\.Pairwise Similarity Distribution\.To facilitate the analysis of preference construction \(Section[4\.7](https://arxiv.org/html/2607.16263#S4.SS7)\), we compute the pairwise Hamming similarity for all sequences in Exp\-1K based on the IMGT\-aligned positions\. As illustrated in Figure[6](https://arxiv.org/html/2607.16263#A1.F6), the similarity distribution exhibits a distinct tri\-modal pattern, characterizing different levels of sequence divergence \(e\.g\., clonal variants vs\. divergent scaffolds\)\. To rigorously test the effect of similarity\-based filtering, we select the two local minima \(valleys\) of this distribution,d1=0\.31d\_\{1\}=0\.31andd2=0\.59d\_\{2\}=0\.59, as the thresholds for the sensitivity analysis in Table[6](https://arxiv.org/html/2607.16263#S4.T6)\. This systematic approach ensures that our thresholds are grounded in the intrinsic structural hierarchy of the antibody library\.
### A\.2Model Configurations
Table[7](https://arxiv.org/html/2607.16263#A1.T7)summarizes the architectural details of the two aLMs and one pLM used as backbones\.
ModelLayersHeadsDimParamsHuggingFace PathAntiBERTa216161024202Malchemab/antiberta2IgBERT30161024420MExscientia/IgBertProtBERT30161024420MRostlab/prot\_bertTable 7:Architectural configurations of the pLM backbones\.
### A\.3Hyperparameter Details
We use the AdamW optimizer with defaultβ1=0\.9,β2=0\.999\\beta\_\{1\}=0\.9,\\beta\_\{2\}=0\.999\. No gradient clipping is applied\. Specific batch sizes and GPU memory usage are listed in Table[A\.3](https://arxiv.org/html/2607.16263#A1.SS3)\.
ModelBatch SizePeak MemoryPrecision\\rowcolorlightgrayStage IAntiBERTa225632\.0GBbf16IgBERT25645\.5GBbf16ProtBERT25645\.5GBbf16\\rowcolorlightgrayStage IIAntiBERTa212832\.1GBbf16IgBERT12845\.7GBbf16ProtBERT12845\.7GBbf16
Table 8:Computational details and batch sizes per GPU \(H100\)\.### A\.4Convergence and Early Stopping Criteria
We observed that different training stages require distinct convergence strategies to maintain a balance between domain adaptation and the risk of overfitting to the relatively scarce labeled data\.
Stage I \(Continual Pre\-training\):For the domain\-alignment phase, the model is trained for a fixed duration of 1 epoch on the Camel\-4M dataset\. Given the large\-scale nature of the unlabeled data, a single pass is sufficient to adapt the model to the antibody\-specific sequence distribution without collapsing the general protein representations learned during initial pre\-training\.
Stage II \(Preference Fine\-tuning\):n the preference optimization phase, we employ an Early Stopping mechanism to prevent over\-optimization on preference pairs\. Training is terminated if the validation metric \(specifically, the geometric mean of Matthews Correlation Coefficient and Kendall’sτ\\tau\) fails to improve for 10 consecutive evaluation steps\. \(1\) For Supervised Baselines: Evaluation occurs once per epoch\. Thus, the patience for early stopping is equivalent to 10 epochs\. \(2\) For DPO Fine\-tuning: Due to the higher density of preference pairs, evaluation occurs every 250 gradient steps\. Consequently, the early stopping patience is equivalent to 2500 training steps\.
### A\.5Probabilistic Sampling for Supervision Balancing
A key challenge in Stage II is the extreme imbalance between the number of strong supervision pairs\|𝒟suppair\|\|\\mathcal\{D\}\_\{\\text\{sup\}\}^\{\\text\{pair\}\}\|and weak supervision pairs\|𝒟weakpair\|\|\\mathcal\{D\}\_\{\\text\{weak\}\}^\{\\text\{pair\}\}\|\. To prevent the weak supervision signal from overwhelming the high\-quality quantitative data, we implement a probabilistic sampling strategy at the\_\_getitem\_\_level of our data loader\.
We define a sampling rate hyperparameterα\\alpha\(set to 1000 in our main experiments\) to oversample the supervised pairs\. The probability of selecting a sample from either the strong supervision set \(SS\) or the weak supervision set \(LL\) is calculated as follows:LetNS=\|𝒟suppair\|N\_\{S\}=\|\\mathcal\{D\}\_\{\\text\{sup\}\}^\{\\text\{pair\}\}\|andNL=\|𝒟weakpair\|N\_\{L\}=\|\\mathcal\{D\}\_\{\\text\{weak\}\}^\{\\text\{pair\}\}\|\. The selection probabilitiesPSP\_\{S\}andPLP\_\{L\}are defined as:P\_S =NS⋅α\(NS⋅α\) \+ NL, P\_L =NL\(NS⋅α\) \+ NLDuring training, for each data request, a random variableu∼Uniform\(0,1\)u\\sim\\text\{Uniform\}\(0,1\)is generated\. Ifu<PSu<P\_\{S\}, a pair is sampled from the supervised set; otherwise, it is sampled from the weakly supervised set\. This ensures that the expected ratio of supervised gradient updates remains controllable and consistent regardless of the raw size of the unlabeled corpus\.
### A\.6Mathematical Definitions of Metrics
Matthews Correlation Coefficient \(MCC\)\.Unlike ROC\-AUC, which can be overly optimistic on imbalanced data, MCC provides a balanced score even if the classes are of very different sizes:
MCC=TP×TN−FP×FN\(TP\+FP\)\(TP\+FN\)\(TN\+FP\)\(TN\+FN\)\\text\{MCC\}=\\frac\{TP\\times TN\-FP\\times FN\}\{\\sqrt\{\(TP\+FP\)\(TP\+FN\)\(TN\+FP\)\(TN\+FN\)\}\}\(1\)
Kendall’sτ\\tau\(Tau\-b\)\.To handle the heavy ties \(Yield=0\) in Exp\-1K, we use the standardτb\\tau\_\{b\}which adjusts for ties:
τb=C−D\(C\+D\+Tx\)\(C\+D\+Ty\),\\tau\_\{b\}=\\frac\{C\-D\}\{\\sqrt\{\(C\+D\+T\_\{x\}\)\(C\+D\+T\_\{y\}\)\}\},\(2\)whereCCis the count of concordant pairs,DDis the count of discordant pairs,TxT\_\{x\}is the number of ties only inxx, andTyT\_\{y\}is the number of ties only inyy\.
SCC on Expressible Subset\.We define the filtered Spearman correlation coefficient \(SCC\) as:
SCCexpr=Spearman\(\{\(yi,y^i\)∣yi\>0\}\)\\text\{SCC\}\_\{\\text\{expr\}\}=\\text\{Spearman\}\(\\\{\(y\_\{i\},\\hat\{y\}\_\{i\}\)\\mid y\_\{i\}\>0\\\}\)\(3\)
Composite Metric \(Avg\)\.We use the geometric mean to penalize models that fail significantly in either classification or ranking:
Avg=MCC×τb\\text\{Avg\}=\\sqrt\{\\text\{MCC\}\\times\\tau\_\{b\}\}\(4\)
## Appendix BBinary Supervision with Weak Positive Labels
In this ablation, we examine whether large\-scale weakly labeled data can be directly exploited by binary classification\. We focus on the Supervised\-Binary baseline using AntiBERTa2 as the backbone\. Based on preliminary experiments, we observe that initializing from the original pretrained weights consistently outperforms continual pretraining for binary supervision, and therefore fix the backbone accordingly\. We augment the labeled training set with subsets of Camel\-4M, which are assumed to be expressible\. To address the resulting class imbalance, we oversample the non\-expressible class and report results under the best\-performing sampling ratio\.
Camel\-4M SizeAUC↑\\uparrowMCC↑\\uparrowNone \(Exp\-1K only\)86\.235\.71K85\.341\.610K85\.951\.9100K87\.856\.71M86\.750\.74M87\.148\.3Ours \(4M\)82\.343\.7Table 9:Effect of weak positive samples on supervised binary training using AntiBERTa2\.For the narrow task of binary expressibility classification, the standard supervised binary network outperforms our preference\-based model \(Table[9](https://arxiv.org/html/2607.16263#A2.T9)\)\. We attribute this to the fact that a binary classifier directly optimizes the decision boundary for a single threshold, whereas our framework treats expressibility as a ranking problem within a much broader sequence space\.
## Appendix CDataset Generation
### C\.1Exp\-1K
All antibody sequences in Exp\-1K were transiently transfected into mammalian expression systems \(CHO or HEK293 cells\)\. Antibody expression was evaluated using SDS\-PAGE–based expression gels, and yield was estimated semi\-quantitatively based on band intensity relative to internal standards\. Reported yields are expressed in mg/L\. Antibodies with no detectable expression band were assigned a yield of<<1 mg/L\.
### C\.2Camel\-4M
Camel\-4M consists of antibody libraries derived from PBMC samples of immunized camelid species, including alpacas and llamas\. Following immunization with diverse therapeutic targets, PBMCs were isolated and used to construct VHH antibody libraries, which were subsequently sequenced using next\-generation sequencing \(MiSeq\) to obtain non\-redundant VHH sequences\. Because these libraries are derived from PBMCs of immunized animals, the resulting sequences predominantly represent mature, antigen\-experienced B cells\. Antibodies produced by such cells have undergone immune selection for proper folding and secretion, resulting in an enrichment of expressible VHHs\.
### C\.3Evidence of why Camelid\-derived sequences are mostly expressible\.
Theoretical evidence\.Camelid\-derived VHHs are well known to exhibit high expression and solubility\(Flajniket al\.,[2011](https://arxiv.org/html/2607.16263#bib.bib43); Lianget al\.,[2015](https://arxiv.org/html/2607.16263#bib.bib44)\)\. Camelid species have uniquely evolved heavy\-chain–only antibodies \(HCAbs\), in which the antigen\-binding domain \(VHH\) functions naturally without a paired light chain\. As part of this evolutionary adaptation, VHHs contain canonical framework mutations, particularly in the former VH–VL interface within FR2, where hydrophobic residues are replaced by more hydrophilic ones to promote proper folding, solubility, and secretion\. In contrast, conventional VH domains derived from VH–VL antibodies lack these canonical FR2 mutations and are evolutionarily optimized to pair with a light chain, making isolated VH\-derived formats, including VH\-based heavy\-chain antibodies, generally more difficult to express\(Zhanget al\.,[2018](https://arxiv.org/html/2607.16263#bib.bib45)\)\. Together, these evolutionary and structural features explain the enrichment of expressible sequences in camelid\-derived VHH libraries\.
Internal evidence\.In addition to prior biological intuition, we provide direct internal experimental evidence supporting the assumption that camelid\-derived antibody sequences are predominantly expressible\. We randomly selected 17 VHH sequences from the camelid\-derived pool and measured their recombinant expression yields under the same experimental protocol used for the Exp\-1K dataset\.
Among these 17 sequences, 16 showed detectable expression \(yield\>0\>0mg/L\), corresponding to an expressibility rate of94\.1%94\.1\\%\. This observation is consistent with the hypothesis that immunization\-derived camel VHH sequences are mostly expressible\.
To further characterize the expression behavior of these sequences, Table[10](https://arxiv.org/html/2607.16263#A3.T10)reports the distribution of measured yield scores across several intervals\. Rather than concentrating near the detection limit, most sequences exhibit non\-trivial expression levels, indicating that camelid\-derived sequences are not only binary\-expressible, but often moderately to highly expressible\.
Yield range \(mg/L\)Number of sequences=0=01\(0,50\)\(0,50\)1\[50,100\)\[50,100\)2\[100,200\)\[100,200\)6\>200\>2007Total17Table 10:Yield score distribution of 17 randomly sampled camelid\-derived VHH sequences\.Similar Articles
AbICL: In-Context Learning for Antigen-Specific Antibody Affinity Ranking
AbICL proposes an in-context learning framework for antigen-specific antibody affinity ranking, combining a pretrained structural encoder with a context ranking head to leverage labeled demonstrations for test-time adaptation without gradient updates.
Learning Task-Specific Antibody Representations via Function-Aware Masking
This paper introduces function-aware masking, a pretraining algorithm for antibody language models that aligns mask placement with functional priors, yielding significant improvements on structure and CDR-related tasks.
Natural-Language-Guided Generator-Agnostic Shortlisting for Protein Binder Design
This paper explores using large language models to generate ranking policies for shortlisting protein binders from candidate pools, showing modest improvements over baseline methods in de novo design workflows.
A Consensus-Based Framework for Relative Preference Evaluation of Large Language Models
This paper introduces a consensus-based framework for evaluating LLM responses by having a panel of models rank anonymized outputs, producing a Relative Intelligence Index (RII) as a proxy for relative quality across domains.
Distributionally Robust Listwise Preference Optimization
This paper proposes a distributionally robust listwise preference optimization method for LLM alignment that handles ranking-label uncertainty, with a tractable objective and strong convergence guarantees.