PFArena: Benchmarking Language Models for Protein Modification
Summary
PFArena introduces a benchmark for evaluating language models on protein modification tasks, comparing protein language models and large language models to assess their capabilities in biological applications.
View Cached Full Text
Cached at: 09/25/26, 09:34 AM
# PFArena: Benchmarking Language Models for Protein Modification Source: [https://arxiv.org/html/2609.28921](https://arxiv.org/html/2609.28921) Figure 1:Benchmark performance across the four protein modification tasks in PFArena\.Method classes are distinguished by cell background color:protein language models,large language models, andagents\. For agents, Claude denotes Claude Opus 5 and GPT denotes GPT\-6 Astra\. Methods are ranked by Recall@4040for taskT1and by Spearman correlation for tasksT2–T4\. Results for the remaining metrics are detailed in the experiments section\.††footnotetext:Correspondence:[zhouhao@air\.tsinghua\.edu\.cn](mailto:[email protected])###### Contents 1. [1Introduction](https://arxiv.org/html/2609.28921#S1) 2. [2PFArena](https://arxiv.org/html/2609.28921#S2)1. [2\.1Task Suite](https://arxiv.org/html/2609.28921#S2.SS1)1. [2\.1\.1T1: Single\-mutant generation](https://arxiv.org/html/2609.28921#S2.SS1.SSS1) 2. [2\.1\.2T2: Measurement\-free multi\-mutant ranking](https://arxiv.org/html/2609.28921#S2.SS1.SSS2) 3. [2\.1\.3T3: Anchor\-informed multi\-mutant ranking](https://arxiv.org/html/2609.28921#S2.SS1.SSS3) 4. [2\.1\.4T4: Single\-mutant\-informed multi\-mutant ranking](https://arxiv.org/html/2609.28921#S2.SS1.SSS4) 2. [2\.2Evaluation Objectives](https://arxiv.org/html/2609.28921#S2.SS2) 3. [2\.3Data Construction](https://arxiv.org/html/2609.28921#S2.SS3)1. [2\.3\.1Assay Collection and Harmonization](https://arxiv.org/html/2609.28921#S2.SS3.SSS1) 2. [2\.3\.2Quality Control](https://arxiv.org/html/2609.28921#S2.SS3.SSS2) 3. [3Experimental Setup](https://arxiv.org/html/2609.28921#S3)1. [3\.1Baselines](https://arxiv.org/html/2609.28921#S3.SS1) 2. [3\.2Metrics](https://arxiv.org/html/2609.28921#S3.SS2) 3. [3\.3Implementation Details](https://arxiv.org/html/2609.28921#S3.SS3) 4. [4Main Results](https://arxiv.org/html/2609.28921#S4)1. [4\.1Overall Results](https://arxiv.org/html/2609.28921#S4.SS1) 2. [4\.2Task\-specific Results](https://arxiv.org/html/2609.28921#S4.SS2)1. [4\.2\.1T1: Single\-mutant generation](https://arxiv.org/html/2609.28921#S4.SS2.SSS1) 2. [4\.2\.2T2: Measurement\-free multi\-mutant ranking](https://arxiv.org/html/2609.28921#S4.SS2.SSS2) 3. [4\.2\.3T3: Anchor\-informed multi\-mutant ranking](https://arxiv.org/html/2609.28921#S4.SS2.SSS3) 4. [4\.2\.4T4: Single\-mutant\-informed multi\-mutant ranking](https://arxiv.org/html/2609.28921#S4.SS2.SSS4) 5. [5Analysis](https://arxiv.org/html/2609.28921#S5)1. [5\.1Comparing PLMs and LLMs](https://arxiv.org/html/2609.28921#S5.SS1)1. [5\.1\.1Divergent Single\-mutant Proposal Strategies](https://arxiv.org/html/2609.28921#S5.SS1.SSS1) 2. [5\.1\.2Utilization of Multi\-mutant Component Additivity](https://arxiv.org/html/2609.28921#S5.SS1.SSS2) 3. [5\.1\.3Common Limitations across Single\-mutant and Multi\-mutant Tasks](https://arxiv.org/html/2609.28921#S5.SS1.SSS3) 2. [5\.2Dive into LLMs](https://arxiv.org/html/2609.28921#S5.SS2)1. [5\.2\.1Feedback\-guided In\-context Adaptation](https://arxiv.org/html/2609.28921#S5.SS2.SSS1) 2. [5\.2\.2Uncertainty Estimation](https://arxiv.org/html/2609.28921#S5.SS2.SSS2) 3. [5\.2\.3Test\-time Scaling](https://arxiv.org/html/2609.28921#S5.SS2.SSS3) 6. [6Related Work](https://arxiv.org/html/2609.28921#S6) 7. [7Conclusion](https://arxiv.org/html/2609.28921#S7) 8. [References](https://arxiv.org/html/2609.28921#bib) 9. [8Evaluation Templates](https://arxiv.org/html/2609.28921#S8)1. [8\.1Single\-mutant Generation](https://arxiv.org/html/2609.28921#S8.SS1) 2. [8\.2Multi\-mutant Ranking](https://arxiv.org/html/2609.28921#S8.SS2) 3. [8\.3Confidence Elicitation for Uncertainty Estimation](https://arxiv.org/html/2609.28921#S8.SS3) 10. [9Experimental Setup](https://arxiv.org/html/2609.28921#S9)1. [9\.1PLM Deployment](https://arxiv.org/html/2609.28921#S9.SS1) 2. [9\.2Test\-time Scaling Methods](https://arxiv.org/html/2609.28921#S9.SS2) 11. [10Extended Experimental Analysis](https://arxiv.org/html/2609.28921#S10)1. [10\.1Evaluation Completeness and Missing\-Query Handling](https://arxiv.org/html/2609.28921#S10.SS1) 2. [10\.2Complementary Strengths and Failure Modes of LLMs and PLMs](https://arxiv.org/html/2609.28921#S10.SS2) ## Abstract Protein modification requires navigating an immense sequence space, yet wet\-lab validation remains low\-throughput and costly\. Although computational paradigms including protein language models \(PLMs\), large language models \(LLMs\), and LLM\-based agents have shown promise in protein modification, their relative efficacy across realistic experimental decision\-making settings remains unclear\. To bridge this gap, we introducePFArena, a benchmark comprising four controlled task interfaces that cover single\-mutant generation and multi\-mutant ranking\. By providing varying levels of mutation fitness data,PFArenareflects four representative research scenarios characterized by differing degrees of prior experimental context\. We assess six PLMs, six LLMs, and five LLM\-based agents using complementary metrics to measure both peak and overall protein modification performance\. Our evaluation reveals that model performance shifts systematically with the availability of target\-specific experimental evidence: PLMs demonstrate proficiency in open\-ended single\-mutant generation by leveraging protein\-specific priors, whereas LLMs and agents perform strongly in multi\-mutant ranking, particularly when target\-specific fitness data are available\. Nevertheless, all model families face fundamental challenges with the increase of search space and mutation depth\. We release our code and benchmark suite to facilitate reproducible research in model\-assisted protein modification\. ## 1Introduction Protein modification is central to protein engineering, fundamentally relying on sequence mutations to alter protein properties\. This process results in an enormous number of possible substitutions and combinations, making experimental characterization costly, time\-consuming, and limited in throughput\. Reliable fitness predictions allow researchers to prioritize mutants with the highest expected value and focus measurements on informative regions of the sequence–fitness landscape, thereby accelerating iterative design–build–test\-learn cycles and supporting the discovery of mutations difficult to identify through empirical screening alone\. Computational models, including protein language models \(PLMs\) and large language models \(LLMs\), have shown substantial potential for protein modification\. PLMs learn sequence, structural, and evolutionary constraints from large\-scale protein data, enabling zero\-shot prediction and prioritization of functional mutants\[[1](https://arxiv.org/html/2609.28921#bib.bib16),[2](https://arxiv.org/html/2609.28921#bib.bib17),[3](https://arxiv.org/html/2609.28921#bib.bib18),[4](https://arxiv.org/html/2609.28921#bib.bib19),[5](https://arxiv.org/html/2609.28921#bib.bib21),[6](https://arxiv.org/html/2609.28921#bib.bib20)\]\. General\-purpose LLMs have also advanced rapidly on protein\-specific prediction tasks\. On ProteinGym Hard, the reported score increased from 37\.7% for Claude Opus 4\.7 to 39\.6% for Claude Opus 4\.8, followed by a further 7\.7\-percentage\-point gain for Claude Opus 5\[[7](https://arxiv.org/html/2609.28921#bib.bib4),[8](https://arxiv.org/html/2609.28921#bib.bib5)\]\. This sustained improvement highlights the growing ability of LLMs to interpret protein sequences through natural\-language interaction\. Nevertheless, it remains unclear when specialized PLMs, general\-purpose LLMs, or LLM\-based agents are preferable across the diverse decision settings encountered in protein modification\. To answer this question, we introducePFArena, an assay\-grounded benchmark that systematically evaluates these model families under four protein modification settings:T1: Single\-mutant generation,T2: Measurement\-free multi\-mutant ranking,T3: Anchor\-informed multi\-mutant rankingandT4: Single\-mutant\-informed multi\-mutant ranking\. These settings vary along two axes: the action space and the available experimental evidence\. They encompass both open\-ended generation and fixed\-pool ranking tasks, and cover conditions with different mutant information supplied\. This design enables a systematic comparison between PLMs and LLMs across single\- and multi\-mutant prioritization, while testing whether LLMs can translate target\-specific experimental evidence into better mutation ranking\. We evaluate six PLMs and six LLMs across four tasks covering 202 unique assays and 293 task instances\. Furthermore, to combine the reasoning capabilities of LLMs with protein\-specific tools, we evaluate five LLM\-based agents built on both the pioneering Biomni framework\[[9](https://arxiv.org/html/2609.28921#bib.bib31)\]and our lightweightAMixKittoolkit, detailed in[Section3\.3](https://arxiv.org/html/2609.28921#S3.SS3)\. Our study yields two main findings: - •Divergent strengths\.No single model family dominates across all four tasks\. PLMs hold an overall advantage in the open\-endedT1setting, where no target\-specific fitness measurements are available and performance depends largely on protein\-specific sequence, structural, and evolutionary priors\. Their predictions also span a broader range of residue\-class substitutions\. In contrast, LLM\-based systems perform strongly in the multi\-mutant ranking tasks \(T2–T4\)\. Their advantage is most evident when comprehensive target\-specific experimental evidence is provided \(T4\)\. Together, these results highlight the complementary strengths of protein\-specific representations and evidence\-guided reasoning\. - •Convergent bottlenecks\.Despite these distinct strengths, all three model families encounter shared limitations\. InT1, even the best\-performing method recovers fewer than 5% of the ground\-truth top\-4040mutations, showing that the single\-mutant search space remains difficult to cover under a limited prediction budget\. AcrossT2–T4, all model families show a substantial drop in ranking accuracy beyond two substitutions, revealing a shared sensitivity to mutation depth\. ## 2PFArena Figure 2:PFArena evaluates protein modification from the perspective of experimental scientists\.The benchmark is organized around four representative research scenarios that correspond to the questions scientists ask when deciding which mutants to study next\.T1: Single\-mutant generationgenerates high\-performing single mutations without candidate input\.T2: Measurement\-free multi\-mutant rankingranks multi\-mutant candidates without prior fitness measurements\.T3: Anchor\-informed multi\-mutant rankingranks multi\-mutant successors conditioned on a tested anchor mutant\.T4: Single\-mutant\-informed multi\-mutant rankingranks multi\-mutant candidates using measured single\-mutant fitness information\. These scenarios span single\- and multi\-mutant discovery under progressively richer experimental evidence, enabling systematic comparison of PLMs, LLMs, and agents in realistic protein\-engineering workflows\. The protein structure depicted here corresponds to PDB entry 1RQC\[[10](https://arxiv.org/html/2609.28921#bib.bib41)\]\.Conventional mutant\-effect benchmarks typically assess whether a model can assign fitness scores to prespecified mutations\[[11](https://arxiv.org/html/2609.28921#bib.bib12),[12](https://arxiv.org/html/2609.28921#bib.bib10)\]\. In contrast,PFArenaevaluates whether PLMs, LLMs, and agents can prioritize mutants across four task interfaces\. These interfaces vary in candidate space, spanning single substitutions and mutation combinations, and the amount and form of target\-specific evidence available in each task\. ### 2\.1Task Suite As summarized in[Figure2](https://arxiv.org/html/2609.28921#S2.F2),PFArenadefines four task interfaces that correspond to common decisions in protein\-engineering workflows\.T1considers the open\-ended generation of single\-mutant candidates from the full legal substitution space, whereasT2–T4consider the ranking of a supplied pool of multi\-mutant candidates under different levels of target\-specific experimental evidence\. Formal task definitions are provided in the following subsections, and the specification tables summarize the information available at the benchmark level\. Because model families expose different native interfaces, this information is presented through model\-appropriate interfaces\. Natural\-language task inputs are provided directly to LLMs and agents, whereas PLMs receive the subset supported by their architectures\. #### 2\.1\.1T1: Single\-mutant generation Task componentSpecificationDefinitionBudgeted generation of legal single substitutions without target\-specific mutation measurements\.InputWild\-type sequencexx, assay contextccand proposed budgetKK\.OutputProposed mutation set of at mostKKsingle substitutions\.InstanceInput: Wild\-type sequence:MSSSLGKE\.\.\.AMV Assay context: \- Primary task class: activity function\. \- Fitness type: organismal or cellular fitness\. \- Readout subclass: growth or selection proxy\. Proposed budgetKK: 40 Output: Proposed mutation set:P557L,K369Y,N434D,\.\.\.\.StatisticsAssays: 123; measured single mutants: 594,695\.Table 1:Task specification forT1: Single\-mutant generation\.In a single\-mutant campaign like deep mutational scanning \(DMS\), researchers assay a broad panel of legal single substitutions around a wild\-type protein and use the measurements to identify mutants with favorable assay\-specific phenotypes\. TaskT1evaluates the analogous budgeted prioritization problem: given the wild\-type sequence and assay context but no target\-specific mutation measurements, which single substitutions should be selected first under a fixed experimental budget? [Table1](https://arxiv.org/html/2609.28921#S2.T1)summarizes the task specification\. Formally, given a wild\-type sequencexx, assay contextcc, and a generation budgetKK, the model implements a generation function: fθ:\(x,c,K\)↦𝒪^,𝒪^⊆𝒱\(x\),\|𝒪^\|≤K,f\_\{\\theta\}:\(x,c,K\)\\mapsto\\widehat\{\\mathcal\{O\}\},\\qquad\\widehat\{\\mathcal\{O\}\}\\subseteq\\mathcal\{V\}\(x\),\\qquad\|\\widehat\{\\mathcal\{O\}\}\|\\leq K, where𝒱\(x\)\\mathcal\{V\}\(x\)denotes the legal single\-substitution space of a canonical wild\-type sequencexx\. For a sequence of lengthLL, this space contains19L19Lpossible substitutions under the standard amino\-acid alphabet\. The model must therefore return a limited set of legal candidates rather than a complete ordering of the entire space\. T1assesses highly selective discovery in a large single\-mutant space under a finite experimental budget\. A successful system should identify at least one highly promising mutation while also exploring a sufficiently broad region of the high\-fitness landscape\. For smaller PLMs, the task can be approached by evaluating all legal substitutions individually and ranking their scores\. For LLMs and agents, exhaustive evaluation is generally impractical because of the associated computational and financial costs, so these systems must directly generate a small candidate set from the full substitution space\. Single\-mutant discovery provides a natural starting point for protein engineering, but many desired phenotypes depend on combinations of substitutions that may reinforce, compensate for, or interfere with one another\. The remaining tasks therefore move from open\-ended single\-mutant generation to fixed\-pool ranking of multi\-mutant candidates with progressively richer levels of target\-specific experimental evidence made available\. #### 2\.1\.2T2: Measurement\-free multi\-mutant ranking Task componentSpecificationDefinitionRanking multi\-mutant candidates without target\-specific mutation measurements\.InputWild\-type sequencexx, assay contextcc, and a score\-blind candidate pool\.OutputComplete ranking of the supplied candidate pool\.InstanceInput: Wild\-type sequence:MLEGKVKW\.\.\.KEA Assay context: \- Primary task class: stability\. \- Fitness type: stability\. \- Readout subclass: folding free energy stability readout\. Candidate pool: \{V28C\+V63L,V28Q\+V63P,V28M\+V63F,\.\.\.\}\. Output: Ranked candidate set:V28C\+V63L,V28M\+V63F,V28Q\+V63P,\.\.\.\.StatisticsAssays: 74; candidate number: 6,288\.Table 2:Task specification forT2: Measurement\-free multi\-mutant ranking\.Task componentSpecificationDefinitionRanking strict successors of a measured anchor mutant\.InputWild\-type sequencexx, assay contextcc, a measured anchor mutantuuwith scoreyuy\_\{u\}, and a score\-blind successor pool𝒞\\mathcal\{C\}\.OutputComplete ranking of the supplied candidate pool\.InstanceInput: Wild\-type sequence:QVQLVQSG\.\.\.VSS Assay context: \- Primary task class: binding\. \- Fitness type: binding\. \- Readout subclass: binding\. Anchor:T28Pwith score \-1\.586\. Successor pool: \{T28P\+S30R\+N59K\+T76A,T28P\+S30R\+N59K\+Q62P\+S75F,T28P\+S30R\+L104V,\.\.\.\}\. Output: Ranked candidate set: T28P\+S30R\+N59K\+Q62P\+S75F,T28P\+S30R\+N59K\+T76A,T28P\+S30R\+L104V,\.\.\.\.StatisticsAssays: 67; candidate number: 3,648\.Table 3:Task specification forT3: Anchor\-informed multi\-mutant ranking\.TaskT2evaluates multi\-mutant prioritization as fixed\-pool ranking rather than open\-ended generation\. Given the wild\-type sequence, assay context and a score\-blind pool sampled from measured multi\-mutants, the model returns a permutation of that pool for downstream experimental evaluation\. [Table2](https://arxiv.org/html/2609.28921#S2.T2)summarizes the task specification\. Formally, given a wild\-type sequencexx, assay contextcc, and a score\-blind multi\-mutant candidate pool𝒞\\mathcal\{C\}, the model implements a ranking function: fθ:\(x,c,𝒞\)↦𝒪^,𝒪^∈Π\(𝒞\)\.f\_\{\\theta\}:\\bigl\(x,c,\\mathcal\{C\}\\bigr\)\\mapsto\\widehat\{\\mathcal\{O\}\},\\qquad\\widehat\{\\mathcal\{O\}\}\\in\\Pi\\bigl\(\\mathcal\{C\}\\bigr\)\. Here,Π\(𝒞\)\\Pi\(\\mathcal\{C\}\)denotes the set of all permutations of the candidate pool\. T2assesses measurement\-free combinatorial reasoning\. Specifically, the model must prioritize variants using the available protein information and assay description while accounting for possible interactions among substitutions\. It does not test whether a model can generate new combinations outside the supplied pool or assign calibrated fitness values to arbitrary mutants\. #### 2\.1\.3T3: Anchor\-informed multi\-mutant ranking TaskT3evaluates successor ranking conditional on one measured sequence background\. In the current release, 65 of 67 anchors are single substitutions; one anchor contains two substitutions and one contains three\. [Table3](https://arxiv.org/html/2609.28921#S2.T3)summarizes the task specification\. Formally, given a wild\-type sequencexx, assay contextcc, and a measured anchor mutantuuwith scoreyuy\_\{u\}, let𝒞\\mathcal\{C\}denote the supplied score\-blind successor pool\. Every candidatev∈𝒞v\\in\\mathcal\{C\}strictly contains the substitutions inuuand introduces at least one additional substitution relative to this measured background\. The model then implements a ranking function: fθ:\(x,c,u,yu,𝒞\)↦𝒪^,𝒪^∈Π\(𝒞\)\.f\_\{\\theta\}:\\bigl\(x,c,u,y\_\{u\},\\mathcal\{C\}\\bigr\)\\mapsto\\widehat\{\\mathcal\{O\}\},\\qquad\\widehat\{\\mathcal\{O\}\}\\in\\Pi\\bigl\(\\mathcal\{C\}\\bigr\)\. T3assesses anchor\-conditioned reasoning\. The measured anchor provides a local experimental reference; thus, the model must determine whether additional substitutions are likely to improve or reduce fitness relative to the measured sequence background\. This setting tests whether information from one experimentally characterized mutant can be transferred to the prioritization of related multi\-mutant candidates\. #### 2\.1\.4T4: Single\-mutant\-informed multi\-mutant ranking WhereasT3exposes one measured anchor,T4exposes a broader set of measured single\-mutant scores\. This context is not required to cover the complete19L19Lsingle\-substitution space\. Instead, the construction requirement is candidate\-specific completeness: every component substitution represented in a candidate has a corresponding measured single\-mutant entry\. [Table4](https://arxiv.org/html/2609.28921#S2.T4)summarizes the task specification\. Formally, given a wild\-type sequencexxand assay contextcc, let𝒮\\mathcal\{S\}map each provided measured single substitutionssto its assay scorey\(s\)y\(s\), and let𝒞\\mathcal\{C\}denote the supplied score\-blind multi\-mutant candidate pool, whose component substitutions are all represented in𝒮\\mathcal\{S\}\. The model then implements a ranking function: fθ:\(x,c,𝒮,𝒞\)↦𝒪^,𝒪^∈Π\(𝒞\)\.f\_\{\\theta\}:\\bigl\(x,c,\\mathcal\{S\},\\mathcal\{C\}\\bigr\)\\mapsto\\widehat\{\\mathcal\{O\}\},\\qquad\\widehat\{\\mathcal\{O\}\}\\in\\Pi\\bigl\(\\mathcal\{C\}\\bigr\)\. T4assesses component\-informed combination ranking\. Given the measured effects of all individual mutations, the model must prioritize multi\-mutant combinations\. Rather than simply adding individual scores, the system must account for non\-additive interactions to translate distributed evidence into accurate rankings\. Task componentSpecificationDefinitionMulti\-mutant ranking with measured context for every component substitution\.InputWild\-type sequencexx, assay contextcc, a measured single\-mutant context map𝒮\\mathcal\{S\}, and a score\-blind multi\-mutant candidate pool𝒞\\mathcal\{C\}\.OutputComplete ranking of the supplied candidate pool\.InstanceInput: Wild\-type sequence:EVKLDETG\.\.\.EIK Assay context: \- Primary task class: stability\. \- Fitness type: abundance or expression\. \- Readout subclass: cellular abundance stability proxy\. Single\-mutant context: \{W108E: \-0\.354,G109P: \-1\.062,M34Q: 0\.605,N35S: 1\.154,\.\.\.\}\. Candidate pool: \{W108E\+G109P,Y102P\+M105K,M34Q\+N35S,\.\.\.\}\. Output: Ranked candidate set:M34Q\+N35S,Y102P\+M105K,W108E\+G109P,\.\.\.\.StatisticsAssays: 29; candidate number: 2,638; visible single\-mutant context rows: 28,609\.Table 4:Task specification forT4: Single\-mutant\-informed multi\-mutant ranking\.ObjectiveMetricT1T2T3T4Global rankingSpearman Correlation×\\times✓\\checkmark✓\\checkmark✓\\checkmarkTop\-weighted rankingNDCG×\\times✓\\checkmark✓\\checkmark✓\\checkmarkPeak discoveryNormalized Maximum Score@KK✓\\checkmark✓\\checkmark✓\\checkmark✓\\checkmarkPeak coverageRecall@KK✓\\checkmark✓\\checkmark✓\\checkmark✓\\checkmarkTable 5:Evaluation objectives, corresponding metrics, and metric applicability across the four tasks\. Ranking metrics are not used forT1because it is a generation task instead of a ranking task\. ### 2\.2Evaluation Objectives [Table5](https://arxiv.org/html/2609.28921#S2.T5)summarizes the objective of each metric and its applicability to each task\. Implementation details of the corresponding metrics are provided in[Section3\.2](https://arxiv.org/html/2609.28921#S3.SS2)\. - •Global ranking with Spearman Correlation\.Measures whether the predicted ordering agrees with the experimental ordering across the complete candidate pool, evaluating a model’s ability to distinguish relative fitness throughout the library rather than only among the candidates placed at the top\. - •Top\-weighted ranking with NDCG\.Measures the quality of the upper part of the predicted list by assigning greater importance to candidates near the top, evaluating whether a model places the most promising variants in positions that are most useful for downstream experimental testing\. - •Peak discovery with Normalized Maximum Score@KK\.Measures the highest assay\-normalized fitness among theKKsubmitted candidates, evaluating whether a model can identify at least one strong candidate within the available budget ofKKproposed mutations\. - •Peak coverage with Recall@KK\.Measures how broadly a model recovers the ground\-truth top\-performing candidates within a prediction budgetKK, complementing peak discovery by assessing whether the model identifies multiple candidates from the high\-fitness region rather than only a single strong candidate\. ### 2\.3Data Construction #### 2\.3\.1Assay Collection and Harmonization We assembled assay\-level protein fitness landscapes from ProteinGym\[[11](https://arxiv.org/html/2609.28921#bib.bib12)\], MaveDB\[[13](https://arxiv.org/html/2609.28921#bib.bib13)\], MegaScale\[[14](https://arxiv.org/html/2609.28921#bib.bib34)\], FLAb\[[15](https://arxiv.org/html/2609.28921#bib.bib35)\], Human Domainome\[[16](https://arxiv.org/html/2609.28921#bib.bib36)\], and CombinGym\[[17](https://arxiv.org/html/2609.28921#bib.bib37)\]\. The collection also includes manually curated target DMS studies for CDKN2A\[[18](https://arxiv.org/html/2609.28921#bib.bib38)\], SLC13A5\[[19](https://arxiv.org/html/2609.28921#bib.bib39)\], and TrpB\[[20](https://arxiv.org/html/2609.28921#bib.bib40)\]\. These sources provide stability, binding, activity or organismal function measurements spanning single\-mutant and combinatorial sequence–fitness landscapes\. An assay is defined by its wild\-type construct, experimental condition, readout, and score semantics, so measurements for one protein remain separate assays when they represent distinct experimental settings or engineering objectives\. Each assay record contains biological and experimental metadata together with one or more assay\-linked A3M alignments generated with MMseqs2\[[21](https://arxiv.org/html/2609.28921#bib.bib23)\]against UniRef100\[[22](https://arxiv.org/html/2609.28921#bib.bib24)\]\. We curated an assay’s primary task class, fitness type and readout subclass based on DMS databases and papers, where primary task class is artificially set to 3 different types: stability, binding, or activity function\. We harmonized mutant representations and score semantics across sources\. A single substitution is encoded asA12V, and a multi\-mutant as a plus\-separated combination such asA12V\+G35L\. We retained mutants composed of the 20 standard amino acids and removed synonymous substitutions, stop codons, deletions, non\-finite scores, repeated mutation sites within a mutant, and mutations that could not be reconstructed from the assay\-specific wild\-type sequence\. Repeated measurements of the same canonical mutant within an assay were averaged\. Assay\-specific score semantics are retained for the score\-construction step below\. For each assay, we calculateDMS\_scorefrom the assay\-level effect values obtained from the source file or reconstructed from the source readout specified in the assay metadata\. Published processed effect columns are imported directly; assays requiring reconstruction are converted according to their measurement semantics\. For example, affinity values reported asKdK\_\{\\mathrm\{d\}\}are expressed as−log10\(Kd\)\-\\log\_\{10\}\(K\_\{\\mathrm\{d\}\}\), and expression ratios are expressed aslog10\(ER\)\\log\_\{10\}\(\\mathrm\{ER\}\)\. The resulting effect values are standardized within each assay using a population z\-score\. For readouts in which lower values indicate better performance, the corresponding effect values are sign\-reversed, so largerDMS\_scorevalues consistently indicate better assay performance\. The resulting assay\-level scores populate the measured single\-mutant tables and the candidate records used byT1–T4\. #### 2\.3\.2Quality Control Every constructed assay–task instance underwent three groups of quality\-control checks covering instance integrity, task consistency, and evaluation reliability: - •Instance integrity\.We first verified the correctness and completeness of each individual instance\. Checks included mutation–sequence consistency, finite fitness values, complete metadata, absence of duplicate mutants or rows, and valid links to the corresponding A3M alignments\. These checks ensure that each instance faithfully represents the underlying assay data, contains information required for downstream evaluation, and avoids malformed or incomplete records that could introduce artificial errors during model inference or metric computation throughout the evaluation pipeline\. - •Task consistency\.We then validated whether each assay was correctly instantiated under the corresponding task definition\. This included checking task\-specific inputs, candidate\-set construction, and mutation constraints, and task\-specific formatting and consistency requirements\. For assays shared across multiple task interfaces, we additionally cross\-checked sequence, measurement, and visual data to ensure that the same underlying assay information remained consistent across tasks and that no discrepancies were introduced during task\-specific preprocessing or instance conversion\. - •Evaluation reliability\.Finally, we examined whether individual instances supported stable and meaningful assessment across all reported evaluation metrics\. We excluded severe tie\-heavyT2–T4instances andT1instances in which exact top\-region ties made top\-KKmembership ambiguous\. Local near\-ties that did not alter the relevant ranking boundary were retained and documented\. We also tested for accidental order leakage by measuring monotonicity and rank correlation between row position andDMS\_score, ensuring that benchmark ordering did not provide unintended information about the target labels or create spurious performance gains unrelated to genuine mutant prioritization ability\. After quality filtering,PFArenacomprises 202 unique assays, 293 assay–task instances, and 607,269 target candidate rows\. Per\-task instance counts are listed in the task specification tables\. Since tasks impose different evidence and candidate constraints, the same assay may contribute to multiple tasks\. At the assay–task level, the current benchmark contains 124 stability, 101 activity or organismal function, and 68 binding instances\. Its source composition comprises 76 MegaScale, 78 ProteinGym, 44 MaveDB, 34 FLAb, 36 Human Domainome, 19 CombinGym and 6 manually curated DMS instances\. ## 3Experimental Setup ### 3\.1Baselines We compare four baseline families: random selection, zero\-shot protein language models, standalone general\-purpose LLMs, and LLM\-based scientific agents\. ##### Random ForT1, the random baseline uniformly samplesKKlegal single substitutions without replacement\. ForT2–T4, it produces a uniformly random permutation of the supplied candidate pool\. ##### Protein Language Models We evaluated six zero\-shot protein language models spanning complementary biological input modalities\. ESM\-2 \(650M\)\[[1](https://arxiv.org/html/2609.28921#bib.bib16)\]and ProGen2\-base \(764M\)\[[2](https://arxiv.org/html/2609.28921#bib.bib17)\]use protein sequences alone; ProSST\-2048\[[3](https://arxiv.org/html/2609.28921#bib.bib18)\]incorporates discrete structure tokens; S3F\[[4](https://arxiv.org/html/2609.28921#bib.bib19)\]integrates sequence, backbone, and surface representations; and VenusREM\[[5](https://arxiv.org/html/2609.28921#bib.bib21)\]additionally incorporates MSA\-derived evolutionary information\. S3F\-MSA combines S3F with an ensemble of five independently trained EVE models\[[6](https://arxiv.org/html/2609.28921#bib.bib20)\]\. ##### LLMs We evaluated six general\-purpose LLMs from diverse model families: GPT\-6 Astra\[[23](https://arxiv.org/html/2609.28921#bib.bib3)\], Claude Opus 5\[[8](https://arxiv.org/html/2609.28921#bib.bib5)\], Gemini 3\.1 Pro\[[24](https://arxiv.org/html/2609.28921#bib.bib6)\], Kimi K3\[[25](https://arxiv.org/html/2609.28921#bib.bib7)\], GLM\-5\.2\[[26](https://arxiv.org/html/2609.28921#bib.bib8)\], and DeepSeek\-V4\-Pro\[[27](https://arxiv.org/html/2609.28921#bib.bib9)\]\. All models were evaluated without any fine\-tuning\. For each task, they received the same prompt template and the same model\-visible inputs defined in[Section2\.1](https://arxiv.org/html/2609.28921#S2.SS1), and were required to return predictions in the prescribed structured format\. ##### Agents We additionally evaluated scientific agent frameworks Biomni\[[9](https://arxiv.org/html/2609.28921#bib.bib31)\]and ourAMixKittoolkit\. They received the same task\-visible inputs and followed the same structured\-output requirements as the standalone LLMs, while retaining the ability to orchestrate the tools and scientific resources provided by their frameworks\. ### 3\.2Metrics We compute all metrics independently for each assay\-level instance and then macro\-average the resulting values, so that every instance contributes equally to the reported performance\. After score harmonization during data construction, largerDMS\_scorevalues consistently indicate better fitness\. We use task\-specific values ofKKto account for differences in task formulation and candidate\-space size\. InT1, models generate candidates from a large single\-mutant space, so we setK=40K=40to provide a sufficiently broad generation budget\. InT2–T4, models rank a smaller supplied candidate pool, so we setK=5K=5to maintain a selective evaluation of the top\-ranked candidates within the reduced candidate space\. ##### Spearman Correlation Measures agreement between the predicted and experimental rankings over the complete candidate pool\. For an instance withnncandidates, we assign each candidate an experimental rank and a predicted rank, and compute the Pearson correlation between the two rank vectors\. Ties in the experimental scores are assigned their average rank\. ##### NDCG Evaluates ranking quality by placing greater weight on high\-fitness candidates near the top of the list\. We first apply min\-max normalization to theDMS\_scorevalues to map them into relevance scoresri=\(si−smin\)/\(smax−smin\)∈\[0,1\]r\_\{i\}=\(s\_\{i\}\-s\_\{\\min\}\)/\(s\_\{\\max\}\-s\_\{\\min\}\)\\in\[0,1\]\. For a predicted permutationπ\\piofnncandidates, NDCG is defined as: NDCG\(π\)=DCG\(π\)IDCG,DCG\(π\)=∑j=1n2rπ\(j\)−1log2\(j\+1\),\\operatorname\{NDCG\}\(\\pi\)=\\frac\{\\operatorname\{DCG\}\(\\pi\)\}\{\\operatorname\{IDCG\}\},\\qquad\\operatorname\{DCG\}\(\\pi\)=\\sum\_\{j=1\}^\{n\}\\frac\{2^\{r\_\{\\pi\(j\)\}\}\-1\}\{\\log\_\{2\}\(j\+1\)\},andIDCG\\operatorname\{IDCG\}represents the ideal maximum possibleDCG\\operatorname\{DCG\}, computed by evaluating the same discounted gain formula after sorting all candidates in decreasing order of true relevancerir\_\{i\}\. ##### Normalized Maximum Score Measures peak discovery among the submitted candidates\. LetℋK\\mathcal\{H\}\_\{K\}denote the legal candidates occurring within the firstKKsubmitted positions\. We define NMS@K=maxi∈ℋKsi−sminsmax−smin\.\\operatorname\{NMS\}@K=\\frac\{\\max\_\{i\\in\\mathcal\{H\}\_\{K\}\}s\_\{i\}\-s\_\{\\min\}\}\{s\_\{\\max\}\-s\_\{\\min\}\}\.ForT1,ℋ40\\mathcal\{H\}\_\{40\}is the submitted single\-mutant set under the fixed generation budget\. In this task,smins\_\{\\min\}andsmaxs\_\{\\max\}denote the minimum and maximum experimental scores over the full measured single\-mutant space\. ForT2–T4,ℋ5\\mathcal\{H\}\_\{5\}consists of the leading candidates in the predicted multi\-mutant ranking, andsmins\_\{\\min\}andsmaxs\_\{\\max\}are computed over the corresponding supplied candidate pool\. ##### Recall Measures coverage of the experimentally high\-fitness region\. Let𝒯K\\mathcal\{T\}\_\{K\}denote the ground\-truth top\-KKcandidates, with all candidates tied at the ground\-truth cutoff included as relevant\. For an instance with candidate\-pool sizenn, we compute Recall@KKas: Recall@K=\|ℋK∩𝒯K\|min\(K,n\)\.\\operatorname\{Recall\}@K=\\frac\{\|\\mathcal\{H\}\_\{K\}\\cap\\mathcal\{T\}\_\{K\}\|\}\{\\min\(K,n\)\}\. The denominator remainsmin\(K,n\)\\min\(K,n\)when ties expand the ground\-truth relevant set\. Recall complements NMS, which depends only on the strongest recovered candidate\. LLM and agent outputs are not fully controllable and may include refusals, incomplete responses, or outputs that cannot be parsed successfully\. The post\-processing procedures for missing or incomplete assay responses are described in[Section10\.1](https://arxiv.org/html/2609.28921#S10.SS1)and applied uniformly across all models\. ### 3\.3Implementation Details We used a common evaluation pipeline across all baseline families\. Task\-visible inputs and required output formats followed[Section2\.1](https://arxiv.org/html/2609.28921#S2.SS1), while the model\-specific inference procedures are described below\. ##### PLM Inference PLM inference used label\-free, zero\-shot mutation scores under fixed model\-specific configurations, without natural\-language assay descriptions or task\-specific auxiliary information\. Wild\-type structures were generated with AlphaFold 3\[[28](https://arxiv.org/html/2609.28921#bib.bib22)\], and evolutionary alignments were constructed against UniRef100\[[22](https://arxiv.org/html/2609.28921#bib.bib24)\]using MMseqs2\[[21](https://arxiv.org/html/2609.28921#bib.bib23)\]\. ForT1, the scores were used to rank all legal single substitutions; forT2–T4, they were used to rank the supplied candidate sets\. Model\-specific checkpoints, scoring definitions, and protocols for multi\-substitution, multi\-chain, and long\-sequence inputs are documented in[Section9\.1](https://arxiv.org/html/2609.28921#S9.SS1)\. ##### LLM Inference Each prompt was instantiated from a task\-specific template, provided in[Section8](https://arxiv.org/html/2609.28921#S8), and submitted as a single non\-streaming request through an asynchronous OpenAI\-compatible client\. The modelsgpt\-6\-astra,claude\-opus\-5,gemini\-3\.1\-pro\-preview,kimi\-k3,glm\-5\.2, anddeepseek\-v4\-prowere required to return a task\-specific structured output: a proposed mutation set forT1and a proposed ranking forT2–T4\. Default maximum completion lengths were set to 65,536 tokens forgemini\-3\.1\-pro\-preview,kimi\-k3andglm\-5\.2, and 32,768 tokens for the remaining models\. Sampling parameters and reasoning\-effort settings were left at their API defaults for all evaluated model configurations\. Responses that could be successfully parsed into the required structure were recorded as results, without repairing, replacing, or filtering predicted mutations before evaluation\. Failed instances were retried for up to 10 complete inference runs, until all pending instances succeeded or a run produced no additional successful results\. ##### Agent Inference We used Biomni\[[9](https://arxiv.org/html/2609.28921#bib.bib31)\]to instantiate each instance as an independent task with a task\-specific agent prompt\. We evaluated two backbone models:gpt\-6\-astraandclaude\-opus\-5\. Unlike standalone LLM inference, each agent operated in a multi\-turn loop with iterative tool use and intermediate feedback\. At each turn, the agent either emitted an<execute\>block to invoke a Biomni tool, inspect public database or model outputs, or run focused Python/Bash code, or emitted a final<solution\>block\. After each tool\-use turn, the observation was returned to the agent before the next turn\. We required at least one execution–observation round before accepting a final answer and allowed at most 40 tool\-use rounds\. Each model call was limited to a maximum of 24,576 output tokens\. Once the tool\-use budget was exhausted, further tool calls were disabled and the agent was required to return the final answer in the specified structured\-JSON format\. GPU\-intensive Biomni tools and eligible code\-execution blocks were dispatched to a remote GPU worker, while the agent remained responsible for selecting actions and producing the final ranking\. Only the structured JSON object within the<solution\>block was parsed as the final response\. Figure 3:Overview of AMixKit, integrating protein features, PLM inference, and mutant construction\. ##### AMixKit Toolset We additionally evaluated agents equipped withAMixKit, a lightweight toolkit that provides standardized protein\-specific operations for evidence retrieval, mutation construction, and PLM\-based prediction\. As illustrated in[Figure3](https://arxiv.org/html/2609.28921#S3.F3),AMixKitcontains seven tools: - •UniProt Features:retrieves functional sites, domains, structural annotations, natural variants, and curated mutagenesis records for a specified protein accession\. - •Assay Features:summarizes assay\-specific evolutionary or structural evidence, including MSA coverage and conservation or AlphaFold pLDDT confidence for an exact benchmark assay\. - •Centroid Distance:ranks residue positions by their distance to the structural centroid, with chain\-aware handling of multichain proteins, to identify structurally central or peripheral sites\. - •Amino Acid Retrieval:retrieves the wild\-type amino acid at specified sequence positions\. - •Build Mutant:constructs and validates single\-substitution annotations from selected positions and target amino acids while enforcing wild\-type sequence consistency\. - •PLM Mutant Suggestion:generates and ranks the 19 possible non\-wild\-type substitutions at selected positions using a specified PLM and reports the highest\-scoring candidates per position\. - •PLM Mutant Scoring:assigns PLM\-based fitness scores to a set of single\- or multi\-mutant candidates\. The PLM tools support ESM\-2, ProGen2, ProSST, S3F, S3F\-MSA, and VenusREM, using cached scores when available and the scoring service otherwise\. Collectively, these tools allow agents to gather biological evidence, construct valid mutations, and incorporate protein\-model predictions in one workflow\.AMixKitagents followed the same multi\-turn execution–observation protocol as the Biomni agents\. We evaluateAMixKitwith both frontier LLMs and our in\-house model, AMix\-2\.1\. Frontier LLMs demonstrate strong capabilities, but may still exhibit over\-refusal or false refusals on biology\-related tasks\. Moreover, their API costs can become substantial when deployed at scale\. Motivated by these limitations, we developed AMix\-2\.1 by scaling AMix\-2\[[29](https://arxiv.org/html/2609.28921#bib.bib1)\]to hundreds of billions of parameters and applying multi\-task agentic RL\. We will release further technical details and open\-source both AMix\-2\.1 andAMixKitin future work\. ## 4Main Results Figure 4:Task\-level and metric\-level rank\-based model capability\.For each task and metric, all methods are ranked according to their performance, withrank=1rank=1indicating the best method\.### 4\.1Overall Results We use[Figure4](https://arxiv.org/html/2609.28921#S4.F4)to summarize benchmark performance at two levels: at themetriclevel, each method is ranked based on each metric for fine\-grained comparison; at thetasklevel, metric\-specific ranks are aggregated into a task\-level mean rank for comprehensive performance\. The random baseline, PLMs, LLMs, and tool\-augmented agents are compared rank\-based, demonstrating how method performance changes with task formulation, available experimental evidence, and evaluation objective\. Two qualitative observations emerge: - •At the task level, relative strengths vary with task formulation and available evidence\.InT1, PLMs lead overall\. Agents lead inT2, whereasT3shows a more interleaved ordering between PLMs and agents\. InT4, agents achieve their clearest overall advantage under comprehensive target\-specific evidence\. This non\-monotonic pattern highlights task formulation alongside experimental evidence availability\. - •At the metric level, overall standing does not guarantee leadership on every metric\.InT1, PLMs hold the strongest overall positions, yet GPT\-6 Astra leads NMS@4040\. InT3, agents rank highest on global and top\-weighted ordering, whereas ProSST\-2048 ranks highest on peak discovery and top\-candidate recovery\. Even where agents lead at the task level, as inT2andT4, the leading agent configuration varies across the four evaluation metrics\. Task\-level mean ranks therefore summarize overall strength but cannot identify the best method for every evaluation objective\. CategoryModelNMS@4040Recall@4040StatisticRandom0\.77490\.77490\.01420\.0142PLMVenusREM0\.79320\.0427S3F\-MSA0\.78160\.78160\.03640\.0364ProSST\-20480\.78470\.78470\.0407S3F0\.78040\.78040\.03640\.0364ESM\-20\.78540\.78540\.03700\.0370ProGen2\-base0\.77090\.77090\.02740\.0274LLMGPT\-6 Astra0\.79660\.02540\.0254Claude Opus 50\.77290\.77290\.02050\.0205Gemini 3\.1 Pro0\.76200\.76200\.02110\.0211Kimi K30\.77390\.77390\.01990\.0199GLM\-5\.20\.77340\.77340\.01850\.0185DeepSeek\-V4\-Pro0\.77030\.77030\.01850\.0185AgentBiomni\+\+GPT\-6 Astra0\.76770\.76770\.03880\.0388Biomni\+\+Claude Opus 50\.76510\.76510\.03070\.0307AMixKit\+\+GPT\-6 Astra0\.78660\.78660\.03860\.0386AMixKit\+\+Claude Opus 50\.78580\.78580\.03540\.0354AMixKit\+\+AMix\-2\.10\.78720\.78720\.03680\.0368Table 6:Performance onT1: Single\-mutant generation\. The best results are inboldand the second\-best results areunderlined\. The same convention is used in the tables below\. ### 4\.2Task\-specific Results #### 4\.2\.1T1: Single\-mutant generation As shown in[Table6](https://arxiv.org/html/2609.28921#S4.T6), single\-mutant generation remains challenging under a fixed generation budget, with limited recovery of top\-ranked mutants and only modest gains in peak discovery over random selection\. Learned methods achieve higher Recall@4040than random selection, but the best value increases from only0\.01420\.0142to0\.04270\.0427, still recovering fewer than 5% of the ground\-truth top\-40 mutations\. The relative gains in NMS@4040are noticeably smaller, rising at most from0\.77490\.7749to0\.79660\.7966, with numerous models showing no competitive edge over random guessing in this setting\. Across model categories, all six PLMs achieve higher Recall@4040than the evaluated standalone LLMs\. PLMs also outperform standalone LLMs in NMS@4040, with the exception of ProGen2\-base and GPT\-6 Astra, whileAMixKitagents outperform all PLMs except VenusREM\. Taken together, PLMs remain competitive in measurement\-free mutant generation, while frontier LLMs andAMixKittool\-use supports identification of highly promising individual candidates\. CategoryModelSpearmanNDCGNMS@55Recall@55StatisticRandom−0\.0215\-0\.02150\.81960\.81960\.71130\.71130\.08110\.0811PLMVenusREM0\.28040\.28040\.86840\.86840\.80030\.80030\.21620\.2162S3F\-MSA0\.28030\.28030\.87000\.87000\.79640\.79640\.20270\.2027ProSST\-20480\.27770\.27770\.86570\.86570\.78290\.78290\.21080\.2108S3F0\.27850\.27850\.86260\.86260\.77510\.77510\.17570\.1757ESM\-20\.25390\.25390\.86340\.86340\.78310\.78310\.18920\.1892ProGen2\-base0\.14050\.14050\.85630\.85630\.78870\.78870\.17840\.1784LLMGPT\-6 Astra0\.27260\.27260\.86350\.86350\.79670\.79670\.2216Claude Opus 50\.28410\.28410\.86530\.86530\.80310\.80310\.19190\.1919Gemini 3\.1 Pro0\.22390\.22390\.85950\.85950\.77230\.77230\.19730\.1973Kimi K30\.19730\.19730\.85070\.85070\.74990\.74990\.19190\.1919GLM\-5\.20\.16070\.16070\.85190\.85190\.77060\.77060\.20000\.2000DeepSeek\-V4\-Pro0\.21870\.21870\.85870\.85870\.77600\.77600\.21350\.2135AgentBiomni\+\+GPT\-6 Astra0\.31940\.87670\.80380\.21620\.2162Biomni\+\+Claude Opus 50\.30640\.30640\.86890\.86890\.80160\.80160\.2351AMixKit\+\+GPT\-6 Astra0\.28510\.28510\.87020\.87020\.77810\.77810\.20270\.2027AMixKit\+\+Claude Opus 50\.33300\.87640\.81200\.2351AMixKit\+\+AMix\-2\.10\.29600\.29600\.87060\.87060\.79110\.79110\.20000\.2000 Table 7:Performance onT2: Measurement\-free multi\-mutant ranking\.CategoryModelSpearmanNDCGNMS@55Recall@55StatisticRandom0\.04110\.04110\.82160\.82160\.76590\.76590\.19700\.1970PLMVenusREM0\.36980\.36980\.88970\.88970\.85780\.3612S3F\-MSA0\.31590\.31590\.88600\.88600\.83820\.83820\.32240\.3224ProSST\-20480\.38360\.89030\.89030\.86200\.3672S3F0\.33570\.33570\.88420\.88420\.85230\.85230\.29550\.2955ESM\-20\.31280\.31280\.88400\.88400\.80990\.80990\.30450\.3045ProGen2\-base0\.18700\.18700\.86060\.86060\.78350\.78350\.27160\.2716LLMGPT\-6 Astra0\.35160\.35160\.88860\.88860\.82580\.82580\.31340\.3134Claude Opus 50\.35400\.35400\.88580\.88580\.82630\.82630\.30150\.3015Gemini 3\.1 Pro0\.28650\.28650\.88140\.88140\.80190\.80190\.28060\.2806Kimi K30\.31310\.31310\.88570\.88570\.81240\.81240\.29250\.2925GLM\-5\.20\.26760\.26760\.87010\.87010\.81490\.81490\.28660\.2866DeepSeek\-V4\-Pro0\.26270\.26270\.86980\.86980\.81060\.81060\.26870\.2687AgentBiomni\+\+GPT\-6 Astra0\.36870\.36870\.88280\.88280\.79250\.79250\.30450\.3045Biomni\+\+Claude Opus 50\.45110\.89850\.85110\.85110\.34630\.3463AMixKit\+\+GPT\-6 Astra0\.38330\.38330\.89070\.83230\.83230\.34330\.3433AMixKit\+\+Claude Opus 50\.36650\.36650\.88750\.88750\.85150\.85150\.34330\.3433AMixKit\+\+AMix\-2\.10\.36160\.36160\.89020\.89020\.85150\.85150\.32540\.3254 Table 8:Performance onT3: Anchor\-informed multi\-mutant ranking\. #### 4\.2\.2T2: Measurement\-free multi\-mutant ranking As shown in[Table8](https://arxiv.org/html/2609.28921#S4.T8), all learned measurement\-free ranking methods outperform the random baseline across all metrics, with the best Spearman correlation improving from−0\.0215\-0\.0215to0\.33300\.3330\. However, the best Recall@55reaches only0\.23510\.2351, recovering fewer than one quarter of the ground\-truth top\-five candidates\. While models can extract useful information regarding multi\-mutant fitness, they cannot reliably reconstruct the ordering or consistently recover the most promising candidates without target\-specific measurements\. Agents achieve the best score on all fourT2metrics\. AMixKit\+\+Claude Opus 5 achieves the highest Spearman correlation, NMS@55, and Recall@55, at0\.33300\.3330,0\.81200\.8120, and0\.23510\.2351, while Biomni\+\+GPT\-6 Astra achieves the highest NDCG at0\.87670\.8767\. The strongest non\-agent competitors are distributed across the LLM and PLM families: Claude Opus 5 is competitive in Spearman and NMS@55, whereas S3F\-MSA leads in NDCG\. #### 4\.2\.3T3: Anchor\-informed multi\-mutant ranking As shown in[Table8](https://arxiv.org/html/2609.28921#S4.T8), models of all categories perform better under the anchor\-informed setting, as compared to the measurement\-free setting\. All evaluated methods outperform random ranking across the four metrics, showing that anchor\-conditioned reasoning can improve multi\-mutant prioritization, while the effects of additional mutations remain difficult to infer reliably\. The leading method depends on the evaluation objective, revealing complementary strengths of agents and PLMs\. Biomni\+\+Claude Opus 5 achieves the highest Spearman correlation at0\.45110\.4511and NDCG at0\.89850\.8985; whereas ProSST\-2048 attains the highest NMS@55and Recall@55, with scores of0\.86200\.8620and0\.36720\.3672, respectively, and outperforms most other agents across metrics\. When an anchor measurement is available, no model family consistently dominates; therefore, the choice of model should be guided by specific experiment objectives\. CategoryModelSpearmanNDCGNMS@55Recall@55StatisticRandom−0\.0204\-0\.02040\.81240\.81240\.63710\.63710\.05520\.0552PLMVenusREM0\.38210\.38210\.88090\.88090\.82120\.82120\.25520\.2552S3F\-MSA0\.36620\.36620\.90050\.90050\.86860\.86860\.28970\.2897ProSST\-20480\.37790\.37790\.87250\.87250\.80780\.80780\.24830\.2483S3F0\.34870\.34870\.88360\.88360\.84100\.84100\.27590\.2759ESM\-20\.31530\.31530\.88280\.88280\.81290\.81290\.22760\.2276ProGen2\-base0\.25680\.25680\.87360\.87360\.83570\.83570\.17930\.1793LLMGPT\-6 Astra0\.59880\.59880\.92740\.92740\.89350\.89350\.44830\.4483Claude Opus 50\.58230\.58230\.92440\.92440\.87230\.87230\.42760\.4276Gemini 3\.1 Pro0\.53290\.53290\.90590\.90590\.85950\.85950\.40000\.4000Kimi K30\.47940\.47940\.89560\.89560\.83370\.83370\.35860\.3586GLM\-5\.20\.49590\.49590\.89840\.89840\.83380\.83380\.38620\.3862DeepSeek\-V4\-Pro0\.53080\.53080\.90590\.90590\.85950\.85950\.40000\.4000AgentBiomni\+\+GPT\-6 Astra0\.63730\.63730\.93470\.91450\.4828Biomni\+\+Claude Opus 50\.64460\.92930\.92930\.90150\.90150\.43450\.4345AMixKit\+\+GPT\-6 Astra0\.63820\.63820\.93360\.91310\.4897AMixKit\+\+Claude Opus 50\.64620\.93360\.90900\.90900\.46900\.4690AMixKit\+\+AMix\-2\.10\.55310\.55310\.91150\.91150\.86420\.86420\.38620\.3862Table 9:Performance onT4: Single\-mutant\-informed multi\-mutant ranking\. #### 4\.2\.4T4: Single\-mutant\-informed multi\-mutant ranking As shown in[Table9](https://arxiv.org/html/2609.28921#S4.T9), providing candidate\-complete single\-mutant fitness context shifts the strongest results toward LLM\-based systems, particularly tool\-augmented agents, which generally outperform PLMs across all metrics\. UnlikeT3, which provides the measured score of one anchor variant, this task supplies evidence for every candidate component, directly testing whether models can integrate distributed assay information\. Nevertheless, single\-mutant scores do not uniquely determine multi\-mutant fitness because non\-additive interactions may alter substitution effects in combined sequences\. Thus, this setting highlights both the value of component\-level experimental evidence and the challenge of extrapolating it to combinatorial variants\. Tool augmentation through both Biomni andAMixKitconsistently improves performance across metrics and backbones, while the best agent depends on the objective\. AMixKit\+\+Claude Opus 5 achieves the highest Spearman correlation at0\.64620\.6462; Biomni\+\+GPT\-6 Astra leads in NDCG at0\.93470\.9347and NMS@55at0\.91450\.9145; and AMixKit\+\+GPT\-6 Astra attains the highest Recall@55at0\.48970\.4897\. Tool\-augmented agents demonstrate the broadest advantage onT4, while the differing leaders for global ordering, top\-weighted ranking, peak discovery, and candidate coverage underscore that these metrics capture complementary ranking aspects\. ## 5Analysis ### 5\.1Comparing PLMs and LLMs As discussed in[Section4](https://arxiv.org/html/2609.28921#S4), PLMs and LLMs exhibit distinct performance trajectories across task environments\. Mechanistically, this divergence stems from their contrasting modes of information processing: PLMs rely primarily on sequential and homologous signals, whereas LLMs demonstrate superior capacity in leveraging contextual and textual auxiliary conditions\. From a biological perspective, the underlying driver of this disparity is task\-dependent, displaying a pronounced demarcation between the single\-mutant task ofT1and multi\-mutant tasks ofT2–T4\. Nevertheless, both model families encounter shared operational bottlenecks\. Specifically,[Section5\.1\.1](https://arxiv.org/html/2609.28921#S5.SS1.SSS1)investigates family\-specific preferences for amino\-acid mutation types;[Section5\.1\.2](https://arxiv.org/html/2609.28921#S5.SS1.SSS2)reveals that the utility of single\-mutant fitness values in LLMs and agentic systems is largely constrained to additive landscapes; and[Section5\.1\.3](https://arxiv.org/html/2609.28921#S5.SS1.SSS3)evaluates how search space complexity governs performance across architectures\. We systematically analyze each aspect in the subsequent sections\. #### 5\.1\.1Divergent Single\-mutant Proposal Strategies Figure 5:Directional six\-class amino\-acid transitions across model families\.The three panels present6×66\\times 6wild\-type\-to\-mutant class transition matrices for PLMs, LLMs, and agents\. The residue classes comprise nonpolar \(NP\), polar uncharged \(P0\), polar positively charged \(P\+\), polar negatively charged \(P\-\), aromatic \(Ar\), and Gly/Pro \(G/P\)\. Each cell denotes the percentage of successfulT1hits within the respective model\-family matrix \(normalized to100%100\\%\)\. Color intensity follows an exponential scale, with frequencies below0\.5%0\.5\\%depicted in white\. Charge\-separated states are preserved, maintaining distinct off\-diagonal transitions for acidic\-to\-basic and basic\-to\-acidic substitutions\.AsT1: Single\-mutant generationlacks target\-specific mutation measurements, models prioritize candidates using the supplied assay context and protein\-specific evolutionary information when available; agents may also incorporate tool\-derived evidence\. We evaluated each model by the proportion of top\-30%30\\%experimental hits within its top\-4040predictions, aggregating results by model family\. Mean success rates reached48\.0%48\.0\\%for PLMs,42\.5%42\.5\\%for LLMs, and50\.4%50\.4\\%for agents\. While overall predictive accuracy remains comparable across architectures, the underlying mutational strategies diverge substantially\. To profile candidate selections, we quantified directed transitions across six physicochemical residue classes among successful variants:nonpolar \(NP\),polar uncharged \(P0\),polar positively charged \(P\+\),polar negatively charged \(P\-\),aromatic \(Ar\), andGly/Pro \(G/P\)\. As shown in[Figure5](https://arxiv.org/html/2609.28921#S5.F5), the6×66\\times 6transition matrices capture all 36 wild\-type\-to\-mutant class shifts, normalized within each model family\. Distinguishing P\+ from P\- preserves critical charge\-reversal dynamics, whereas Gly and Pro are grouped together for compactness\. Class\-preserving substitutions along the main diagonal accounted for25\.8%25\.8\\%of successful PLM proposals, compared to53\.3%53\.3\\%for LLMs and36\.8%36\.8\\%for agents\. Successful LLM proposals were thus predominantly conservative, whereas PLM successes spanned a broader spectrum of class\-altering transitions, with LLM\-based agents occupying an intermediate regime\. This disparity is not an artifact of classification grouping, as intra\-G/P transitions contribute minimally \(0\.21%0\.21\\%,0\.06%0\.06\\%, and0\.18%0\.18\\%, respectively\)\. Practically, these patterns indicate that PLMs explore a wider substitution landscape, whereas LLM\-based architectures preferentially propose mutations that conserve fundamental physicochemical properties\. Both behaviors reflect plausible biological heuristics conditioned on available inputs\. Sequence\-based PLMs effectively identify substitutions compatible with evolutionary alignments and structural stability, which frequently transcend rigid residue classes\. The conservative substitutions observed for standalone LLMs are consistent with general biochemical heuristics, although these results do not identify the source of their internal priors\. AMixKit agents can additionally call PLMs through tools\. However, neither prior explicitly captures assay\-specific determinants \(e\.g\., active\-site geometry, interface kinetics, or conformational dynamics\) that ultimately dictate measured phenotypes\. Consequently, these trends reflect contrasting search strategies rather than an intrinsic hierarchy in model quality\. #### 5\.1\.2Utilization of Multi\-mutant Component Additivity Figure 6:Conditional utilization of single\-mutant evidence in multi\-mutant ranking\.Panels \(a\)–\(c\) show PLMs, LLMs and agents, respectively, with each point corresponding to oneT4assay\. The horizontal axis is the Spearman correlation between the sum of measured component single\-mutant fitness and the measured multi\-mutant fitness\. This is an information\-transfer diagnostic, not a direct estimate of a mechanistic epistasis coefficient\.Circlesmark assays withρsum,combination≥0\\rho\_\{\\mathrm\{sum,combination\}\}\\geq 0andsquaresmark assays below this operational threshold\. The line marks a linear fit, whileρ\\rhois the ranking correlation between the diagnostic and model performance\.For multi\-mutant ranking tasks, the experimental context supplied to the model increases progressively\. A key observation is that while PLMs, LLMs and agents exhibit comparable performance when single\-mutant evidence is limited, LLMs and agents gain a distinct advantage as component\-level measurements become available\. AcrossT2: Measurement\-free multi\-mutant rankingandT3: Anchor\-informed multi\-mutant ranking, the three model families achieve mean assay\-level Spearman correlations of0\.2710\.271vs\.0\.3310\.331\(PLMs\),0\.2230\.223vs\.0\.3110\.311\(LLMs\), and0\.3170\.317vs\.0\.3830\.383\(agents\)\. However, this performance gap widens dramatically inT4: Single\-mutant\-informed multi\-mutant ranking—where measured single\-mutant effects are explicitly provided—with LLMs and agents achieving0\.5350\.535and0\.6350\.635, substantially outperforming PLMs at0\.3490\.349\. This performance trajectory suggests that LLM\-based architectures leverage individual componentDMS\_scorevalues through an additive heuristic\. To test this hypothesis, we computed the sum of measured single\-mutantDMS\_scorevalues for each multi\-mutant candidate inT4and evaluated its correlation with actual combination fitness\. As depicted in[Figure6](https://arxiv.org/html/2609.28921#S5.F6), a high correlation indicates that combination rankings can be reliably approximated via additive components, whereas a low correlation denotes poor transferability\. The predictive accuracy of LLMs \(ρ=0\.733\\rho=0\.733,P=6\.12×10−6P=6\.12\\times 10^\{\-6\}\) and agents \(ρ=0\.649\\rho=0\.649,P=1\.39×10−4P=1\.39\\times 10^\{\-4\}\) strongly correlates with the applicability of this additive heuristic, whereas PLMs display a much weaker coupling \(ρ=0\.378\\rho=0\.378,P=0\.043P=0\.043\)\. On the non\-additive regime ofρsum,combination<0\\rho\_\{\\mathrm\{sum,combination\}\}<0, comprising 7 of 29 assays, LLM performance drops sharply to0\.0150\.015, compared to0\.2420\.242for PLMs and0\.3140\.314for agents\. Conversely, on the high\-transfer split, mean performance reaches0\.7010\.701for LLMs,0\.7370\.737for agents, but only0\.3840\.384for PLMs\. These findings reveal a conditional dependency on component evidence: LLMs and agents excel when combination fitness is well\-approximated by additive single\-mutant values, but LLMs collapse when this assumption fails\. Although agentic workflows provide partial robustness in non\-additive regimes, neither LLMs nor agents show evidence of modeling complex non\-additive residue interactions\. PLMs are less coupled to this diagnostic and therefore show smaller performance degradation, but not a general advantage on the low\-transfer assays\. Consequently, the dominance of LLM\-based systems inT4reflects effective utilization of transferable component evidence rather than a comprehensive understanding of residue coupling\. #### 5\.1\.3Common Limitations across Single\-mutant and Multi\-mutant Tasks Figure 7:Common limitations across model families\.\(a\)Sequence search space\. VerifiedT1top\-40 hit counts, stratified by the number of legal variants per assay into three bins \(≤\\leq2,000; 2,001–8,000;\>\>8,000\)\. Each violin shows the distribution of hit counts across assays for one family \(PLMs, LLMs or agents\), with a horizontal line marking the median; the three families are offset within each bin and share a common color coding\.\(b\)Mutation depth\. Mean assay\-level Spearman ranking accuracy for double mutants \(solid\) versus candidates with more than two substitutions \(hatched\), shown per family with the mean value labeled above each bar\. PLMs, LLMs and agents use the same family colors in both panels\.Whether through distinct strategies for selecting amino\-acid mutations, the application of single\-mutation fitness values, or the response mechanisms to additive effects, these differences ultimately reflect variations in the models’ capabilities across diverse tasks\. Yet all models face the same two dilemmas: sensitivity to sequence length, and a decline in prediction accuracy for variants of higher\-orders\. ##### High\-dimensional Search Spaces Navigating the search space of a protein sequence presents a severe challenge\. InT1: Single\-mutant generation, a single assay features a median of2,9642\{,\}964valid single\-point substitutions \(ranging from1,0931\{,\}093to22,53622\{,\}536\)\. However, each model proposes only its top4040candidates, thereby sampling a mere1\.35%1\.35\\%of the median search space\. As depicted in[Figure7\(a\)](https://arxiv.org/html/2609.28921#S5.F7), the average number of top\-40 hits in the single\-mutant scenario ofT1is only1\.211\.21, with even the highest\-performing case recovering only1313\. Furthermore, this recovery rate degrades rapidly with increasing sequence length and search space size\. This weak performance only marginally outperforms random sampling, and indicates that current models fall far short of truly mastering sequence space or providing reliable candidate fitness evaluations\. Extending this task to multi\-mutant proposal across full\-length sequences will exacerbate this challenge, as the combinatorial search space expands exponentially\. ##### High\-order Combinations The mutation depth within a variant \(i\.e\., the number of substituted sites per combination\) significantly impacts model ranking performance\. Across the multi\-mutant ranking tasks ofT2–T4, ranking accuracy declines monotonically as mutation count increases \(ρ=−0\.430\\rho=\-0\.430,P=4\.75×10−9P=4\.75\\times 10^\{\-9\}\), as shown in[Figure7\(b\)](https://arxiv.org/html/2609.28921#S5.F7)\. Specifically, inT4: Single\-mutant\-informed multi\-mutant ranking, the mean Spearman correlation for LLMs drops from0\.6690\.669on22mutations to0\.5140\.514for33–44mutations, and down to0\.2220\.222for variants with over 5 mutated sites, demonstrating superior efficacy on lower\-order combinations\. PLMs exhibit a parallel reduction from0\.3530\.353to0\.2360\.236and0\.0510\.051, while LLM\-based agents also show decay from0\.6740\.674to0\.5360\.536and0\.3860\.386\. Notably, even when augmented with single\-mutant fitness values—which substantially improve double\-mutant ranking capabilities—models fail to generalize to higher\-order variants involving complex epistatic interactions\. This pervasive performance drop highlights the fundamental difficulty of modeling high\-order combinations\. Together, these findings highlight how sequence dimension and mutation depth compound search pressure on computational models\. Given how heavily performance degrades even within the controlled settings ofT1–T4, unconstrained multi\-site generation—a common scenario in practical protein engineering—presents an exponentially higher barrier\. Addressing real\-world bio\-design problems requires robust predictive capability across high\-dimensional search spaces and higher\-order combinations\. This underscores the core motivation behind our task design: models must maintain reliable analytical and judgment capabilities when subjected to simultaneous scaling in sequence length and mutation depth\.PFArenareveals the shortcomings of existing models along these crucial dimensions\. ### 5\.2Dive into LLMs #### 5\.2\.1Feedback\-guided In\-context Adaptation Figure 8:Cumulative performance comparison of round\-wise single mutant generation\.Single\-roundrefers to the originalT1: Single\-mutant generationsetting, where all 40 mutations are proposed in a single pass, and is indicated with dotted horizontal lines\.Multi\-roundrefers to the results after a complete four\-round feedback\-guided generation protocol, with 10 mutations proposed each round, where proposed mutants with an assay\-recordedDMS\_scoreare used as scored context for the next model pass, and is depicted in solid lines\.To emulate a sequential wet\-laboratory campaign, in which mutation fitness is measured in batches and later candidates are selected using earlier outcomes, we extended the single\-roundT1: Single\-mutant generationinto a four\-round adaptive search\. This experiment tests whether assay\-specific fitness measurements supplied in context help models refine subsequent generations and identify higher\-fitness mutants within a fixed screening budget\. In each round, the model proposed 10 single mutants unselected in earlier rounds\. For the next round, the ground\-truthDMS\_scorefitness of all previously proposed mutants was added to the prompt as in\-context feedback, if the fitness value of the mutant is available\. Both settings therefore used the same total budget of 40 proposed mutations, but only the multi\-round setting allowed later generations to condition on intermediate fitness measurements\. As shown in[Figure8](https://arxiv.org/html/2609.28921#S5.F8), the multi\-round protocol shows monotonic improvements in both NMS@4040and Recall@4040across rounds for every model\. Relative to single\-round generation, the final Round\-4 endpoint NMS@4040increased from0\.79660\.7966to0\.83670\.8367for GPT\-6 Astra, highest across all models, and increased by0\.01680\.0168for Claude Opus 5,0\.02440\.0244for Gemini 3\.1 Pro, and0\.00760\.0076for DeepSeek\-V4\-Pro\. Recall@4040increased by0\.02110\.0211,0\.01750\.0175,0\.01280\.0128, and0\.00870\.0087respectively, with GPT\-6 Astra achieving the highest recovery of0\.04650\.0465\. Averaged across the four models, multi\-round feedback\-guided generation improved NMS@4040by0\.02220\.0222and Recall@4040by0\.01500\.0150\. All Round\-4 results exceeded the random baseline on both metrics, whereas three of the four single\-round models remained below random sampling in NMS@4040\. Across rounds, GPT\-6 Astra exceeded its single\-round NMS@4040baseline after two rounds, Claude Opus 5 and Gemini 3\.1 Pro after three rounds, and DeepSeek\-V4\-Pro after four rounds\. The first three models exceeded their single\-round baselines in Recall@4040after three rounds, whereas DeepSeek\-V4\-Pro after the fourth\. As the cumulative candidate budget increases across rounds, the earlier rounds alone do not establish the value of feedback; the decisive comparison is the Round\-4 endpoint against single\-round generation\. The consistent performance gains demonstrate the effectiveness of feedback\-guided in\-context adaptation, particularly for GPT\-6 Astra, which substantially surpassed its already strong single\-round baseline, highlighting its capability in knowledge\-based work tasks\. Notably, LLMs frequently revisited previously high\-scoring sites with alternative amino\-acid mutations, indicating that further exploration of promising positions in subsequent rounds may help identify mutants of higher fitness\. #### 5\.2\.2Uncertainty Estimation Figure 9:Confidence distribution and quality alignment\.\(a\)GPT\-6 Astra’s self\-reported confidence acrossT1–T4\.\(b\)Within\-setting Spearman correlations between confidence and ranking\-quality metrics\.We evaluated self\-reported confidence from GPT\-6 Astra to examine whether LLMs can estimate their own predictive uncertainty\. For queries inT1–T4, the model was prompted to assign an overall confidence score from00to100100reflecting its confidence in identifying and prioritizing high\-fitness mutations\. Following five API policy refusals and two safety exclusions inT1, confidence was available for116116,7474,6767, and2929queries inT1–T4, respectively\. We treated this score as an inverse measure of relative query\-level uncertainty and assessed its association with predictive performance using Spearman correlation\. The confidence elicitation prompt used in our experiments is provided in Appendix[8\.3](https://arxiv.org/html/2609.28921#S8.SS3)\. As shown in[Figure9\(a\)](https://arxiv.org/html/2609.28921#S5.F9), confidence varied substantially across settings, with mean scores of22\.522\.5,35\.635\.6,38\.338\.3, and63\.063\.0forT1–T4, respectively, and no response reaching9090\.T1: Single\-mutant generationelicited relatively conservative judgments, whereasT4: Single\-mutant\-informed multi\-mutant rankingproduced a relatively high\-confidence distribution\. These absolute values should not be interpreted as success probabilities or directly compared across settings, which differ in data and evaluation objectives\. The more meaningful test is whether confidence distinguishes higher\- from lower\-quality predictions within each setting\. Self\-reported confidence provides a partial, task\- and context\-dependent estimate of ranking reliability\.[Figure9\(b\)](https://arxiv.org/html/2609.28921#S5.F9)shows positive evidence inT3: Anchor\-informed multi\-mutant ranking, where confidence aligned with all four quality dimensions \(ρ≈0\.32\\rho\\approx 0\.32–0\.490\.49\) after multiple\-testing correction\. InT2: Measurement\-free multi\-mutant ranking, confidence also aligned with all four dimensions \(ρ≈0\.28\\rho\\approx 0\.28–0\.570\.57\), with stronger associations for global ordering than top\-five selection\.T1showed weak\-to\-moderate associations with its top\-40 metrics \(ρ=0\.35\\rho=0\.35for NMS@4040and0\.230\.23for Recall@4040\), which remained significant after filtering dirty outputs\.T4showed the largest correlations for full\-list Spearman and NDCG quality \(ρ=0\.69\\rho=0\.69and0\.670\.67\), whereas neither top\-five metric passed multiple\-testing correction\. Assay\-level examples reinforce this limitation: one protein\-stability query inT4received high confidence \(7878\) despite limited top\-five recovery \(Recall@55=0\.20\.2\), whereas a fluorescence query received lower confidence \(4545\) but substantially better recovery \(Recall@55=0\.80\.8\)\. These cases show that confidence does not uniformly track top\-candidate recovery, even when it aligns with global ranking quality\. It can therefore support coarse query triage within a fixed task, but should not be treated as a generally reliable or cross\-assay uncertainty estimate\. Figure 10:Test\-time scaling performance across sample counts\.The panels show NMS@4040\(left\) and Recall@4040\(right\) as the number of sampled mutation lists increases, with the dotted line representing the sample average\.SettingMethodNMS@4040↑\\uparrowRecall@4040↑\\uparrowInvalid per Assay↓\\downarrowVanillaSample Avg0\.76960\.76960\.01860\.01862\.28052\.2805RRF0\.78470\.01991\.4472CleanSample Avg0\.77710\.77710\.01890\.01891\.0531RRF0\.78450\.01971\.10571\.1057In\-AssaySample Avg0\.77450\.77450\.01720\.01720\.00000\.0000RRF0\.78420\.01910\.00000\.0000Table 10:Validity\-controlled RRF analysis atn=16n=16\. Clean requires 40 unique and sequence\-valid mutations\. In\-Assay further requires all mutations to appear in the assay’s DMS table\. #### 5\.2\.3Test\-time Scaling To test whether additional inference\-time computation improves single\-mutation generation without experimental feedback, we independently sampled 16 responses from DeepSeek\-V4\-Pro using exactly the same input for eachT1: Single\-mutant generationquery\. We extracted a ranked top\-40 mutation list from each response and applied aggregation to nested prefixes withn∈\{1,2,4,8,16\}n\\in\\\{1,2,4,8,16\\\}\. Aggregation used no ground\-truthDMS\_score, candidate labels, or evaluator metrics, and the mean performance of the corresponding raw samples at each prefix size served as the Sample Avg baseline\. We compared five TTS aggregation strategies\.Best\-of\-Nselected the complete sampled ranking most consistent with the other samples\.Approval Votingprioritized mutations that appeared repeatedly across samples, whereasBorda FusionandReciprocal Rank Fusion \(RRF\)additionally incorporated within\-sample rank using linear and reciprocal weighting, respectively\.Hierarchical RRFfurther pooled evidence across different mutations at the same residue position\. In the main scaling experiments, all methods operated directly on the extracted candidate strings without explicit sequence\-validity or DMS\-membership filtering\. Complete method details and hyperparameters used in our experiments are provided in[Section9\.2](https://arxiv.org/html/2609.28921#S9.SS2)\. As shown in[Figure10](https://arxiv.org/html/2609.28921#S5.F10), increasing sample count generally improved performance\. Approval scaled most consistently, reaching0\.78760\.7876NMS@4040and0\.02220\.0222Recall@4040atn=16n=16, versus0\.76960\.7696and0\.01860\.0186for Sample Avg\. Borda and RRF also improved with more samples, while Best\-of\-N was less stable and Hierarchical RRF mainly benefited smallnn\. All five methods outperformed Sample Avg atn=16n=16\. One source of improvement was invalid\-candidate suppression\. LLM\-proposed mutations may belocally invaliddue to sequence inconsistency or formatting errors, orassay invalidwhen they do not appear in the assay’s DMS table\. In the main setting, TTS aggregation substantially suppressed invalid outputs\. As shown in[Table10](https://arxiv.org/html/2609.28921#S5.T10), RRF reduced invalid predictions from2\.28052\.2805to1\.44721\.4472, without explicit validity filtering\. To determine whether this fully explains the gain, we further evaluatedCleanandIn\-Assaysettings\. Clean retained only samples with 40 unique and locally valid candidates, while In\-Assay further required all candidates to appear in the assay’s DMS table\. UnderClean, RRF improved NMS from0\.77710\.7771to0\.78450\.7845and recall from0\.01890\.0189to0\.01970\.0197, despite slightly more assay\-absent predictions \(1\.05311\.0531to1\.10571\.1057\)\. UnderIn\-Assay, where both methods had zero invalid predictions, RRF still improved NMS from0\.77450\.7745to0\.78420\.7842and recall from0\.01720\.0172to0\.01910\.0191\. Thus, although invalid\-candidate suppression significantly contributes to the TTS gain, it does not fully account for the improvement; even when validity is controlled, aggregation helps identify and prioritize higher\-quality valid candidates across repeated samples from the same query\. Despite the gains from TTS at larger sample counts, a substantial oracle gap remains\. Atn=16n=16, best\-sample oracles reached0\.84740\.8474NMS@4040and0\.04610\.0461Recall@4040, compared with0\.78760\.7876and0\.02220\.0222for the best practical TTS method\. This gap suggests that repeated sampling can already produce substantially better predictions, but reliably identifying them without experimental feedback remains challenging\. ## 6Related Work ##### Protein Fitness Benchmarks Protein fitness benchmarks have established standardized evaluation of computational models for predicting experimentally measured mutation effects\. FLIP\[[12](https://arxiv.org/html/2609.28921#bib.bib10)\]evaluates fitness\-landscape inference under low\-resource and extrapolative generalization settings, while FLIP2\[[30](https://arxiv.org/html/2609.28921#bib.bib11)\]extends this framework to unseen mutation identities, positions, mutation counts, wild\-type proteins, and high\-fitness regions\. ProteinGym\[[11](https://arxiv.org/html/2609.28921#bib.bib12)\]further enables large\-scale evaluation of zero\-shot and supervised mutation\-effect predictors across diverse DMS assays and clinical labels\. Together, these benchmarks characterize how accurately PLMs score and rank predefined mutants across heterogeneous fitness landscapes\. Recent benchmarks have extended protein fitness evaluation to general\-purpose LLMs and LLM\-based agents\. ProteinGym\-LLM\[[31](https://arxiv.org/html/2609.28921#bib.bib14)\]evaluates whether LLMs can rank a fixed set of protein mutants from the wild\-type sequence and assay description, providing a direct assessment of LLM\-based mutation prioritization\. BioDesignBench\[[32](https://arxiv.org/html/2609.28921#bib.bib15)\]evaluates tool\-using LLM agents across expert\-curated protein\-design workflows, with an emphasis on candidate generation, tool use, and iterative evaluation\. These studies extend protein engineering benchmarks from specialized fitness predictors to general\-purpose reasoning and agentic design systems\. Despite these advances, existing benchmarks typically focus on either fixed\-list mutant ranking or end\-to\-end protein\-design workflows, leaving systematic comparison across different mutation\-discovery settings underexplored\.PFArenaaddresses this gap by evaluating PLMs, general\-purpose LLMs, and LLM\-based agents through one single\-mutant generation task \(T1\) and three multi\-mutant ranking tasks under measurement\-free \(T2\), anchor\-informed \(T3\), and single\-mutant\-informed \(T4\) conditions, applying consistent task interfaces and evaluation criteria across all model families\. ##### Protein Fitness Methods Target\-specific supervised methods provide the most direct approach to protein fitness prediction by learning sequence–fitness relationships from experimentally measured mutants\. The Low\-NNframework\[[33](https://arxiv.org/html/2609.28921#bib.bib25)\]uses pretrained UniRep representations and as few as 24 assayed mutants to guide protein engineering, while Hsu et al\.\[[34](https://arxiv.org/html/2609.28921#bib.bib26)\]combine site\-specific amino\-acid features with evolutionary density estimates to predict fitness from limited measurements\. Because these methods are trained against the phenotype of interest, they can directly adapt to the assay\-specific objective, but their performance depends on the number, diversity, and sequence\-space coverage of available labels\. Zero\-shot PLMs avoid target\-specific training by deriving mutation scores from general biological priors learned from sequence, evolutionary, and structural data\. Evolutionary models such as EVmutation\[[35](https://arxiv.org/html/2609.28921#bib.bib27)\], DeepSequence\[[36](https://arxiv.org/html/2609.28921#bib.bib28)\], and EVE\[[6](https://arxiv.org/html/2609.28921#bib.bib20)\]infer family\-specific constraints from homologous sequences\. Protein language models, including ESM\-1v\[[37](https://arxiv.org/html/2609.28921#bib.bib29)\], ESM\-2\[[1](https://arxiv.org/html/2609.28921#bib.bib16)\], ProGen2\[[2](https://arxiv.org/html/2609.28921#bib.bib17)\], Tranception\[[38](https://arxiv.org/html/2609.28921#bib.bib30)\], and DPLM\-Evo\[[39](https://arxiv.org/html/2609.28921#bib.bib2)\]instead learn sequence compatibility across large protein corpora\. Structure\-aware models such as ProSST\[[3](https://arxiv.org/html/2609.28921#bib.bib18)\], S3F\[[4](https://arxiv.org/html/2609.28921#bib.bib19)\], and VenusREM\[[5](https://arxiv.org/html/2609.28921#bib.bib21)\]further incorporate geometric and evolutionary information\. These models can score mutations without assay\-specific labels, but their scores primarily reflect general protein plausibility rather than the phenotype and experimental evidence associated with a particular engineering objective\. General\-purpose LLMs and scientific agents have also shown substantial potential for protein mutation discovery\. Successive Claude releases have demonstrated continued progress in protein mutation\-effect prediction, with Claude Opus 5 further improving upon previous generations on ProteinGym Hard\[[8](https://arxiv.org/html/2609.28921#bib.bib5)\]\. LLMs have also been incorporated into budget\-constrained protein optimization procedures\[[40](https://arxiv.org/html/2609.28921#bib.bib32)\], suggesting that their biological knowledge can support mutation selection beyond direct fitness scoring\. Scientific agents further extend these capabilities through retrieval and tool use: ProtAgents\[[41](https://arxiv.org/html/2609.28921#bib.bib33)\]coordinates specialized agents for protein analysis and design, while Biomni\[[9](https://arxiv.org/html/2609.28921#bib.bib31)\]integrates biomedical databases, scientific software, and code execution\. These developments motivate systematic evaluation of whether LLMs and agents can convert assay context, experimental evidence, and specialized protein\-model outputs into effective mutation priorities\. Given the rapidly expanding set of PLMs, LLMs, and scientific agents, our evaluation focuses on representative methods spanning sequence\-based, structure\-based, evolution\-based, and tool\-based paradigms\. ## 7Conclusion PFArenaprovides a unified evaluation of PLMs, LLMs, and LLM\-based agents across realistic protein modification decisions\. Taken together, our results show that the preferred model family depends on the task interface\. Protein\-specific representations are particularly effective for open\-ended single\-mutant search, whereas general\-purpose reasoning and tool use become more useful for ranking mutation combinations, especially when measured single\-mutant context is available\. This complementarity does not eliminate shared bottlenecks: high\-fitness candidate recovery remains sparse under limited prediction budgets, and ranking becomes substantially less reliable beyond double mutants\. Our additional analyses show that feedback\-guided adaptation and test\-time scaling can improve LLM\-based inference\. We release our code and benchmark suite to support reproducible progress on these open problems\. ##### Limitations Despite its broad coverage,PFArenahas several limitations\. - •Potential label leakage\.All assays inPFArenaare derived from publicly available datasets and publications\. Their sequences, mutation labels, or experimental results may therefore have appeared in the pretraining corpora of the evaluated LLMs, making it impossible to fully exclude memorization or other forms of data contamination that could influence the evaluation results\. - •Stochasticity and single\-run evaluation\.Because evaluating frontier LLMs and tool\-augmented agents is computationally and financially expensive, each LLM or agent configuration was run only once for main evaluation\. The reported results may therefore depend on sampling randomness\. - •Missing outputs, refusals, and comparison fairness\.Some LLMs, particularly proprietary systems such as GPT and Claude, refuse to answer a subset of assays because of their safety policies\. To retain a fixed evaluation set and avoid selectively excluding difficult cases, we replace such outputs with a deterministic random baseline result, as detailed in[Section10\.1](https://arxiv.org/html/2609.28921#S10.SS1)\. This protocol measures the end\-to\-end usability of each deployed system, but it also conflates underlying protein\-reasoning ability with provider\-specific refusal policies and output reliability\. - •Retrospective rather than prospective evaluation\.PFArenais constructed from previously measured assays\. Improvements in benchmark metrics consequently do not establish higher prospective wet\-lab hit rates or guarantee that the selected mutations will succeed in a new experimental campaign\. Prospective validation will be required to determine whether the observed performance gains translate into practical reductions in experimental cost and design cycles\. - •Dataset coverage and heterogeneity\.The assays vary in experimental protocol, measurement noise, sequence coverage, candidate\-library construction, and phenotype definition\. Harmonization and quality control reduce but cannot eliminate these differences\. Moreover, publicly available datasets may overrepresent well\-studied proteins, and assay types that are convenient to measure, while underrepresenting negative results, rare protein families, and complex cellular or organism\-level phenotypes\. The conclusions may therefore not generalize to all protein\-engineering settings\. ## Contributions Project Lead Yawen Ouyang1,2 Co\-first Authors Yawen Ouyang1,2, Xinbo Zhang1,2, Ziyuan Ma1,4 Task Design and Data Collection Ziyuan Ma1,4, Xinbo Zhang1,2, Yawen Ouyang1,2 Main Results Yawen Ouyang1,2, Yixin Wu1,2, Wenbin Liao1,5, Xinbo Zhang1,2 Analysis Ziyuan Ma1,4, Yixin Wu1,2, Wenjie Li1, Feiran Zhang1,6, Yawen Ouyang1,2 Other Contributors Lihao Wang1,2, Hao Wang2,3, Xiaoqing Zheng6, Xuefeng Yan5, Lei Bai1, Ya\-Qin Zhang3, Shuyi Zhang4, Wei\-Ying Ma3,7, Dahua Lin1, Bowen Zhou1 Correspondence Hao Zhou1,2,3 ### Affiliation 1Shanghai Artificial Intelligence Laboratory 2Generative Symbolic Intelligence Lab \(GenSI\), Tsinghua University 3Institute for AI Industry Research \(AIR\), Tsinghua University 4School of Pharmaceutical Sciences, Tsinghua University 5School of Information Science and Engineering, East China University of Science and Technology 6College of Computer Science and Artificial Intelligence, Fudan University 7City University of Hong Kong ## Acknowledgments This work is supported by Shanghai Artificial Intelligence Laboratory and NSFC \(Grant No\. 62406170\)\. ## References - \[1\]Z\. Lin, H\. Akin, R\. Rao, B\. Hie, Z\. Zhu, W\. Lu, N\. Smetanin, R\. Verkuil, O\. Kabeli, Y\. Shmueli, A\. dos Santos Costa, M\. Fazel\-Zarandi, T\. Sercu, S\. Candido, and A\. Rives\(2023\)Evolutionary\-scale prediction of atomic\-level protein structure with a language model\.Science379\(6637\),pp\. 1123–1130\.External Links:[Document](https://dx.doi.org/10.1126/science.ade2574),[Link](https://doi.org/10.1126/science.ade2574)Cited by:[§1](https://arxiv.org/html/2609.28921#S1.p2.1),[§3\.1](https://arxiv.org/html/2609.28921#S3.SS1.SSS0.Px2.p1.1),[§6](https://arxiv.org/html/2609.28921#S6.SS0.SSS0.Px2.p2.1),[Table 11](https://arxiv.org/html/2609.28921#S9.T11.3.2.1.1)\. - \[2\]\(2023\)ProGen2: exploring the boundaries of protein language models\.Cell Systems14\(11\),pp\. 968–978\.e3\.External Links:[Document](https://dx.doi.org/10.1016/j.cels.2023.10.002),[Link](https://doi.org/10.1016/j.cels.2023.10.002)Cited by:[§1](https://arxiv.org/html/2609.28921#S1.p2.1),[§3\.1](https://arxiv.org/html/2609.28921#S3.SS1.SSS0.Px2.p1.1),[§6](https://arxiv.org/html/2609.28921#S6.SS0.SSS0.Px2.p2.1),[Table 11](https://arxiv.org/html/2609.28921#S9.T11.3.3.1.1)\. - \[3\]M\. Li, Y\. Tan, X\. Ma, B\. Zhong, H\. Yu, Z\. Zhou, W\. Ouyang, B\. Zhou, P\. Tan, and L\. Hong\(2024\)ProSST: protein language modeling with quantized structure and disentangled attention\.InAdvances in Neural Information Processing Systems,A\. Globerson, L\. Mackey, D\. Belgrave, A\. Fan, U\. Paquet, J\. Tomczak, and C\. Zhang \(Eds\.\),Vol\.37,pp\. 35700–35726\.External Links:[Document](https://dx.doi.org/10.52202/079017-1126),[Link](https://proceedings.neurips.cc/paper_files/paper/2024/file/3ed57b293db0aab7cc30c44f45262348-Paper-Conference.pdf)Cited by:[§1](https://arxiv.org/html/2609.28921#S1.p2.1),[§3\.1](https://arxiv.org/html/2609.28921#S3.SS1.SSS0.Px2.p1.1),[§6](https://arxiv.org/html/2609.28921#S6.SS0.SSS0.Px2.p2.1),[Table 11](https://arxiv.org/html/2609.28921#S9.T11.3.4.1.1)\. - \[4\]Z\. Zhang, P\. Notin, Y\. Huang, A\. Lozano, V\. Chenthamarakshan, D\. Marks, P\. Das, and J\. Tang\(2024\)Multi\-scale representation learning for protein fitness prediction\.InAdvances in Neural Information Processing Systems,A\. Globerson, L\. Mackey, D\. Belgrave, A\. Fan, U\. Paquet, J\. Tomczak, and C\. Zhang \(Eds\.\),Vol\.37,pp\. 101456–101473\.External Links:[Document](https://dx.doi.org/10.52202/079017-3217),[Link](https://proceedings.neurips.cc/paper_files/paper/2024/file/b7d795e655c1463d7299688d489e8ef4-Paper-Conference.pdf)Cited by:[§1](https://arxiv.org/html/2609.28921#S1.p2.1),[§3\.1](https://arxiv.org/html/2609.28921#S3.SS1.SSS0.Px2.p1.1),[§6](https://arxiv.org/html/2609.28921#S6.SS0.SSS0.Px2.p2.1),[Table 11](https://arxiv.org/html/2609.28921#S9.T11.3.5.1.1),[Table 11](https://arxiv.org/html/2609.28921#S9.T11.3.7.1.1)\. - \[5\]Y\. Tan, R\. Wang, B\. Wu, L\. Hong, and B\. Zhou\(2025\)From high\-throughput evaluation to wet\-lab studies: advancing mutation effect prediction with a retrieval\-enhanced model\.Bioinformatics41\(Supplement\_1\),pp\. i401–i409\.External Links:[Document](https://dx.doi.org/10.1093/bioinformatics/btaf189),[Link](https://doi.org/10.1093/bioinformatics/btaf189)Cited by:[§1](https://arxiv.org/html/2609.28921#S1.p2.1),[§3\.1](https://arxiv.org/html/2609.28921#S3.SS1.SSS0.Px2.p1.1),[§6](https://arxiv.org/html/2609.28921#S6.SS0.SSS0.Px2.p2.1),[Table 11](https://arxiv.org/html/2609.28921#S9.T11.3.6.1.1)\. - \[6\]J\. Frazer, P\. Notin, M\. Dias, A\. Gomez, J\. K\. Min, K\. Brock, Y\. Gal, and D\. S\. Marks\(2021\)Disease variant prediction with deep generative models of evolutionary data\.Nature599\(7883\),pp\. 91–95\.External Links:[Document](https://dx.doi.org/10.1038/s41586-021-04043-8),[Link](https://doi.org/10.1038/s41586-021-04043-8)Cited by:[§1](https://arxiv.org/html/2609.28921#S1.p2.1),[§3\.1](https://arxiv.org/html/2609.28921#S3.SS1.SSS0.Px2.p1.1),[§6](https://arxiv.org/html/2609.28921#S6.SS0.SSS0.Px2.p2.1),[Table 11](https://arxiv.org/html/2609.28921#S9.T11.3.7.1.1)\. - \[7\]Anthropic\(2026\)Claude opus 4\.8 system card\.Note:[https://www\-cdn\.anthropic\.com/0b4915911bb0d19eca5b5ee635c80fef830a37ea\.pdf](https://www-cdn.anthropic.com/0b4915911bb0d19eca5b5ee635c80fef830a37ea.pdf)Cited by:[§1](https://arxiv.org/html/2609.28921#S1.p2.1)\. - \[8\]Anthropic\(2026\)Introducing claude opus 5\.Note:[https://www\.anthropic\.com/news/claude\-opus\-5](https://www.anthropic.com/news/claude-opus-5)Cited by:[§1](https://arxiv.org/html/2609.28921#S1.p2.1),[§3\.1](https://arxiv.org/html/2609.28921#S3.SS1.SSS0.Px3.p1.1),[§6](https://arxiv.org/html/2609.28921#S6.SS0.SSS0.Px2.p3.1)\. - \[9\]K\. Huang, S\. Zhang, H\. Wang, Y\. Qu, Y\. Lu, Y\. Roohani, R\. Li, L\. Qiu, J\. Zhang, Y\. Di,et al\.\(2025\)Biomni: a general\-purpose biomedical AI agent\.bioRxiv\.Note:PreprintExternal Links:2025\.05\.30\.656746,[Document](https://dx.doi.org/10.1101/2025.05.30.656746),[Link](https://doi.org/10.1101/2025.05.30.656746)Cited by:[§1](https://arxiv.org/html/2609.28921#S1.p4.1),[§3\.1](https://arxiv.org/html/2609.28921#S3.SS1.SSS0.Px4.p1.1),[§3\.3](https://arxiv.org/html/2609.28921#S3.SS3.SSS0.Px3.p1.1),[§6](https://arxiv.org/html/2609.28921#S6.SS0.SSS0.Px2.p3.1)\. - \[10\]M\. A\. Robien, K\. T\. Nguyen, A\. Kumar, I\. Hirsh, S\. Turley, D\. Pei, and W\. G\. J\. Hol\(2004\)An improved crystal form of Plasmodium falciparum peptide deformylase\.Protein Science13\(4\),pp\. 1155–1163\.External Links:[Document](https://dx.doi.org/10.1110/ps.03456404)Cited by:[Figure 2](https://arxiv.org/html/2609.28921#S2.F2),[Figure 2](https://arxiv.org/html/2609.28921#S2.F2.9.5)\. - \[11\]P\. Notin, A\. Kollasch, D\. Ritter, L\. van Niekerk, S\. Paul, H\. Spinner, N\. Rollins, A\. Shaw, R\. Orenbuch, R\. Weitzman, J\. Frazer, M\. Dias, D\. Franceschi, Y\. Gal, and D\. S\. Marks\(2023\)ProteinGym: large\-scale benchmarks for protein fitness prediction and design\.InAdvances in Neural Information Processing Systems,External Links:[Link](https://mlanthology.org/neurips/2023/notin2023neurips-proteingym/)Cited by:[§2\.3\.1](https://arxiv.org/html/2609.28921#S2.SS3.SSS1.p1.1),[§2](https://arxiv.org/html/2609.28921#S2.p1.1),[§6](https://arxiv.org/html/2609.28921#S6.SS0.SSS0.Px1.p1.1)\. - \[12\]C\. Dallago, J\. Mou, K\. E\. Johnston, B\. J\. Wittmann, N\. Bhattacharya, S\. Goldman, A\. Madani, and K\. K\. Yang\(2021\)FLIP: benchmark tasks in fitness landscape inference for proteins\.InNeural Information Processing Systems Datasets and Benchmarks Track,External Links:[Link](https://datasets-benchmarks-proceedings.neurips.cc/paper/2021/hash/2b44928ae11fb9384c4cf38708677c48-Abstract-round2.html)Cited by:[§2](https://arxiv.org/html/2609.28921#S2.p1.1),[§6](https://arxiv.org/html/2609.28921#S6.SS0.SSS0.Px1.p1.1)\. - \[13\]D\. Esposito, J\. Weile, J\. Shendure, L\. M\. Starita, A\. T\. Papenfuss, F\. P\. Roth, D\. M\. Fowler, and A\. F\. Rubin\(2019\)MaveDB: an open\-source platform to distribute and interpret data from multiplexed assays of variant effect\.Genome Biology20\(1\),pp\. 223\.External Links:[Document](https://dx.doi.org/10.1186/s13059-019-1845-6),[Link](https://doi.org/10.1186/s13059-019-1845-6)Cited by:[§2\.3\.1](https://arxiv.org/html/2609.28921#S2.SS3.SSS1.p1.1)\. - \[14\]K\. Tsuboyama, J\. Dauparas, J\. Chen, E\. Laine, Y\. Mohseni Behbahani, J\. J\. Weinstein, N\. M\. Mangan, S\. Ovchinnikov, and G\. J\. Rocklin\(2023\)Mega\-scale experimental analysis of protein folding stability in biology and design\.Nature620,pp\. 434–444\.External Links:[Document](https://dx.doi.org/10.1038/s41586-023-06328-6)Cited by:[§2\.3\.1](https://arxiv.org/html/2609.28921#S2.SS3.SSS1.p1.1)\. - \[15\]M\. Chungyoun, J\. Ruffolo, and J\. J\. Gray\(2024\)FLAb: benchmarking tasks in fitness landscape inference for antibodies\.bioRxiv\.External Links:[Document](https://dx.doi.org/10.1101/2024.01.13.575504)Cited by:[§2\.3\.1](https://arxiv.org/html/2609.28921#S2.SS3.SSS1.p1.1)\. - \[16\]A\. Beltran, X\. Jiang, Y\. Shen, and B\. Lehner\(2025\)Site\-saturation mutagenesis of 500 human protein domains\.Nature637\(8047\),pp\. 885–894\.External Links:[Document](https://dx.doi.org/10.1038/s41586-024-08370-4)Cited by:[§2\.3\.1](https://arxiv.org/html/2609.28921#S2.SS3.SSS1.p1.1)\. - \[17\]Y\. Chen, L\. Fu, X\. Lu, W\. Li, Y\. Gao, Y\. Wang, Z\. Ruan, and T\. Si\(2026\)CombinGym: a benchmark platform for machine learning\-assisted design of combinatorial protein variants\.bioRxiv\.External Links:[Document](https://dx.doi.org/10.64898/2026.03.24.714074)Cited by:[§2\.3\.1](https://arxiv.org/html/2609.28921#S2.SS3.SSS1.p1.1)\. - \[18\]H\. Kimura, K\. Lahouel, C\. Tomasetti, and N\. J\. Roberts\(2024\)Functional characterization of all cdkn2a missense variants and comparison to in silico models of pathogenicity\.eLife13,pp\. RP95347\.External Links:[Document](https://dx.doi.org/10.7554/eLife.95347)Cited by:[§2\.3\.1](https://arxiv.org/html/2609.28921#S2.SS3.SSS1.p1.1)\. - \[19\]W\. Wang, E\. Ferrada, C\. Klimek, T\. Osthushenrich, A\. MacNamara, T\. Wiedmer, and G\. Superti\-Furga\(2025\)Large\-scale experimental assessment of variant effects on the structure and function of the citrate transporter slc13a5\.Science Advances11\(26\),pp\. eadx3011\.External Links:[Document](https://dx.doi.org/10.1126/sciadv.adx3011)Cited by:[§2\.3\.1](https://arxiv.org/html/2609.28921#S2.SS3.SSS1.p1.1)\. - \[20\]K\. E\. Johnston, P\. J\. Almhjell, E\. J\. Watkins\-Dulaney, G\. Liu, N\. J\. Porter, J\. Yang, and F\. H\. Arnold\(2024\)A combinatorially complete epistatic fitness landscape in an enzyme active site\.Proceedings of the National Academy of Sciences121\(32\),pp\. e2400439121\.External Links:[Document](https://dx.doi.org/10.1073/pnas.2400439121)Cited by:[§2\.3\.1](https://arxiv.org/html/2609.28921#S2.SS3.SSS1.p1.1)\. - \[21\]M\. Steinegger and J\. Söding\(2017\)MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets\.Nature Biotechnology35\(11\),pp\. 1026–1028\.External Links:[Document](https://dx.doi.org/10.1038/nbt.3988),[Link](https://doi.org/10.1038/nbt.3988)Cited by:[§2\.3\.1](https://arxiv.org/html/2609.28921#S2.SS3.SSS1.p1.1),[§3\.3](https://arxiv.org/html/2609.28921#S3.SS3.SSS0.Px1.p1.1),[§9\.1](https://arxiv.org/html/2609.28921#S9.SS1.p3.1)\. - \[22\]B\. E\. Suzek, Y\. Wang, H\. Huang, P\. B\. McGarvey, C\. H\. Wu, and UniProt Consortium\(2015\)UniRef clusters: a comprehensive and scalable alternative for improving sequence similarity searches\.Bioinformatics31\(6\),pp\. 926–932\.External Links:[Document](https://dx.doi.org/10.1093/bioinformatics/btu739),[Link](https://doi.org/10.1093/bioinformatics/btu739)Cited by:[§2\.3\.1](https://arxiv.org/html/2609.28921#S2.SS3.SSS1.p1.1),[§3\.3](https://arxiv.org/html/2609.28921#S3.SS3.SSS0.Px1.p1.1),[§9\.1](https://arxiv.org/html/2609.28921#S9.SS1.p3.1)\. - \[23\]Openai\(2026\)GPT\-6 astra: a new generation of intelligence\.Note:[https://openai\.com/index/gpt\-6\-astra/](https://openai.com/index/gpt-6-astra/)Cited by:[§3\.1](https://arxiv.org/html/2609.28921#S3.SS1.SSS0.Px3.p1.1)\. - \[24\]G\. DeepMind\(2026\)Gemini 3\.1 pro\.Note:[https://deepmind\.google/models/gemini/pro](https://deepmind.google/models/gemini/pro)Cited by:[§3\.1](https://arxiv.org/html/2609.28921#S3.SS1.SSS0.Px3.p1.1)\. - \[25\]K\. Team\(2026\)Kimi k3: open frontier intelligence\.External Links:2607\.24653,[Link](https://arxiv.org/abs/2607.24653)Cited by:[§3\.1](https://arxiv.org/html/2609.28921#S3.SS1.SSS0.Px3.p1.1)\. - \[26\]Z\.ai\(2026\)GLM\-5\.2: built for long\-horizon tasks\.Note:[https://z\.ai/blog/glm\-5\.2](https://z.ai/blog/glm-5.2)Cited by:[§3\.1](https://arxiv.org/html/2609.28921#S3.SS1.SSS0.Px3.p1.1)\. - \[27\]D\. AI\(2026\)DeepSeek\-v4: towards highly efficient million\-token context intelligence\.Note:[https://huggingface\.co/deepseek\-ai/DeepSeek\-V4\-Pro](https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro)Cited by:[§3\.1](https://arxiv.org/html/2609.28921#S3.SS1.SSS0.Px3.p1.1)\. - \[28\]J\. Abramson, J\. Adler, J\. Dunger, R\. Evans, T\. Green, A\. Pritzel, O\. Ronneberger, L\. Willmore, A\. J\. Ballard, J\. Bambrick, S\. W\. Bodenstein, D\. A\. Evans, C\. Hung, M\. O’Neill, D\. Reiman, K\. Tunyasuvunakool, Z\. Wu, A\. Žemgulytė, E\. Arvaniti, C\. Beattie, O\. Bertolli, A\. Bridgland, A\. Cherepanov, M\. Congreve, A\. I\. Cowen\-Rivers, A\. Cowie, M\. Figurnov, F\. B\. Fuchs, H\. Gladman, R\. Jain, Y\. A\. Khan, C\. M\. R\. Low, K\. Perlin, A\. Potapenko, P\. Savy, S\. Singh, A\. Stecula, A\. Thillaisundaram, C\. Tong, S\. Yakneen, E\. D\. Zhong, M\. Zielinski, A\. Žídek, V\. Bapst, P\. Kohli, M\. Jaderberg, D\. Hassabis, and J\. M\. Jumper\(2024\)Accurate structure prediction of biomolecular interactions with AlphaFold 3\.Nature630\(8016\),pp\. 493–500\.External Links:[Document](https://dx.doi.org/10.1038/s41586-024-07487-w),[Link](https://doi.org/10.1038/s41586-024-07487-w)Cited by:[§3\.3](https://arxiv.org/html/2609.28921#S3.SS3.SSS0.Px1.p1.1),[§9\.1](https://arxiv.org/html/2609.28921#S9.SS1.p4.1)\. - \[29\]K\. Qiu, Y\. Wu, L\. Wang, Y\. Ouyang, J\. Yu, Z\. Zhou, C\. Lv, D\. Xue, Y\. Song, X\. Zhang,et al\.\(2026\)AMix\-2: establishing protein as a native modality in large language models\.arXiv preprint arXiv:2605\.30963\.Cited by:[§3\.3](https://arxiv.org/html/2609.28921#S3.SS3.SSS0.Px4.p4.1)\. - \[30\]K\. Didi, S\. Alamdari, A\. X\. Lu, B\. Wittmann, K\. E\. Johnston, A\. A\. Amini, A\. Madani, M\. Czeneszew, C\. Dallago, and K\. K\. Yang\(2026\)FLIP2: expanding protein fitness landscape benchmarks for real\-world machine learning applications\.InForty\-third International Conference on Machine Learning,External Links:[Link](https://flip.protein.properties/)Cited by:[§6](https://arxiv.org/html/2609.28921#S6.SS0.SSS0.Px1.p1.1)\. - \[31\]R\. K\. Arora, L\. T\. Chen, M\. Du, D\. Marks, and G\. Church\(2026\)PG\-LLM: benchmarking general\-purpose language models for protein variant ranking\.bioRxiv\.External Links:[Document](https://dx.doi.org/10.64898/2026.07.27.741045),[Link](https://www.biorxiv.org/content/10.64898/2026.07.27.741045v1)Cited by:[§6](https://arxiv.org/html/2609.28921#S6.SS0.SSS0.Px1.p2.1)\. - \[32\]J\. Kim and P\. Romero\(2026\)Evaluating LLM\-driven protein design: agents lack iterative evaluation depth\.Note:GitHub repositoryBioDesignBenchExternal Links:[Link](https://github.com/RomeroLab/BioDesignBench)Cited by:[§6](https://arxiv.org/html/2609.28921#S6.SS0.SSS0.Px1.p2.1)\. - \[33\]S\. Biswas, G\. Khimulya, E\. C\. Alley, K\. M\. Esvelt, and G\. M\. Church\(2021\)Low\-N protein engineering with data\-efficient deep learning\.Nature Methods18\(4\),pp\. 389–396\.External Links:[Document](https://dx.doi.org/10.1038/s41592-021-01100-y),[Link](https://doi.org/10.1038/s41592-021-01100-y)Cited by:[§6](https://arxiv.org/html/2609.28921#S6.SS0.SSS0.Px2.p1.1)\. - \[34\]C\. Hsu, H\. Nisonoff, C\. Fannjiang, and J\. Listgarten\(2022\)Learning protein fitness models from evolutionary and assay\-labeled data\.Nature Biotechnology40\(7\),pp\. 1114–1122\.External Links:[Document](https://dx.doi.org/10.1038/s41587-021-01146-5),[Link](https://doi.org/10.1038/s41587-021-01146-5)Cited by:[§6](https://arxiv.org/html/2609.28921#S6.SS0.SSS0.Px2.p1.1)\. - \[35\]T\. A\. Hopf, J\. B\. Ingraham, F\. J\. Poelwijk, C\. P\. I\. Schärfe, M\. Springer, C\. Sander, and D\. S\. Marks\(2017\)Mutation effects predicted from sequence co\-variation\.Nature Biotechnology35\(2\),pp\. 128–135\.External Links:[Document](https://dx.doi.org/10.1038/nbt.3769),[Link](https://doi.org/10.1038/nbt.3769)Cited by:[§6](https://arxiv.org/html/2609.28921#S6.SS0.SSS0.Px2.p2.1)\. - \[36\]A\. J\. Riesselman, J\. B\. Ingraham, and D\. S\. Marks\(2018\)Deep generative models of genetic variation capture the effects of mutations\.Nature Methods15\(10\),pp\. 816–822\.External Links:[Document](https://dx.doi.org/10.1038/s41592-018-0138-4),[Link](https://doi.org/10.1038/s41592-018-0138-4)Cited by:[§6](https://arxiv.org/html/2609.28921#S6.SS0.SSS0.Px2.p2.1)\. - \[37\]J\. Meier, R\. Rao, R\. Verkuil, J\. Liu, T\. Sercu, and A\. Rives\(2021\)Language models enable zero\-shot prediction of the effects of mutations on protein function\.InAdvances in Neural Information Processing Systems,Vol\.34,pp\. 29287–29303\.External Links:[Link](https://proceedings.neurips.cc/paper/2021/hash/f51338d736f95dd42427296047067694-Abstract.html)Cited by:[§6](https://arxiv.org/html/2609.28921#S6.SS0.SSS0.Px2.p2.1)\. - \[38\]P\. Notin, M\. Dias, J\. Frazer, J\. Marchena\-Hurtado, A\. N\. Gomez, D\. S\. Marks, and Y\. Gal\(2022\)Tranception: protein fitness prediction with autoregressive transformers and inference\-time retrieval\.InProceedings of the 39th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.162,pp\. 16990–17017\.External Links:[Link](https://proceedings.mlr.press/v162/notin22a.html)Cited by:[§6](https://arxiv.org/html/2609.28921#S6.SS0.SSS0.Px2.p2.1)\. - \[39\]X\. Wang, L\. Hong, J\. Ye, Z\. Zheng, S\. Huang, and Q\. Gu\(2026\)Towards a generative protein evolution machine with DPLM\-evo\.InICLR 2026 Workshop on Generative and Experimental Perspectives for Biomolecular Design,External Links:[Link](https://openreview.net/forum?id=xqnTGzQmfB)Cited by:[§6](https://arxiv.org/html/2609.28921#S6.SS0.SSS0.Px2.p2.1)\. - \[40\]Y\. Wang, J\. He, Y\. Du, X\. Chen, J\. C\. Li, L\. Liu, X\. Xu, and S\. Hassoun\(2025\)Large language model is secretly a protein sequence optimizer\.InICLR 2025 Workshop on Learning Meaningful Representations of Life,External Links:2501\.09274,[Link](https://arxiv.org/abs/2501.09274)Cited by:[§6](https://arxiv.org/html/2609.28921#S6.SS0.SSS0.Px2.p3.1)\. - \[41\]A\. Ghafarollahi and M\. J\. Buehler\(2024\)ProtAgents: protein discovery via large language model multi\-agent collaborations combining physics and machine learning\.Digital Discovery3\(7\),pp\. 1389–1409\.External Links:[Document](https://dx.doi.org/10.1039/D4DD00013G),[Link](https://doi.org/10.1039/D4DD00013G)Cited by:[§6](https://arxiv.org/html/2609.28921#S6.SS0.SSS0.Px2.p3.1)\. \\beginappendix ## 8Evaluation Templates To support standardized and reproducible evaluation, we provide the prompt templates used for the evaluation of general\-purpose language models onPFArena\. Agentic models use similar core templates, supplemented with additional instructions for tool calling\. ### 8\.1Single\-mutant Generation TaskT1: Single\-mutant generationprovides the language model with thewildtype\_sequenceof the evaluated protein, itssequence\_length, and its UniProt identifier \(uniprot\_id, e\.g\., “P11413”\)\. A concise description of its experimental assay context is also passed to the model, including theprimary\_task\_class\(e\.g\., “activity\_function”\),fitness\_type\(e\.g\., “enzymatic\_activity”\), andassay\_readout\_subclass\(e\.g\., “activity\_proxy”\)\. The model is instructed to generate a list of plausible mutations and rank them according to their expected fitness based on its own internal knowledge\. Youareanexpertproteinengineerandcomputationalbiologistspecializingindeepmutationalscanning\(DMS\)andmutation\-effectprediction\. \#\#\#TASKGOAL Givenawild\-typeproteinsequenceanditsexperimentalassaycontext,predictthe\*\*top40singlepointmutations\*\*\(WTposMUTformat\)thatoptimizethetargetfitnessmetric,orderedfromhighestexpectedfitnesstolowestexpectedfitness\. \#\#\#REASONING&EVIDENCEBOUNDARIES 1\.\*\*BiochemicalDeductions\*\*:Analyzeresiduechemistry,conservation,secondarystructurepropensities,stericpacking,hydrophobiccores,electrostaticinteractions,andsequencemotifswithinthesuppliedwild\-typesequence\. 2\.\*\*AssayAlignment\*\*:Aligneveryrankedmutationstrictlywiththesuppliedassayreadout\.Forexample,ifevaluatingstability/abundance,prioritizemutationsthatimprovehydrophobicpackingorthermostabilitywithoutdisruptingnecessarystructuraldynamics\. 3\.\*\*NoHallucination\*\*:DoNOTinventorclaimspecificnumericalmodelscores,experimentalPDBcoordinates,literaturemeasurements,oralignmentsthatarenotlogicallyderivablefromsequencebiochemistry\. 4\.\*\*Single\-ResponseConstraint\*\*:Youhavenoexternaltools,browsing,orfollow\-upturns\.Completetheanalysisinthissingleresponse\. \#\#\#STRICTMUTATION&FORMATCONSTRAINTS 1\.\*\*RankingSize\*\*:TheoutputlistMUSTcontain\*\*exactly40mutations\*\*,orderedfrombesttoworst\. 2\.\*\*Format\*\*:EverymutantMUSTberepresentedinstandard1\-indexed‘WTposMUT‘format\(e\.g\.,‘H24R‘,‘A15V‘\)\. 3\.\*\*Alphabet\*\*:Both‘WT‘and‘MUT‘mustbestandard20aminoacidsingle\-lettercodes:‘A,C,D,E,F,G,H,I,K,L,M,N,P,Q,R,S,T,V,W,Y‘\. 4\.\*\*Uniqueness\*\*:EverymutationstringMUSTappearexactlyonce\(noduplicates,noomissionswithinyourtop\-40list\)\. 5\.\*\*StrictValidation\*\*: \-Everymutationpositionfollows‘1<=position<=Length‘\. \-‘WT‘MUSTstrictlymatchthecharacterat‘wildtype\_sequence\[position\-1\]‘\. \-‘MUT‘MUSTbestrictlydifferentfrom‘WT‘\(nosynonymous/no\-opmutations\)\. 6\.\*\*SearchSpace\*\*:Considerallvalidsingleaminoacidsubstitutionsacrossthefullwild\-typesequence,thenreturnonlythetop40\. 7\.\*\*ProhibitedFormats\*\*:Multi\-sitemutations,insertions\(‘ins‘\),deletions\(‘del‘\),stopcodons\(‘\*‘\),HGVSnotations,ornumericalconfidencescores\. — \#\#\#OUTPUTFORMAT Returna\*\*validJSONobjectONLY\*\*withNOmarkdowncodeblockwrappers,prefix,orconversationaltext\.UsethefollowingexactJSONschema: \{ "ranking":\[ "<bestmutant\>", "<second\-bestmutant\>", "…" \] \} — \#\#\#INSTANCEDATA 1\.ASSAYCONTEXT&TARGETMETRICS \-\*\*UniProtID\*\*:\{uniprot\_id\} \-\*\*PrimaryTaskClass\*\*:\{primary\_task\_class\} \-\*\*FitnessMetricType\*\*:\{fitness\_type\} \-\*\*AssayReadoutSubclass\*\*:\{readout\_subclass\} 2\.INPUTWILD\-TYPESEQUENCE\(Length:\{sequence\_length\}\) ‘\{wildtype\_sequence\}‘ ### 8\.2Multi\-mutant Ranking For the mutation\-ranking tasks ofT2–T4, the protein and assay metadata included inT1is similarly provided\. Additionally, the model receives the total amount of candidates to ranknum\_candidates, and a shuffled list ofcandidate\_mutants, each specified with the original amino\-acid and position in the wild\-type sequence followed by the mutant, in a “WTposMUT” format like “E19H”\. ##### Measurement\-free Ranking No additional experimental measurements are provided forT2: Measurement\-free multi\-mutant ranking\. The model ranks the candidate mutants using only the shared information on the protein and assay context\. Youareanexpertproteinengineerandcomputationalbiologistspecializingindeepmutationalscanning\(DMS\)andmutation\-effectprediction\. \#\#\#TASKGOAL Givenawild\-typeproteinsequence,itsexperimentalassaycontext,andaspecificlistofcandidatemulti\-mutations,\*\*rankallcandidatemulti\-mutationsfrombesttoworst\*\*accordingtotheirexpectedtargetfitnessmetric\. \#\#\#REASONING&EVIDENCEBOUNDARIES 1\.\*\*BiochemicalDeductions\*\*:Comparethecandidatesbasedonresiduechemistry,conservation,secondarystructurepropensities,stericpacking,hydrophobiccores,electrostaticinteractions,andsequencemotifswithinthesuppliedwild\-typesequence\. 2\.\*\*AssayAlignment\*\*:Evaluaterelativeeffectsofthesespecificsubstitutionsonthesuppliedassayreadout\.Rankmutationsthatbetterpreserveorenhancestructural/functionalrequirementsabovethosethatintroducesevereclashes,chargemismatch,orinstability\. 3\.\*\*NoHallucination\*\*:DoNOTinventorclaimspecificnumericalmodelscores,experimentalPDBcoordinates,literaturemeasurements,oralignmentsthatarenotlogicallyderivablefromsequencebiochemistry\. 4\.\*\*Single\-ResponseConstraint\*\*:Youhavenoexternaltools,browsing,orfollow\-upturns\.Completetherankinginthissingleresponse\. \#\#\#STRICTMUTATION&FORMATCONSTRAINTS 1\.\*\*ClosedSetPrinciple\*\*:YouMUSTONLYrankthemutationsprovidedinthelistabove\.DoNOTintroducenewmutations,insertions,deletions,orwild\-typestrings\. 2\.\*\*Multi\-MutationCandidates\*\*:Acandidatecontainsmultiplemutationsjoinedby\+\(e\.g\.,K23A\+A40P\+T52S\)\.Whenrankingmulti\-mutationcandidates,accountfortheircombinedeffects\. 3\.\*\*ExactCopy&Completeness\*\*: \-TheoutputlistMUSTcontain\*\*exactly‘total\_candidates‘items\*\*\. \-EverycandidatefromtheinputlistMUSTappear\*\*exactlyonce\*\*\(noduplicates,noomissions\)\. \-EachmutationstringMUSTbecopied\*\*exactlyasprovided\*\*\. 4\.\*\*StrictValidation\*\*: \-Everymutationpositionfollows‘1<=position<=Length‘\. \-‘WT‘MUSTstrictlymatchthecharacterat‘wildtype\_sequence\[position\-1\]‘\. \-‘MUT‘MUSTbestrictlydifferentfrom‘WT‘\(nosynonymous/no\-opmutations\)\. — \#\#\#OUTPUTFORMAT Returna\*\*validJSONobjectONLY\*\*withNOmarkdowncodeblockwrappers,prefix,orconversationaltext\.UsethefollowingexactJSONschema: \{ "ranking":\[ "<bestmutantcopiedexactlyfromtheprovidedlist\>", "<second\-bestmutantcopiedexactlyfromtheprovidedlist\>", "…" \] \} — \#\#\#INSTANCEDATA 1\.ASSAYCONTEXT&TARGETMETRICS \-\*\*UniProtID\*\*:\{uniprot\_id\} \-\*\*PrimaryTaskClass\*\*:\{primary\_task\_class\} \-\*\*FitnessMetricType\*\*:\{fitness\_type\} \-\*\*AssayReadoutSubclass\*\*:\{readout\_subclass\} 2\.INPUTWILD\-TYPESEQUENCE\(Length:\{sequence\_length\}\) ‘\{wildtype\_sequence\}‘ 3\.CANDIDATEMUTATIONSTORANK\(TotalCandidates:\{num\_candidates\}\) \{candidate\_mutants\} ##### Anchor\-informed Ranking In taskT3: Anchor\-informed multi\-mutant ranking, the model additionally receives ananchor\_mutantshared by all candidates and its ground\-truthanchor\_DMS\_score\. This measurement provides a common experimental reference for ranking the candidate mutations\. Youareanexpertproteinengineerandcomputationalbiologistspecializingindeepmutationalscanning\(DMS\)andmutation\-effectprediction\. \#\#\#TASKGOAL Givenawild\-typeproteinsequence,itsexperimentalassaycontext,andaspecificlistofcandidatemulti\-mutations,\*\*rankallcandidatemulti\-mutationsfrombesttoworst\*\*accordingtotheirexpectedtargetfitnessmetric\. <<Theground\-truthDMSscoreofananchormutantappearingineverycandidate,whichmaybesingle\-siteormulti\-site,willalsobeprovided\.\>\> \#\#\#REASONING&EVIDENCEBOUNDARIES 1\.\*\*BiochemicalDeductions\*\*:Comparethecandidatesbasedonresiduechemistry,conservation,secondarystructurepropensities,stericpacking,hydrophobiccores,electrostaticinteractions,andsequencemotifswithinthesuppliedwild\-typesequence\. 2\.\*\*AssayAlignment\*\*:Evaluaterelativeeffectsofthesespecificsubstitutionsonthesuppliedassayreadout\.Rankmutationsthatbetterpreserveorenhancestructural/functionalrequirementsabovethosethatintroducesevereclashes,chargemismatch,orinstability\. 3\.\*\*NoHallucination\*\*:DoNOTinventorclaimspecificnumericalmodelscores,experimentalPDBcoordinates,literaturemeasurements,oralignmentsthatarenotlogicallyderivablefromsequencebiochemistry\. 4\.\*\*Single\-ResponseConstraint\*\*:Youhavenoexternaltools,browsing,orfollow\-upturns\.Completetherankinginthissingleresponse\. \#\#\#STRICTMUTATION&FORMATCONSTRAINTS 1\.\*\*ClosedSetPrinciple\*\*:YouMUSTONLYrankthemutationsprovidedinthelistabove\.DoNOTintroducenewmutations,insertions,deletions,orwild\-typestrings\. 2\.\*\*Multi\-MutationCandidates\*\*:Acandidatecontainsmultiplemutationsjoinedby\+\(e\.g\.,K23A\+A40P\+T52S\)<<,withtheanchormutantincluded\>\>\.Whenrankingmulti\-mutationcandidates,accountfortheircombinedeffects\. 3\.\*\*ExactCopy&Completeness\*\*: \-TheoutputlistMUSTcontain\*\*exactly‘total\_candidates‘items\*\*\. \-EverycandidatefromtheinputlistMUSTappear\*\*exactlyonce\*\*\(noduplicates,noomissions\)\. \-EachmutationstringMUSTbecopied\*\*exactlyasprovided\*\*\. 4\.\*\*StrictValidation\*\*: \-Everymutationpositionfollows‘1<=position<=Length‘\. \-‘WT‘MUSTstrictlymatchthecharacterat‘wildtype\_sequence\[position\-1\]‘\. \-‘MUT‘MUSTbestrictlydifferentfrom‘WT‘\(nosynonymous/no\-opmutations\)\. — \#\#\#OUTPUTFORMAT Returna\*\*validJSONobjectONLY\*\*withNOmarkdowncodeblockwrappers,prefix,orconversationaltext\.UsethefollowingexactJSONschema: \{ "ranking":\[ "<bestmutantcopiedexactlyfromtheprovidedlist\>", "<second\-bestmutantcopiedexactlyfromtheprovidedlist\>", "…" \] \} — \#\#\#INSTANCEDATA 1\.ASSAYCONTEXT&TARGETMETRICS \-\*\*UniProtID\*\*:\{uniprot\_id\} \-\*\*PrimaryTaskClass\*\*:\{primary\_task\_class\} \-\*\*FitnessMetricType\*\*:\{fitness\_type\} \-\*\*AssayReadoutSubclass\*\*:\{readout\_subclass\} 2\.INPUTWILD\-TYPESEQUENCE\(Length:\{sequence\_length\}\) ‘\{wildtype\_sequence\}‘ <<3\.ANCHORMUTANTCONTEXT \-\*\*AnchorMutant\*\*:\{anchor\_mutant\} \-\*\*AnchorDMSScore\*\*:\{anchor\_DMS\_score\}\>\> 4\.CANDIDATEMUTATIONSTORANK\(TotalCandidates:\{num\_candidates\}\) \{candidate\_mutants\} ##### Single\-mutant\-informed Ranking TaskT4: Single\-mutant\-informed multi\-mutant rankingadditionally provides the ground\-truth fitness scores of every component mutation appearing in the candidates assingle\_mutant\_DMS\_score\. Each mutation–score pair is formatted as “<mutant\>: <dms\_score\>”, and the pairs are joined by newline characters \(“\\n”\)\. These measurements provide mutation\-level evidence for ranking the multi\-mutant combinations\. Youareanexpertproteinengineerandcomputationalbiologistspecializingindeepmutationalscanning\(DMS\)andmutation\-effectprediction\. \#\#\#TASKGOAL Givenawild\-typeproteinsequence,itsexperimentalassaycontext,andaspecificlistofcandidatemulti\-mutations,\*\*rankallcandidatemulti\-mutationsfrombesttoworst\*\*accordingtotheirexpectedtargetfitnessmetric\. <<Theground\-truthsingle\-mutantDMSscoresofeverycomponentappearinginanycandidatewillalsobeprovided\.\>\> \#\#\#REASONING&EVIDENCEBOUNDARIES 1\.\*\*BiochemicalDeductions\*\*:Comparethecandidatesbasedonresiduechemistry,conservation,secondarystructurepropensities,stericpacking,hydrophobiccores,electrostaticinteractions,andsequencemotifswithinthesuppliedwild\-typesequence\. 2\.\*\*AssayAlignment\*\*:Evaluaterelativeeffectsofthesespecificsubstitutionsonthesuppliedassayreadout\.Rankmutationsthatbetterpreserveorenhancestructural/functionalrequirementsabovethosethatintroducesevereclashes,chargemismatch,orinstability\. 3\.\*\*NoHallucination\*\*:DoNOTinventorclaimspecificnumericalmodelscores,experimentalPDBcoordinates,literaturemeasurements,oralignmentsthatarenotlogicallyderivablefromsequencebiochemistry\. 4\.\*\*Single\-ResponseConstraint\*\*:Youhavenoexternaltools,browsing,orfollow\-upturns\.Completetherankinginthissingleresponse\. \#\#\#STRICTMUTATION&FORMATCONSTRAINTS 1\.\*\*ClosedSetPrinciple\*\*:YouMUSTONLYrankthemutationsprovidedinthelistabove\.DoNOTintroducenewmutations,insertions,deletions,orwild\-typestrings\. 2\.\*\*Multi\-MutationCandidates\*\*:Acandidatecontainsmultiplemutationsjoinedby\+\(e\.g\.,K23A\+A40P\+T52S\)\.Whenrankingmulti\-mutationcandidates,accountfortheircombinedeffects\. 3\.\*\*ExactCopy&Completeness\*\*: \-TheoutputlistMUSTcontain\*\*exactly‘total\_candidates‘items\*\*\. \-EverycandidatefromtheinputlistMUSTappear\*\*exactlyonce\*\*\(noduplicates,noomissions\)\. \-EachmutationstringMUSTbecopied\*\*exactlyasprovided\*\*\. 4\.\*\*StrictValidation\*\*: \-Everymutationpositionfollows‘1<=position<=Length‘\. \-‘WT‘MUSTstrictlymatchthecharacterat‘wildtype\_sequence\[position\-1\]‘\. \-‘MUT‘MUSTbestrictlydifferentfrom‘WT‘\(nosynonymous/no\-opmutations\)\. — \#\#\#OUTPUTFORMAT Returna\*\*validJSONobjectONLY\*\*withNOmarkdowncodeblockwrappers,prefix,orconversationaltext\.UsethefollowingexactJSONschema: \{ "ranking":\[ "<bestmutantcopiedexactlyfromtheprovidedlist\>", "<second\-bestmutantcopiedexactlyfromtheprovidedlist\>", "…" \] \} — \#\#\#INSTANCEDATA 1\.ASSAYCONTEXT&TARGETMETRICS \-\*\*UniProtID\*\*:\{uniprot\_id\} \-\*\*PrimaryTaskClass\*\*:\{primary\_task\_class\} \-\*\*FitnessMetricType\*\*:\{fitness\_type\} \-\*\*AssayReadoutSubclass\*\*:\{readout\_subclass\} 2\.INPUTWILD\-TYPESEQUENCE\(Length:\{sequence\_length\}\) ‘\{wildtype\_sequence\}‘ <<3\.SINGLEMUTANTCONTEXT \{single\_mutant\_dms\_scores\}\>\> 4\.CANDIDATEMUTATIONSTORANK\(TotalCandidates:\{num\_candidates\}\) \{candidate\_mutants\} ### 8\.3Confidence Elicitation for Uncertainty Estimation To elicit list\-level confidence, we inserted the confidence elicitation instruction immediately before the original output contract\. The original JSON ranking schema was retained, with oneconfidencefield added after the ranking array; the non\-numerical placeholder<your confidence\>was used to avoid anchoring the model to an example value\. The same confidence instruction was used in all settings\. InT1, the original prohibition on “numerical confidence scores” was changed to “per\-mutation confidence scores” to permit the required list\-level value\. All other task instructions, ranking constraints, assay information, sequences, candidate data, and contextual scores remained unchanged\. Youareanexpertproteinengineerandcomputationalbiologistspecializingindeepmutationalscanning\(DMS\)andmutation\-effectprediction\. \#\#\#TASKGOAL Givenawild\-typeproteinsequenceanditsexperimentalassaycontext,predictthe\*\*top40singlepointmutations\*\*\(WTposMUTformat\)thatoptimizethetargetfitnessmetric,orderedfromhighestexpectedfitnesstolowestexpectedfitness\. \#\#\#REASONING&EVIDENCEBOUNDARIES 1\.\*\*BiochemicalDeductions\*\*:Analyzeresiduechemistry,conservation,secondarystructurepropensities,stericpacking,hydrophobiccores,electrostaticinteractions,andsequencemotifswithinthesuppliedwild\-typesequence\. 2\.\*\*AssayAlignment\*\*:Aligneveryrankedmutationstrictlywiththesuppliedassayreadout\.Forexample,ifevaluatingstability/abundance,prioritizemutationsthatimprovehydrophobicpackingorthermostabilitywithoutdisruptingnecessarystructuraldynamics\. 3\.\*\*NoHallucination\*\*:DoNOTinventorclaimspecificnumericalmodelscores,experimentalPDBcoordinates,literaturemeasurements,oralignmentsthatarenotlogicallyderivablefromsequencebiochemistry\. 4\.\*\*Single\-ResponseConstraint\*\*:Youhavenoexternaltools,browsing,orfollow\-upturns\.Completetheanalysisinthissingleresponse\. \#\#\#STRICTMUTATION&FORMATCONSTRAINTS 1\.\*\*RankingSize\*\*:TheoutputlistMUSTcontain\*\*exactly40mutations\*\*,orderedfrombesttoworst\. 2\.\*\*Format\*\*:EverymutantMUSTberepresentedinstandard1\-indexed‘WTposMUT‘format\(e\.g\.,‘H24R‘,‘A15V‘\)\. 3\.\*\*Alphabet\*\*:Both‘WT‘and‘MUT‘mustbestandard20aminoacidsingle\-lettercodes:‘A,C,D,E,F,G,H,I,K,L,M,N,P,Q,R,S,T,V,W,Y‘\. 4\.\*\*Uniqueness\*\*:EverymutationstringMUSTappearexactlyonce\(noduplicates,noomissionswithinyourtop\-40list\)\. 5\.\*\*StrictValidation\*\*: \-Everymutationpositionfollows‘1<=position<=Length‘\. \-‘WT‘MUSTstrictlymatchthecharacterat‘wildtype\_sequence\[position\-1\]‘\. \-‘MUT‘MUSTbestrictlydifferentfrom‘WT‘\(nosynonymous/no\-opmutations\)\. 6\.\*\*SearchSpace\*\*:Considerallvalidsingleaminoacidsubstitutionsacrossthefullwild\-typesequence,thenreturnonlythetop40\. 7\.\*\*ProhibitedFormats\*\*:Multi\-sitemutations,insertions\(‘ins‘\),deletions\(‘del‘\),stopcodons\(‘\*‘\),HGVSnotations,orper\-mutationconfidencescores\. — <<\#\#\#CONFIDENCE Afterproducingtheranking,reportanoverallconfidencescorefrom0to100indicatinghowconfidentyouareinthepredictivequalityofthetop\-40listforthisspecificproteinandassay\.Thescoreshouldreflectconfidencethatthelistcontainsandprioritizesgenuinelyhigh\-fitnessmutations,ratherthanconfidenceinformattingorinstructionfollowing\. Reportexactlyonelist\-level‘confidence‘JSONnumberin\[0,100\]\.Donotaddper\-mutationconfidencescoresoranyotherfields\.Thisconfidenceinstructionmustnotchangetherankingoritem\-countrules\.\>\> \#\#\#OUTPUTFORMAT Returna\*\*validJSONobjectONLY\*\*withNOmarkdowncodeblockwrappers,prefix,orconversationaltext\.UsethefollowingexactJSONschema: \{ "ranking":\[ "<bestmutant\>", "<second\-bestmutant\>", "…" \], <<"confidence":<yourconfidence\>\>\> \} — \#\#\#INSTANCEDATA 1\.ASSAYCONTEXT&TARGETMETRICS \-\*\*UniProtID\*\*:<uniprot\_id\> \-\*\*PrimaryTaskClass\*\*:<primary\_task\_class\> \-\*\*FitnessMetricType\*\*:<fitness\_type\> \-\*\*AssayReadoutSubclass\*\*:<assay\_readout\_subclass\> 2\.INPUTWILD\-TYPESEQUENCE\(Length:<sequence\_length\>\) ‘<wildtype\_sequence\>‘ ## 9Experimental Setup ### 9\.1PLM Deployment [Table11](https://arxiv.org/html/2609.28921#S9.T11)summarizes the model\-specific inputs, checkpoints, inference configurations, and scoring definitions\. Neural\-network inference used FP32 precision, and all scores were oriented such that larger values indicate more favorable candidates\. Each run was validated for input consistency, complete candidate coverage, finite outputs, and agreement between aggregate and component scores\. ModelInputs and CheckpointCore ConfigurationCandidate ScoreESM\-2\[[1](https://arxiv.org/html/2609.28921#bib.bib16)\]Wild\-type sequence;esm2\_t33\_650M\_UR50DMasked inference with a 1,024\-token context \(at most 1,022 residues\); deterministic mutation\-centered windows for longer chainsSum of wild\-type\-context masked\-marginal log odds,logp\(ximut\)−logp\(xiwt\)\\log p\(x\_\{i\}^\{\\mathrm\{mut\}\}\)\-\\log p\(x\_\{i\}^\{\\mathrm\{wt\}\}\), over substituted sitesProGen2\-base\[[2](https://arxiv.org/html/2609.28921#bib.bib17)\]Complete wild\-type and mutant sequences;progen2\-baseForward and reversed causal sequence scoring with a 2,048\-token context; all evaluated chains were processed at full lengthDifference between the bidirectional sequence scores of the complete mutant and wild\-type chainsProSST\-2048\[[3](https://arxiv.org/html/2609.28921#bib.bib18)\]Wild\-type sequence and AlphaFold 3 structure;AI4Protein/ProSST\-2048Official GVP quantizer with a 2,048\-code structural vocabulary; all evaluated chains were processed at full length within the 2,046\-residue limitSum of structure\-conditioned wild\-type\-context marginal log odds over substituted sitesS3F\[[4](https://arxiv.org/html/2609.28921#bib.bib19)\]Wild\-type sequence, AlphaFold 3 structure, and the corresponding molecular surface; released S3F checkpointMutation\-centered windows of at most 1,022 residues; sequence logits replace structure\-conditioned logits where AF3 pLDDT is below 70Sum of masked\-marginal log\-odds contributions over substituted sitesVenusREM\[[5](https://arxiv.org/html/2609.28921#bib.bib21)\]Wild\-type sequence, ProSST\-2048 structural tokens, and a UniRef100/MMseqs2 MSAaa\_seq\_alnretrieval;α=0\.8\\alpha=0\.8, sampling ratio1\.01\.0, and one sampling pass; full\-chain inference for all evaluated contextsSum of mutant\-minus\-wild\-type values from the fused sequence, structure, and MSA representationS3F\-MSA\[[4](https://arxiv.org/html/2609.28921#bib.bib19),[6](https://arxiv.org/html/2609.28921#bib.bib20)\]S3F score and an EVE ensemble trained on the corresponding wild\-type\-chain MSAFive EVE seeds; 400,000 optimization steps, batch size 256, learning rate10−410^\{\-4\}; 20,000 Monte Carlo samples per seed at scoring timeEqual average of query\-wise standardized S3F and EVE scores, where the EVE score is the negative mean evolution index across seedsTable 11:Deployment and scoring configurations of the protein\-model baselines\.Multi\-substitution scoring followed each baseline’s formulation\. For ESM\-2, ProSST\-2048, S3F, and VenusREM, substitutions on the same chain were evaluated in a shared wild\-type context and their sitewise contributions were summed\. ProGen2\-base instead evaluated the complete mutant sequence, whereas the EVE component of S3F\-MSA assigned a joint score to the complete within\-chain mutation set\. For candidates spanning multiple chains, each naturally mutated chain was scored separately and the chain\-level contributions were summed\. Thus, the protocol preserves native chain boundaries but does not explicitly model inter\-chain epistasis\. The MSAs used by VenusREM and EVE were generated against UniRef100\[[22](https://arxiv.org/html/2609.28921#bib.bib24)\]with MMseqs2\[[21](https://arxiv.org/html/2609.28921#bib.bib23)\]\. One A3M alignment was associated with each wild\-type chain; lowercase insertion symbols were removed where required by the parser, while alignment gaps were retained\. For EVE, the focus\-column and sequence\-fragment gap thresholds were 1\.0 and 0\.5, respectively, and the sequence\-reweighting threshold was 0\.01 for viral proteins and 0\.2 otherwise\. Five EVE models were associated with each input context\. Existing EVE ensembles were reused only when the wild\-type sequence, processed MSA, and reweighting configuration were identical; all benchmark mutation sets were scored anew\. Structure tokens, molecular surfaces, and MSA\-derived models were prepared once per wild\-type context rather than for every candidate mutation\. We generated a consistent set of wild\-type structures with AlphaFold 3 \(AF3\)\[[28](https://arxiv.org/html/2609.28921#bib.bib22)\]for the structure\-dependent baselines\. Each natural chain was modeled independently as a monomer using random seed 1, 10 recycles, and five diffusion samples\. The corresponding UniRef100/MMseqs2 unpaired MSA was supplied directly; templates and additional database searches were disabled\. The highest\-ranked sample provided the PDB input used to derive ProSST structural tokens and S3F molecular surfaces\. The canonical mmCIF files, confidence outputs, sample rankings, and all five samples were retained for reproducibility\. For the benchmark release, AF3 was run for all unique wild\-type chains\. Each result was required to reproduce the exact target sequence and residue mapping and to contain a complete protein backbone and valid confidence arrays\. Of these, 184 structures passed the primary confidence criteria; the remaining 12 had lower pTM or mean pLDDT and were retained with a review flag because they remained sequence\-consistent, structurally complete, and free of detected clashes\. These flags were propagated with the structural resources so that confidence\-dependent analyses can be performed without changing the benchmark coverage\. ### 9\.2Test\-time Scaling Methods For eachT1: Single\-mutant generationassay, let theii\-th independently sampled top\-40 ranking be Ri=\[mi1,mi2,…,mi40\],R\_\{i\}=\[m\_\{i1\},m\_\{i2\},\\ldots,m\_\{i40\}\], and letri\(m\)r\_\{i\}\(m\)denote the one\-indexed rank of mutationmminRiR\_\{i\}\. We evaluatedn∈\{1,2,4,8,16\}n\\in\\\{1,2,4,8,16\\\}using nested prefixes of the same 16 samples\. Sample Avg was computed by evaluating each eligible raw ranking independently and averaging its assay\-level metric\. All aggregation methods received the same extracted candidate strings\. Candidates were deduplicated within each ranking while preserving their first occurrence\. In the main scaling experiments, no explicit sequence\-validity or DMS\-membership filtering was applied\. Incomplete rankings were padded to 40 entries using sample\-specific invalid placeholders, preventing missing outputs from creating artificial agreement across samples\. The aggregated top\-40 output was mapped and scored only by the common evaluator\. ##### Best\-of\-N We first computed the cross\-sample reciprocal\-rank consensus of each candidate, c\(m\)=∑i:m∈Ri1k\+ri\(m\),c\(m\)=\\sum\_\{i:m\\in R\_\{i\}\}\\frac\{1\}\{k\+r\_\{i\}\(m\)\}, withk=10k=10\. Each complete ranking received the mean consensus score qi=1\|Ri\|∑m∈Ric\(m\),q\_\{i\}=\\frac\{1\}\{\|R\_\{i\}\|\}\\sum\_\{m\\in R\_\{i\}\}c\(m\), and the ranking with the largestqiq\_\{i\}was selected\. Thus, Best\-of\-N always returned one intact sampled ranking and never recombined candidates across samples\. Remaining ties were resolved in favor of the lower sample index\. ##### Approval Voting Each mutation received one vote from every top\-40 ranking in which it appeared, sapproval\(m\)=∑i𝕀\[m∈Ri\]\.s\_\{\\mathrm\{approval\}\}\(m\)=\\sum\_\{i\}\\mathbb\{I\}\[m\\in R\_\{i\}\]\. Candidates were sorted by vote count, followed by their mean observed rank, best observed rank, and mutation string\. The 40 highest\-ranked candidates from the union were returned\. ##### Borda Fusion A mutation at rankrrreceived41−r41\-rpoints, giving sBorda\(m\)=∑i:m∈Ri\(41−ri\(m\)\)\.s\_\{\\mathrm\{Borda\}\}\(m\)=\\sum\_\{i:m\\in R\_\{i\}\}\\left\(41\-r\_\{i\}\(m\)\\right\)\. This score jointly reflects occurrence frequency and a linear preference for mutations placed near the top of each sampled ranking\. ##### Reciprocal Rank Fusion RRF assigned each mutation the score sRRF\(m\)=∑i:m∈Ri1k\+ri\(m\),s\_\{\\mathrm\{RRF\}\}\(m\)=\\sum\_\{i:m\\in R\_\{i\}\}\\frac\{1\}\{k\+r\_\{i\}\(m\)\}, wherek=10k=10\. Compared with Borda Fusion, reciprocal weighting places relatively greater emphasis on the highest\-ranked candidates\. ##### Hierarchical RRF To aggregate evidence shared by different substitutions at the same residue, we definedp\(m\)p\(m\)as the residue position of mutationmmandripos\(p\)r\_\{i\}^\{\\mathrm\{pos\}\}\(p\)as the first rank at which positionppoccurred inRiR\_\{i\}\. Position\-level support was spos\(p\)=∑i:∃m∈Ri,p\(m\)=p1k\+ripos\(p\),s\_\{\\mathrm\{pos\}\}\(p\)=\\sum\_\{i:\\exists m\\in R\_\{i\},\\ p\(m\)=p\}\\frac\{1\}\{k\+r\_\{i\}^\{\\mathrm\{pos\}\}\(p\)\}, and the final score was shier\(m\)=sRRF\(m\)\+λspos\(p\(m\)\),s\_\{\\mathrm\{hier\}\}\(m\)=s\_\{\\mathrm\{RRF\}\}\(m\)\+\\lambda\\,s\_\{\\mathrm\{pos\}\}\\\!\\left\(p\(m\)\\right\), withk=10k=10andλ=0\.35\\lambda=0\.35\. Candidates whose strings could not be parsed into residue positions retained their exact\-mutation RRF support but received no position\-level contribution\. For Borda, RRF, and Hierarchical RRF, ties were resolved by higher occurrence frequency, better mean observed rank, and then mutation string\. We used no per\-position diversity cap, so multiple substitutions at the same residue could appear in the final top\-40 ranking\. ## 10Extended Experimental Analysis ### 10\.1Evaluation Completeness and Missing\-Query Handling CategoryModelBaselineAdaptiveT1T2T3T4StatisticTotal123746729123LLMGPT\-6 Astra70005Claude Opus 5126619Gemini 3\.1 Pro00000Kimi K32321—GLM\-5\.20420—DeepSeek\-V4\-Pro00100AgentBiomni\+\+GPT\-6 Astra9111—Biomni\+\+Claude Opus 50410—AMixKit\+\+GPT\-6 Astra2000—AMixKit\+\+Claude Opus 51201—AMixKit\+\+AMix\-2\.10001—Table 12:Number of assay\-level queries replaced by the deterministic SHA\-256 random baseline during evaluation, for the baselinePFArenasetting and the multi\-round adaptive search setting\. A dash indicates that the corresponding task was not run for that model\.Despite repeated retries, some model–task pairs involving general\-purpose LLMs produced no prediction that satisfied the required output format and evaluation constraints\. These failures occurred particularly when the content\-filtering policies of certain general\-purpose LLMs restricted specific inputs, observed most frequently for Claude Opus 5, GPT\-6 Astra, Kimi K3 and GLM\-5\.2, with frequency varying across models and tasks\. ##### Fallback Substitutions To handle this issue, we distinguished successful predictions from missing assay\-level queries rather than assigning a score of zero to failed queries\. Missing queries were replaced by a deterministic SHA\-256 random baseline\. For each missing query, the evaluator collected all measured mutants in the corresponding ground\-truth table and computed its SHA\-256 digest\. The measured mutants were ordered lexicographically by these hexadecimal digests\. ForT1, the first 40 mutants in this order were used as the substitute prediction list\. ForT2–T4, the same procedure was applied independently to each candidate query, after which the evaluator applied the corresponding task\-specific ranking budget\. The numbers of substituted queries are summarized in[Table12](https://arxiv.org/html/2609.28921#S10.T12)\. Because this fallback ordering is independent of theDMS\_scoreranking requirement, the reported aggregate metrics combine model predictions from successfully completed queries with the random\-baseline lists for missing queries\. Consequently, a larger number of missing queries will generally depress the aggregate score\. ##### Matched\-query Analysis To isolate intrinsic model performance from the effects of incomplete outputs, we curated the multi\-round adaptive search evaluation of[Section5\.2\.1](https://arxiv.org/html/2609.28921#S5.SS2.SSS1)on all 123 ground\-truth assay queries to a common cohort of 113 assays, as shown in[Figure11](https://arxiv.org/html/2609.28921#S10.F11)\. The common cohort contained only assays for which every model produced a valid cumulative prediction prefix across all four rounds; no random substitution was used for these assays\. By contrast, the original multi\-round evaluation retained all 123 assays and used the deterministic random fallback whenever a model lacked a complete prediction\. At the fourth round, complete\-prefix coverage ranged from 114 assays for Claude Opus 5 and 118 assays for GPT\-6 Astra to all complete 123 assays for Gemini 3\.1 Pro and DeepSeek\-V4\-Pro\. Across the common cohort, all models showed progressive improvements in both NMS@4040and Recall@4040over successive rounds\. At round four, common\-cohort NMS@4040exceeded the corresponding full\-cohort value by0\.0180\.018for GPT\-6 Astra and by approximately0\.0150\.015for the other three models\. Recall@4040was boosted by approximately0\.00110\.0011to0\.00230\.0023across models\. The higher common\-cohort scores for Gemini 3\.1 Pro and DeepSeek\-V4\-Pro, which had complete coverage of all 123 assays, suggest that the 10 assays excluded from the matched cohort due to incomplete predictions from other models were more difficult on average\. These common\-cohort results therefore provide a cleaner comparison of intrinsic model quality, whereas the full\-cohort curves preserve performance under the original evaluation protocol\. Figure 11:Matched\-query analysis of round\-wise single\-mutant generation\. Marker lines show cumulative performance on all 123 assay queries, with deterministic random\-baseline substitutions used for missing queries\. Thin curves show performance on the common cohort of 113 assays completed by every model in all four rounds\. Shaded regions indicate the performance gap between the full and matched evaluations\. ### 10\.2Complementary Strengths and Failure Modes of LLMs and PLMs We tested whether PLMs and LLMs identify the same useful mutants and make the same ranking errors\. GPT\-6 Astra represented the LLM family, and VenusREM represented the PLM family\. The comparison had two parts:T1measured overlap and experimental quality among generated single\-mutant candidates, whereasT2–T4measured whether one model could correct extreme ranking errors made by the other\. This design separates complementary exploration inT1from complementary error correction in multi\-mutant ranking\. ##### T1: Single\-mutant Exploration GPT\-6 Astra and VenusREM selected largely different single\-mutant candidates, but their small consensus set was the most reliable\. As shown in[Figure12](https://arxiv.org/html/2609.28921#S10.F12), across 116 paired assays, VenusREM and GPT\-6 Astra produced 4,640 and 4,501 evaluable top\-4040selections, respectively, with 273 shared mutant–assay pairs\. The shared candidates therefore represented only a small fraction of the selections made by either model\. Nevertheless, the consensus set had the highest experimental quality: the median true rank score was 0\.79 for shared candidates, compared with 0\.66 for VenusREM\-only candidates and 0\.67 for GPT\-6 Astra\-only candidates, where 1 denotes the best experimental rank\. Thus, agreement was uncommon but informative: consensus candidates can provide a high\-confidence shortlist, whereas model\-specific candidates expand the explored sequence space\. Figure 12:Complementary single\-mutant selection inT1\.Left, overlap between evaluable top\-4040selections from VenusREM \(PLM\) and GPT\-6 Astra \(LLM\) across 116 paired assays\. Right, experimental fitness\-rank distributions of PLM\-only, shared and LLM\-only selections\. Rank scores were normalized within each assay from 1 \(best\) to 0 \(worst\)\. n denotes mutant–assay selections\. ##### T2–T4: Ranking Errors We next examined whether LLMs and PLMs make complementary ranking errors in the multi\-mutant tasks\. For each task, experimental fitness and model predictions were converted into within\-assay rank scores, with 1 denoting the best mutant\. In[Figure13](https://arxiv.org/html/2609.28921#S10.F13), thexx\-axis shows the true fitness rank score and theyy\-axis shows the predicted rank score\. The diagonal corresponds to perfect rank agreement, while vertical arrows connect the two predictions for the same mutant\. We defined an extreme error as placing a true bottom\-30% mutant in the predicted top 10% or a true top\-30% mutant in the predicted bottom 10%\. A correction was counted when the other model assigned the same mutant a substantially less extreme rank, moving it closer to its measured rank\. This paired, case\-level analysis distinguishes shared failures from model\-specific errors that can be rescued by the other model\. The direction\-specific counts further show that the corrected cases include both low\-fitness over\-rankings and high\-fitness under\-rankings, indicating that the complementarity is not restricted to a single type of ranking failure\. Figure 13:Complementary correction of extreme ranking errors by GPT\-6 Astra and VenusREM acrossT2–T4\.The left column shows mutants for which GPT\-6 Astra corrected extreme VenusREM ranking errors, and the right column shows the converse\. True fitness rank scores range from 0 \(worst\) to 1 \(best\), whereas predicted rank scores range from 0 \(bottom\) to 1 \(top\)\. Orange crosses and blue circles denote VenusREM and GPT\-6 Astra predictions, respectively; grey arrows connect predictions for the same mutant\. Shaded regions indicate true bottom\-30% mutants predicted in the top 10% or true top\-30% mutants predicted in the bottom 10%\. Central labels report the total number of mutants in each direction\.The error\-correction patterns inT2andT3indicate that VenusREM remains at least as reliable as GPT\-6 Astra when limited or localized experimental evidence is available, while they still provide complementary predictions\. InT2, VenusREM corrected 45 extreme GPT\-6 Astra errors, including 24 low\-fitness over\-rankings and 21 high\-fitness under\-rankings, whereas GPT\-6 Astra corrected 30 VenusREM errors, including 24 and 6 of these two types, respectively\. InT3, the two models corrected the same total number of extreme errors, with GPT\-6 Astra correcting 16 low\-fitness over\-rankings and 5 high\-fitness under\-rankings, and VenusREM correcting 11 and 10, respectively\. The availability of complete single\-mutant context reverses the direction of error correction in favor of GPT\-6 Astra\. InT4, GPT\-6 Astra corrected 39 extreme VenusREM errors, including 29 low\-fitness over\-rankings and 10 high\-fitness under\-rankings, whereas VenusREM corrected 24 GPT\-6 Astra errors, including 13 and 11 of these two types, respectively\. This reversal mirrors the higher aggregate performance of LLM\-based systems inT4, where measured fitness values for the component single mutations are available as additional evidence for ranking multi\-mutant candidates\. The results therefore support a context\-dependent shift in model utility rather than a universal superiority of either model family\. Taken together, VenusREM and GPT\-6 Astra exhibit complementary ranking errors acrossT2–T4, with each correcting both low\-fitness over\-rankings and high\-fitness under\-rankings made by the other\. These patterns motivate combined or agent\-mediated approaches, but this descriptive analysis does not establish that an ensemble would improve aggregate performance\.
Similar Articles
ProtSent: Protein Sentence Transformers
This article introduces ProtSent, a contrastive fine-tuning framework for protein language models that improves embedding quality for downstream tasks like remote homology detection and structural retrieval.
AutoProteinEngine: A Large Language Model Driven Agent Framework for Multimodal AutoML in Protein Engineering
Introduces AutoProteinEngine (AutoPE), an LLM-driven agent framework that enables biologists without deep learning expertise to perform multimodal AutoML for protein engineering via natural language, showing improvements over zero-shot and manual fine-tuning approaches.
RepBench: Compiling Benchmarks into Capability Representations for Large Language Models
RepBench is a benchmark-grounded data layer for representation probing in LLMs, built from 13,427 benchmark papers and 353 public datasets, yielding 46,149 probe texts across 94 capabilities. It evaluates capability vector transfer across models and readout methods.
Natural-Language-Guided Generator-Agnostic Shortlisting for Protein Binder Design
This paper explores using large language models to generate ranking policies for shortlisting protein binders from candidate pools, showing modest improvements over baseline methods in de novo design workflows.
Agentic BAIM-LLM Evaluation (ABLE): Benchmarking LLM Use of Protein Design Tools
ABLE is a benchmark for evaluating LLM agents' ability to use biological AI models like ProteinMPNN and AlphaFold3 in protein design workflows. It assesses 15 frontier models, finding that Claude Sonnet 4 and Gemini 3 Pro achieve the highest scores, while some models refuse all tasks.