Can Tabular In-Context Learners Generalize to Biomolecular Property Prediction?

arXiv cs.LG Papers

Summary

This paper investigates whether tabular in-context learning models, pretrained on synthetic causal tables, can generalize to predict biomolecular properties from limited labeled data. The authors find that these models are competitive for protein fitness regression but that representation choice is crucial for small-molecule classification.

arXiv:2606.31126v1 Announce Type: new Abstract: Predicting biomolecular properties from limited labeled data is a central bottleneck in protein engineering and small-molecule design. As strong pretrained encoders now supply rich fixed-length representations, the difficulty has shifted from representation learning to building a data-efficient predictor for the few-shot regime. Tabular foundation models such as TabPFN3 and TabICL are unlikely candidates for this role: they are in-context learners pretrained on synthetic tables drawn from random causal graphs, a generative prior with no obvious correspondence to the processes that produce protein sequences or molecular graphs. That this tabular, causal inductive bias should transfer to biomolecular data at all is unintuitive, yet we find it does. Treating each method as a predictor-representation pair, we evaluate across two domains. Over a fixed ESMC representation, tabular in-context learning is consistently competitive for protein fitness regression on ProteinGym and a diverse esterase dataset. For small-molecule classification with ECFP/RDKit descriptors, no single pairing dominates across TDC ADMET, MoleculeNet, FS-Mol, and DrugOOD; representation choice becomes a primary determinant, as expected when the predictor's own prior is indifferent to molecular structure. We conclude that tabular foundation models are strong performers on biomolecular prediction tasks, but that their performance depends strongly on the sequence or molecular representation used.
Original Article
View Cached Full Text

Cached at: 07/01/26, 05:33 AM

# 1 Introduction
Source: [https://arxiv.org/html/2606.31126](https://arxiv.org/html/2606.31126)
\\logo

CSIRO\\reporttitleCan Tabular In\-Context Learners Generalize to Biomolecular Property Prediction?\\reporttitlefontsize28pt\\reportsubtitle\\reportauthorsDavy Guan, Lu Zhang, Asiri Wijesinghe, Allen Zhu, He Zhao, Helen Power, F\. Hafna Ahmed, Andrew Warden, Cheng Soon Ong, Daniel M\. Steinberg\\reportnumber\\clientname\\clientcontact\\commercialinconfidence\\reportdateJune 30, 2026\\businessunit\\businessuniturlhttps://www\.csiro\.au/en/work\-with\-us/industries/technology\\contactnameDaniel M\. Steinberg\\contactphone\\contactemaildan\.steinberg@csiro\.au

\\openingpage

Citation

Guan D, Zhang L, Wijesinghe A, Zhu A, Zhao H, Power H, Ahmed FH, Warden A, Ong CS and Steinberg DM \(2026\) Can Tabular In\-Context Learners Generalize to Biomolecular Property Prediction? CSIRO, Australia\.

Copyright

© Commonwealth Scientific and Industrial Research Organisation 2026\. To the extent permitted by law, all rights are reserved and no part of this publication covered by copyright may be reproduced or copied in any form or by any means except with the written permission of CSIRO\.

Important disclaimer

CSIRO advises that the information contained in this publication comprises general statements based on scientific research\. The reader is advised and needs to be aware that such information may be incomplete or unable to be used in any specific situation\. No reliance or actions must therefore be made on that information without seeking prior expert professional, scientific and technical advice\. To the extent permitted by law, CSIRO \(including its employees and consultants\) excludes all liability to any person for any consequences, including but not limited to all losses, damages, costs, expenses and any other compensation, arising directly or indirectly from using this publication \(in part or in whole\) and any information or material contained in it\.

Acknowledgement of Country

CSIRO acknowledges the Traditional Owners of the lands, seas and waters of the areas that we live and work on across Australia and pays its respects to Elders past and present\.

###### Abstract

Predicting biomolecular properties from limited labeled data is a central bottleneck in protein engineering and small\-molecule design\. As strong pretrained encoders now supply rich fixed\-length representations, the difficulty has shifted from representation learning to building a data\-efficient predictor for the few\-shot regime\. Tabular foundation models such as TabPFN3 and TabICL are unlikely candidates for this role: they are in\-context learners pretrained on synthetic tables drawn from random causal graphs, a generative prior with no obvious correspondence to the processes that produce protein sequences or molecular graphs\. That this tabular, causal inductive bias should transfer to biomolecular data at all is unintuitive, yet we find it does\. Treating each method as a predictor–representation pair, we evaluate across two domains\. Over a fixed ESMC representation, tabular in\-context learning is consistently competitive for protein fitness regression on ProteinGym and a diverse esterase dataset\. For small\-molecule classification with ECFP/RDKit descriptors, no single pairing dominates across TDC ADMET, MoleculeNet, FS\-Mol, and DrugOOD; representation choice becomes a primary determinant, as expected when the predictor’s own prior is indifferent to molecular structure\. We conclude that tabular foundation models are strong performers on biomolecular prediction tasks, but that their performance depends strongly on the sequence or molecular representation used\.

Predicting the functional consequences of biomolecular data from limited labels is a central problem in protein engineering, variant\-effect prediction, and small\-molecule design\. Deep mutational scanning \(DMS\) benchmarks such as ProteinGym have greatly improved the scale and rigor of evaluation, but each assay still induces a distinct low\-data supervised task in which only a small number of experimentally measured variants are available for a new protein or condition\(Notinet al\.,[2023](https://arxiv.org/html/2606.31126#bib.bib8); Elnaggaret al\.,[2022](https://arxiv.org/html/2606.31126#bib.bib29)\)\. Recently, large biological foundation models have made sequence representation substantially stronger: protein language models trained on evolutionary corpora provide rich fixed\-length embeddings that can be reused across downstream tasks\(Elnaggaret al\.,[2022](https://arxiv.org/html/2606.31126#bib.bib29); Linet al\.,[2023](https://arxiv.org/html/2606.31126#bib.bib22); Hayeset al\.,[2025](https://arxiv.org/html/2606.31126#bib.bib9)\)\. This shifts part of the modeling bottleneck from representation learning to the design of a data\-efficient predictor that can exploit those embeddings in the few\-shot regime\.

Tabular foundation models offer a compelling approach to this bottleneck\. Rather than training a new predictor from scratch for each dataset, they amortize Bayesian\-style prediction offline over synthetic tabular tasks and perform inference by conditioning on labeled support examples at test time, without task\-specific gradient updates\(Hollmannet al\.,[2023](https://arxiv.org/html/2606.31126#bib.bib1),[2025](https://arxiv.org/html/2606.31126#bib.bib2); Quet al\.,[2025](https://arxiv.org/html/2606.31126#bib.bib3),[2026](https://arxiv.org/html/2606.31126#bib.bib4)\)\. Recent work has extended this paradigm to larger contexts, broader task families, and more realistic data regimes\. Science\-domain evaluations in metabolomics, analytical chemistry, and materials informatics suggest that these models can transfer effectively once domain\-specific structure has been encoded as feature vectors\(Wuet al\.,[2026](https://arxiv.org/html/2606.31126#bib.bib10); Granittoet al\.,[2026](https://arxiv.org/html/2606.31126#bib.bib11); Liet al\.,[2026](https://arxiv.org/html/2606.31126#bib.bib12)\)\.

It remains an open question whether this transfer holds for two domains with distinct representation challenges and output types: protein fitness regression and small\-molecule property classification\. Protein fitness prediction requires mapping high\-dimensional sequence embeddings to continuous fitness values across hundreds of structurally diverse assays with as few as𝒪​\(10\)\\mathcal\{O\}\(10\)labeled examples\. Small\-molecule property classification poses a different challenge: representations range from fixed fingerprint vectors to learned molecular graph embeddings, and evaluation must simultaneously capture in\-distribution accuracy, low\-shot sample efficiency, and out\-of\-distribution generalization across scaffold, assay, and molecular\-size shifts\.

We study both settings as an evaluation problem using existing pretrained models and fingerprinting techniques\. For protein fitness regression, we encode sequences with ESMC\(Hayeset al\.,[2025](https://arxiv.org/html/2606.31126#bib.bib9)\)and benchmark TabICL and TabPFN3 against supervised baselines across ProteinGym and a newly published diverse esterase family dataset\(Ahmedet al\.,[2026](https://arxiv.org/html/2606.31126#bib.bib25)\)\. For small\-molecule property classification, we evaluate learner–representation pairs, pairing tabular models with molecular descriptor views and comparing against graph\-based baselines\. We benchmark across TDC ADMET\(Huanget al\.,[2021](https://arxiv.org/html/2606.31126#bib.bib15)\), MoleculeNet\(Wuet al\.,[2018](https://arxiv.org/html/2606.31126#bib.bib16)\), FS\-Mol\(Stanleyet al\.,[2021](https://arxiv.org/html/2606.31126#bib.bib14)\), and DrugOOD\(Jiet al\.,[2023](https://arxiv.org/html/2606.31126#bib.bib17)\), with DrugOOD providing an explicit OOD generalization test\.

Our contribution is primarily empirical and methodological\. We provide a benchmarked audit of tabular foundation models in scientific prediction settings that have not yet been evaluated systematically, and we do so under protocols designed to support cautious, track\-appropriate claims\. We keep feature pipelines fixed where possible, separate few\-shot from full\-train conclusions, and treat OOD behavior, official ProteinGym holdouts, benchmark coverage, and support\-set sensitivity as first\-class parts of the evaluation\.

## 2Background and Related Work

The prior\-data fitted network \(PFN\) framework trains a transformer offline on synthetic tabular tasks so that prediction can be performed at inference time by conditioning on labelled in\-context examples, without task\-specific gradient updates\(Mülleret al\.,[2022](https://arxiv.org/html/2606.31126#bib.bib23)\)\. TabPFN established this formulation for small classification problems and showed that amortized inference over synthetic Bayesian priors can match or exceed carefully tuned baselines on real benchmarks\(Hollmannet al\.,[2023](https://arxiv.org/html/2606.31126#bib.bib1),[2025](https://arxiv.org/html/2606.31126#bib.bib2)\)\. The 2026 release\(Grinsztajnet al\.,[2026](https://arxiv.org/html/2606.31126#bib.bib26)\), referred to throughout this paper as TabPFN3, extends the framework with improved scalability and regression support\. TabICL scales the paradigm through a two\-stage architecture: a first stage applies column\-then\-row attention to compress each sample into a fixed\-dimension row embedding, and a second stage applies a transformer across row embeddings for in\-context prediction\(Quet al\.,[2025](https://arxiv.org/html/2606.31126#bib.bib3),[2026](https://arxiv.org/html/2606.31126#bib.bib4)\)\. We refer to the TabICL v2 implementation used in this study as TabICL for readability\.

ProteinGym standardizes evaluation across 217 DMS assays spanning diverse proteins and fitness phenotypes, enabling systematic comparison across assay types and substitution depths\(Notinet al\.,[2023](https://arxiv.org/html/2606.31126#bib.bib8)\)\. The dominant supervised strategy pairs a large pretrained sequence encoder with a lightweight task\-specific predictor\. Protein language models including ProtTrans\(Elnaggaret al\.,[2022](https://arxiv.org/html/2606.31126#bib.bib29)\), ESM2\(Linet al\.,[2023](https://arxiv.org/html/2606.31126#bib.bib22)\), and ESMC\(Hayeset al\.,[2025](https://arxiv.org/html/2606.31126#bib.bib9)\)are trained on evolutionary sequence corpora so that their embeddings capture structural and functional constraints learned from natural variation\.

Small\-molecule property prediction spans ADMET profiling, activity prediction, and out\-of\-distribution generalization across chemical series\. MoleculeNet\(Wuet al\.,[2018](https://arxiv.org/html/2606.31126#bib.bib16)\), TDC ADMET\(Huanget al\.,[2021](https://arxiv.org/html/2606.31126#bib.bib15)\), FS\-Mol\(Stanleyet al\.,[2021](https://arxiv.org/html/2606.31126#bib.bib14)\), and DrugOOD\(Jiet al\.,[2023](https://arxiv.org/html/2606.31126#bib.bib17)\)provide complementary regimes covering standard supervised settings, few\-shot learning with task\-level variation, and systematic distribution shift across scaffolds, assays, and molecular sizes\. Fixed descriptors such as ECFP fingerprints\(Rogers and Hahn,[2010](https://arxiv.org/html/2606.31126#bib.bib18)\)and RDKit features expose tabular feature vectors, while graph models such as ChemProp\(Yanget al\.,[2019](https://arxiv.org/html/2606.31126#bib.bib13)\)and CheMeleon\-initialized ChemProp\(Burnset al\.,[2025](https://arxiv.org/html/2606.31126#bib.bib19)\)model molecular structure directly\.

## 3Using Tabular ICL Models for Biomolecules

Across both domains we follow a single recipe\. A frozen domain encoder maps each biomolecule to a fixed\-length feature vector, and that vector becomes one row of an input table for a tabular in\-context learner\. The predictors, TabICL and TabPFN3, predict the label of a query molecule by conditioning on labeled support rows in context\. Protein fitness is treated as regression and reported by MSE and Spearman rank correlation\. Small\-molecule property prediction is treated as binary classification and reported by ROC\-AUC\.

#### Protein representations\.

Protein sequences are encoded with ESMC\(Hayeset al\.,[2025](https://arxiv.org/html/2606.31126#bib.bib9)\), a 300M\-parameter protein language model and successor to ESM2\(Linet al\.,[2023](https://arxiv.org/html/2606.31126#bib.bib22)\)\. We use the resulting fixed\-length 960\-dimensional per\-sequence embedding as the default feature vector\. The same embedding is fed unchanged to every protein method so that comparisons isolate the predictor under a fixed representation\. For the largest ProteinGym assays where full\-feature inference exceeded available memory or wall\-time limits, we additionally report PCA\-compressed ESMC rescue runs; these rows are marked explicitly in the artifact notes and used only to complete otherwise infeasible assays\.

#### Molecule features\.

For small molecules, tabular learners receive either ECFP/Morgan fingerprints\(Rogers and Hahn,[2010](https://arxiv.org/html/2606.31126#bib.bib18)\), RDKit two\-dimensional physicochemical descriptors, or their concatenation\. The descriptor is treated as an integral part of the model pair rather than a separable preprocessing step, because a tabular learner with ECFP features and the same learner with RDKit descriptors can behave differently\. As structural alternatives to fixed descriptors, graph baselines consume molecular graphs directly through ChemProp\(Yanget al\.,[2019](https://arxiv.org/html/2606.31126#bib.bib13)\)trained from scratch and a foundation\-model\-initialized variant that fine\-tunes ChemProp from the CheMeleon checkpoint\(Burnset al\.,[2025](https://arxiv.org/html/2606.31126#bib.bib19)\)\.

## 4Protein Prediction Experiments

We evaluate TabICL and TabPFN3 in two protein prediction settings\. The first is ProteinGym’s 217 DMS tasks, and the second is the PpEST esterase\-family dataset introduced inAhmedet al\.\([2026](https://arxiv.org/html/2606.31126#bib.bib25)\)\. We test these models in full\-training tasks and few\-shot settings with 8 to 64 training samples to approximate low\-throughput wet\-lab measurement regimes\.

#### Datasets\.

ProteinGym\(Notinet al\.,[2023](https://arxiv.org/html/2606.31126#bib.bib8)\)is a large\-scale collection of DMS assays curated to benchmark computational models of protein fitness\. Each assay quantifies the effect of amino acid substitutions on a biologically relevant property\. We use the 217 supervised\-substitution assays available in the staged evaluation files and report both random 5\-fold results and the official random, modulo, and contiguous holdout schemes\. In the random scheme, each mutation is randomly assigned to one of five different folds\. For the modulo scheme, the position of a mutation is assigned to a particular fold \(modulo 5\)\. In the contiguous scheme, 5 segments of contiguous positions are obtained by splitting the sequence, and mutations are assigned to segments based on position\.PpESTis a diverse esterase dataset containing 1513 ancestrally related sequences after excluding approximately 600 sequences partially selected using ML methods\. It includes catalytic reaction rates for three ester substrates and thermostability measurements at two temperatures\(Ahmedet al\.,[2026](https://arxiv.org/html/2606.31126#bib.bib25)\)\.

#### Baselines\.

We compare TabICL and TabPFN3 with ridge regression, RBF sampler, HistGradientBoostingRegressor \(HGBR\)\(Pedregosaet al\.,[2011](https://arxiv.org/html/2606.31126#bib.bib27)\), and fine\-tuned ESMC with a task\-specific MLP regression head\. All baselines receive the same ESMC embeddings as inputs\. A compatible retrained sequence CNN for PpEST is reported as an appendix check rather than inserted into the main comparison\.

#### Evaluation metrics\.

Protein fitness prediction is evaluated by mean squared error \(MSE\) and the Spearman rank correlation coefficient\. Spearman is the primary ranking metric for comparison with ProteinGym\-style variant\-effect evaluation, while MSE captures calibrated regression error on the assay scale\. For few\-shot learning\-curve plots, we compute raw MSE and Spearman at each support size and plot a monotone best\-so\-far envelope over increasing support sizes:

MSEk\+\\displaystyle\\textrm\{MSE\}^\{\+\}\_\{k\}=min⁡\{MSEk,MSEk−1\+\},MSE1\+=MSE1,\\displaystyle=\\min\\left\\\{\\textrm\{MSE\}\_\{k\},~\\textrm\{MSE\}^\{\+\}\_\{k\-1\}\\right\\\},\\quad\\textrm\{MSE\}^\{\+\}\_\{1\}=\\textrm\{MSE\}\_\{1\},\(1\)Spearmank\+\\displaystyle\\textrm\{Spearman\}^\{\+\}\_\{k\}=max⁡\{Spearmank,Spearmank−1\+\},Spearman1\+=Spearman1\.\\displaystyle=\\max\\left\\\{\\textrm\{Spearman\}\_\{k\},~\\textrm\{Spearman\}^\{\+\}\_\{k\-1\}\\right\\\},\\quad\\textrm\{Spearman\}^\{\+\}\_\{1\}=\\textrm\{Spearman\}\_\{1\}\.\(2\)Herekkindexes the few\-shot training size,n∈\{n1,…,nk,…,nK\}n\\in\\\{n\_\{1\},\\ldots,n\_\{k\},\\ldots,n\_\{K\}\\\}andnk<nk\+1n\_\{k\}<n\_\{k\+1\}\. This plotting convention avoids visually overemphasizing non\-nested support\-set noise; raw means and standard deviations are retained in source tables\.

#### Experimental settings\.

In the full\-train regime, about 20% of samples are selected for testing\. For ProteinGym, we use five\-fold cross\-validation and report the mean score averaged across assays\. For PpEST, we randomly select 20% of samples as the test set\. In the few\-shot regime, training sizes are 8, 16, 32, and 64\. We run 30 repeated support\-set draws for each size using stratified quantile\-bin sampling to encourage coverage of the observed fitness range\.

### 4\.1ProteinGym Results

Table[1](https://arxiv.org/html/2606.31126#S4.T1)reports model performance on ProteinGym under the random 5\-fold cross\-validation protocol used for the main reproduced comparison\.

Table 1:ProteinGym random 5\-fold full\-data summary across 217 assays\. Values are rounded to three decimals for ProteinGym\-style reporting\.TabPFN3 achieves the strongest aggregate result, with mean Spearman 0\.767 and mean MSE 0\.351 across 217 assays\. TabICL remains close, with mean Spearman 0\.753 and mean MSE 0\.376\. HistGradientBoosting is the strongest classical baseline, followed by Ridge and FT ESM\. The RBF sampler fails in this setting, yielding near\-zero Spearman and the highest MSE, indicating that the median\-heuristic feature scaling is poorly calibrated for the 960\-dimensional ESMC embedding space\.

![Refer to caption](https://arxiv.org/html/2606.31126v1/x1.png)Figure 1:ProteinGym random 5\-fold method comparison across 217 assays\. Bars report mean Spearman correlation and mean MSE across assays\. TabPFN3 is strongest, while TabICL remains close and above the classical baselines\.Figure[1](https://arxiv.org/html/2606.31126#S4.F1)summarizes the same random 5\-fold comparison visually\. The two largest assays, HIS7\_YEAST\_Pokusaeva\_2019 and SPG1\_STRSG\_Olson\_2014, required special handling because full\-feature inference exceeded practical memory and time limits on the available hardware\. For those assays, the complete TabPFN3 row uses PCA128 features with 8 estimators and the complete TabICL row uses PCA rescue\.

![Refer to caption](https://arxiv.org/html/2606.31126v1/x2.png)Figure 2:Assay\-level paired difference in validation performance between TabPFN3 and TabICL on ProteinGym\. Positive values indicate higher Spearman correlation for TabPFN3 on the same assay\.The paired assay\-level comparison in Figure[2](https://arxiv.org/html/2606.31126#S4.F2)shows that the aggregate improvement is broad but modest rather than driven by a small number of isolated wins\. This supports treating TabPFN3 as the stronger ProteinGym configuration while still describing TabICL as close and reproducible\.

Table 2:ProteinGym supervised\-substitution performance by official holdout scheme across 217 assays\. Standard deviations are computed across assays\.Table[2](https://arxiv.org/html/2606.31126#S4.T2)extends the comparison to the official random, modulo, and contiguous ProteinGym fold schemes\. The harder modulo and contiguous splits reduce performance for both models, indicating that random\-split numbers should not be treated as the only estimate of ProteinGym generalization\. Across official schemes, TabPFN3 is modestly but consistently ahead of TabICL, while both methods retain competitive rank\-correlation performance over fixed ESMC embeddings\.

![Refer to caption](https://arxiv.org/html/2606.31126v1/x3.png)Figure 3:ProteinGym performance across official random, modulo, and contiguous holdout schemes\. TabPFN3 is modestly higher than TabICL in each scheme, but both methods drop on the harder modulo and contiguous splits\.Figure[3](https://arxiv.org/html/2606.31126#S4.F3)visualizes the official split comparison and makes the generalization gap explicit\. The relative ordering is stable, but absolute performance is strongly split\-dependent\.

Table 3:ProteinGym external context\. TabPFN3 and TabICL rows are local random 5\-fold results; public leader\-board values are included only for context until external ProteinGym evaluation is complete\.Table[3](https://arxiv.org/html/2606.31126#S4.T3)places the local TabPFN3 and TabICL random\-split aggregates alongside top\-ranked public ProteinGym leader\-board methods\. We treat this table as external context rather than as a new public ranking claim, because the TabICL and TabPFN3 submission packages must still be evaluated through ProteinGym’s external process\.

![Refer to caption](https://arxiv.org/html/2606.31126v1/x4.png)Figure 4:ProteinGym few\-shot performance as support\-set size increases\. Curves show best\-so\-far task\-averaged trends for Spearman and MSE for TabPFN3, TabICL, and supervised baselines\.Table 4:ProteinGym few\-shot summary at the smallest and largest support sizes\. Values are mean±\\pmstandard deviation across assays; MSE is computed on the assay\-standardized target scale\. Thek=64k=64aggregate contains 216 assays because one assay has only 63 variants\.Figure[4](https://arxiv.org/html/2606.31126#S4.F4)and Table[4](https://arxiv.org/html/2606.31126#S4.T4)show best\-so\-far Spearman and MSE as a function of the number of training samplesn∈\{8,16,32,64\}n\\in\\\{8,16,32,64\\\}\. TabPFN3 is strongest at both endpoint support sizes, improving from Spearman 0\.361 and MSE 0\.888 atk=8k=8to Spearman 0\.537 and MSE 0\.639 atk=64k=64\. TabICL remains close, with Spearman 0\.327 and MSE 0\.995 atk=8k=8, and Spearman 0\.505 and MSE 0\.708 atk=64k=64\. The learning curves support the interpretation that tabular predictors can exploit local structure in ESMC embedding space from very limited labeled examples, while the table records cross\-assay variability that is intentionally omitted from the main curves for legibility\.

![Refer to caption](https://arxiv.org/html/2606.31126v1/x5.png)Figure 5:ProteinGym assay\-size diagnostic\. Mean Spearman is plotted against assay variant count for TabPFN3 and TabICL\.Figure[5](https://arxiv.org/html/2606.31126#S4.F5)shows that assay size alone does not explain the variation in full\-data performance, motivating more detailed diagnostics based on representation geometry and task structure\.

### 4\.2PpEST Results

Table[5](https://arxiv.org/html/2606.31126#S4.T5)reports full\-training results on the five PpEST targets\.

Table 5:Full\-training performance on the five PpEST endpoints\. “Sp\.” denotes Spearman; FT ESM denotes fine\-tuned ESM\. Lower MSE and higher Spearman indicate better performance\.TabICL and TabPFN3 provide the strongest overall results among the evaluated methods\. TabICL has the lowest MSE on Octanoate, Butyrate, and 60∘C Butyrate, while TabPFN3 has the lowest MSE on Acetate and 90∘C Butyrate\. TabPFN3 has the highest Spearman correlation on four of five assay endpoints, while TabICL remains strongest on 60∘C Butyrate\. Fine\-tuned ESM is consistently the strongest non\-tabular baseline\.

![Refer to caption](https://arxiv.org/html/2606.31126v1/x6.png)Figure 6:Few\-shot PpEST performance measured by Spearman correlation\. Bold curves are monotone best\-so\-far envelopes over support\-set size, while hollow markers show raw means\.Figure[6](https://arxiv.org/html/2606.31126#S4.F6)shows that TabICL and TabPFN3 exhibit strong rank preservation in the few\-shot regime\. Across many PpEST endpoints, one of the two tabular foundation models achieves the highest or near\-highest Spearman correlation, indicating that they effectively preserve the ordering of protein fitness values even when only a small number of labeled examples are available\.

![Refer to caption](https://arxiv.org/html/2606.31126v1/x7.png)Figure 7:Few\-shot PpEST performance measured by MSE\. Bold curves are monotone best\-so\-far envelopes over support\-set size, while hollow markers show raw means\.Figure[7](https://arxiv.org/html/2606.31126#S4.F7)presents the same few\-shot runs measured by MSE\. The bold curves show the monotone best\-so\-far envelope used for visual comparison, while the hollow markers show the raw mean at each support size\. This preserves the expected improvement trend as more labels are added without hiding the fact that individual support\-set draws can fluctuate in the low\-data regime\. Taken together, PpEST supports the broader protein conclusion: fixed ESMC embeddings paired with tabular foundation models are effective for diverse esterase\-family prediction, with TabICL and TabPFN3 trading off small differences in regression error and rank preservation\.

## 5Small\-Molecule Prediction Experiments

We evaluate whether the behavior of tabular in\-context learners transfers from protein sequence embeddings to small\-molecule classification\. The molecular benchmark covers four public benchmark families: TDC ADMET, MoleculeNet, FS\-Mol, and DrugOOD\-LBAP\-core\-IC50\. We report learner–representation pairs rather than pure architectures: TabICL, TabPFN3, and XGBoost use fixed molecular descriptors, whereas ChemProp uses molecular graphs and ChemProp\+CheMeleon uses graph foundation\-model fine\-tuning\.

#### Datasets\.

TDC ADMET\(Huanget al\.,[2021](https://arxiv.org/html/2606.31126#bib.bib15)\)contains 13 curated small\-molecule classification endpoints for absorption, distribution, metabolism, excretion, and toxicity prediction\.MoleculeNet\(Wuet al\.,[2018](https://arxiv.org/html/2606.31126#bib.bib16)\)contributes a broad set of molecular classification endpoints, including BBBP, BACE, HIV, ClinTox, SIDER, Tox21, ToxCast, MUV, and PCBA\. The current classification registry contains 806 MoleculeNet endpoints; full\-train TabPFN3 excludes PCBA endpoints in this report snapshot because those jobs were resource\-infeasible under the available rescue attempts\.FS\-Mol\(Stanleyet al\.,[2021](https://arxiv.org/html/2606.31126#bib.bib14)\)contains 157 few\-shot molecular property tasks\.DrugOOD\-LBAP\-core\-IC50\(Jiet al\.,[2023](https://arxiv.org/html/2606.31126#bib.bib17)\)contains three OOD classification tasks defined by assay, scaffold, and molecular\-size shifts\.

Table 6:Small\-molecule benchmark inventory\. MoleculeNet full\-train TabPFN3 excludes PCBA endpoints in this snapshot after high\-memory rescue attempts proved resource\-infeasible\.Table[6](https://arxiv.org/html/2606.31126#S5.T6)summarizes the molecular benchmark inventory\. The key distinction from the protein setting is that the molecular experiments combine several benchmark families with different task construction rules and split semantics, so conclusions are reported at the benchmark\-family level rather than collapsed into a single global molecular score\.

#### Baselines and metrics\.

The fixed\-feature tabular pairs are TabICL, TabPFN3, and XGBoost\(Chen and Guestrin,[2016](https://arxiv.org/html/2606.31126#bib.bib28)\), each paired with ECFP, RDKit, or ECFP\+RDKit features\. The graph pairs are ChemProp trained from scratch and ChemProp\+CheMeleon\. The primary metric is ROC\-AUC\. For TDC ADMET, MoleculeNet, and FS\-Mol, we report benchmark test ROC\-AUC\. For DrugOOD, the main OOD metric is OOD\-test ROC\-AUC, and we also report the ID/OOD generalization gap, defined as ID\-test ROC\-AUC minus OOD\-test ROC\-AUC\. Undefined ROC\-AUC values can arise when an evaluation split contains a single class; these rows are treated as undefined split/metric cases rather than model failures\. For few\-shot learning\-curve plots, we use the monotone best\-so\-far envelope

AUCk\+\\displaystyle\\textrm\{AUC\}^\{\+\}\_\{k\}=max⁡\{AUCk,AUCk−1\+\},AUC1\+=AUC1,\\displaystyle=\\max\\left\\\{\\textrm\{AUC\}\_\{k\},~\\textrm\{AUC\}^\{\+\}\_\{k\-1\}\\right\\\},\\quad\\textrm\{AUC\}^\{\+\}\_\{1\}=\\textrm\{AUC\}\_\{1\},\(3\)wherekkindexes the few\-shot training size,n∈\{n1,…,nk,…,nK\}n\\in\\\{n\_\{1\},\\ldots,n\_\{k\},\\ldots,n\_\{K\}\\\}andnk<nk\+1n\_\{k\}<n\_\{k\+1\}\. Raw support\-size means are retained in source tables\.

#### Experimental settings\.

In the few\-shot regime, support sets are sampled from the training split only, while validation and test splits remain fixed\. We use support sizesn∈\{8,16,32,64,128,256,512\}n\\in\\\{8,16,32,64,128,256,512\\\}where feasible and five random seeds by default\. In full\-train experiments, all available training examples are used where feasible\. DrugOOD uses train\-once/evaluate\-many semantics: each fitted model is evaluated on the ID and OOD test splits without split\-specific retraining artifacts\.

## 6Small\-Molecule Results

![Refer to caption](https://arxiv.org/html/2606.31126v1/x8.png)Figure 8:Few\-shot small\-molecule learning curves\. The x\-axis is the number of labeled support molecules\. The y\-axis is test ROC\-AUC for TDC ADMET, MoleculeNet, and FS\-Mol, and OOD\-test ROC\-AUC for DrugOOD\. Curves show monotone best\-so\-far envelopes of task\-averaged performance across support sizes; raw means are retained in source tables\.Figure[8](https://arxiv.org/html/2606.31126#S6.F8)shows that fixed\-feature tabular foundation models remain competitive in low\-label molecular prediction, but the ranking depends strongly on both benchmark family and representation\. At the largest support size shown, TabICL with RDKit descriptors is strongest on MoleculeNet, TabPFN3 with ECFP\+RDKit is strongest on FS\-Mol, and TabICL RDKit and TabPFN3 ECFP\+RDKit are effectively tied on TDC ADMET\. On DrugOOD, the best few\-shot OOD ROC\-AUC values are close among TabICL ECFP\+RDKit, TabICL ECFP, and TabPFN3 ECFP\+RDKit\.

![Refer to caption](https://arxiv.org/html/2606.31126v1/x9.png)Figure 9:Full\-train small\-molecule benchmark summary\. Bars show mean ROC\-AUC across tasks; DrugOOD uses OOD\-test ROC\-AUC and other benchmark families use benchmark test ROC\-AUC\.In the full\-train regime, TabPFN3 and TabICL are closely matched on the descriptor\-based tasks where both are available\. Figure[9](https://arxiv.org/html/2606.31126#S6.F9)shows that TabPFN3 RDKit slightly exceeds TabICL RDKit on TDC ADMET in the current aggregate, while TabPFN3 ECFP\+RDKit is strongest on FS\-Mol\. MoleculeNet full\-train TabPFN3 should be interpreted cautiously because PCBA endpoints could not be completed\.

Table 7:Best full\-train pair by benchmark family in the current report snapshot\. MoleculeNet TabPFN3 coverage excludes PCBA endpoints and should not be described as complete MoleculeNet coverage\.Table[7](https://arxiv.org/html/2606.31126#S6.T7)reinforces the same conclusion: no single learner–representation pair dominates all molecular benchmark families\. TabPFN3 RDKit is strongest on TDC ADMET and the completed non\-PCBA MoleculeNet snapshot, TabPFN3 ECFP\+RDKit is strongest on FS\-Mol, and ChemProp\+CheMeleon is strongest on DrugOOD OOD ROC\-AUC\.

![Refer to caption](https://arxiv.org/html/2606.31126v1/x10.png)Figure 10:DrugOOD full\-train ID/OOD generalization gap\. Bars report ID\-test ROC\-AUC minus OOD\-test ROC\-AUC for each model pair on assay, scaffold, and molecular\-size shifts\. Larger positive values indicate a larger performance drop under OOD shift\.DrugOOD shows that in\-distribution performance is not sufficient to characterize molecular generalization\. Figure[10](https://arxiv.org/html/2606.31126#S6.F10)shows positive ID/OOD gaps across model families, and the magnitude of the gap depends on both representation and pretraining regime\.

![Refer to caption](https://arxiv.org/html/2606.31126v1/x11.png)Figure 11:Paired few\-shot comparison of TabPFN3 and TabICL on molecules\. Positive values indicate higher ROC\-AUC for TabPFN3 after matching benchmark family, task, molecular representation, support\-set size, and seed\.The paired TabPFN3–TabICL comparison in Figure[11](https://arxiv.org/html/2606.31126#S6.F11)gives a more controlled view than aggregate ranking alone\. TabPFN3 does not uniformly dominate TabICL: the paired deltas are small overall and change sign by benchmark family and representation\. These results support a cautious molecular conclusion: newer tabular foundation models can improve some few\-shot settings, but representation choice and task distribution remain first\-order determinants of performance\.

## 7Discussion

Our results suggest tabular in\-context learning is a practical interface between pretrained biological representations and property prediction from only low numbers of labels\. The strongest results arise when the representation already exposes task\-relevant structure across proteins and small molecules: ESMC embeddings for protein fitness and RDKit/ECFP\-style descriptors for molecular classification\. This supports the view that tabular foundation models should be evaluated as learner–representation pairs rather than architecture\-only methods\. Mechanistically, this pattern is consistent with an interpolation view of ICL: when nearby points in representation space have similar labels or rankings, an in\-context tabular predictor can exploit local support\-set geometry without updating model weights\.

The protein experiments show that tabular in\-context learners can be highly competitive on fixed ESMC embeddings, particularly in few\-shot and full\-training summaries of ProteinGym\. The official ProteinGym modulo and contiguous schemes are harder than random splits for both TabICL and TabPFN3, so split choice materially affects the strength of the claim\. PpEST provides a complementary diverse\-sequence test case in which performance varies across assay conditions: TabICL and TabPFN3 trade off small differences in regression error and rank preservation, and both are stronger than the non\-tabular baselines in the main comparison\.

The small\-molecule experiments show a related but less uniform pattern\. TabICL and TabPFN3 are competitive with XGBoost and graph models in several few\-shot settings, but no single model dominates across TDC ADMET, MoleculeNet, FS\-Mol, and DrugOOD\. Representation choice is often as important as learner choice: RDKit, ECFP, and their concatenation can change rankings within the same model family\. DrugOOD also shows that in\-distribution performance does not fully predict OOD behavior under assay, scaffold, or molecular\-size shifts\.

#### Limitations\.

Several limitations should remain explicit\. First, the ProteinGym submission packages for TabICL and TabPFN3 are complete for the fold columns present in the staged local ProteinGym CV files, but two large assays require PCA rescue rows rather than the same full\-feature configuration as the remaining assays\. Second, the hardest ProteinGym modulo and contiguous holdout schemes reduce performance for both tabular in\-context learners, so random\-split numbers should not be interpreted as the only estimate of generalization\. Third, MoleculeNet full\-train TabPFN3 excludes PCBA endpoints because those jobs were resource\-infeasible under available rescue attempts\. Fourth, molecular ROC\-AUC can be undefined for single\-class evaluation splits, especially in highly imbalanced endpoints; these rows should be treated as split/metric limitations rather than model failures\. Finally, support\-set sensitivity is an analysis axis in its own right and should not be mixed into headline few\-shot curves\.

## 8Conclusion

Tabular foundation models are viable predictors for biomolecular property prediction, but they should not be evaluated as representation\-free black boxes\. Their strongest use is as part of a learner–representation pair: ESMC plus TabPFN3 or TabICL for protein regression, and descriptor\-specific TabPFN3 or TabICL pairings for molecular classification\. This framing makes the positive result more credible and the limitations easier to audit\.

## References

- F\. H\. Ahmed, A\. Bender, A\. Wijesinghe, A\. Zhu, L\. Zhang, L\. Gebbie, A\. Marsh, C\. Ishitate, W\. Holdsworth, C\. Jones, A\. C\. Warden, H\. Power, C\. S\. Ong, D\. Steinberg, and R\. E\. Speight \(2026\)Data\-efficient exploration of enzyme function using family\-specific machine learning\.bioRxiv\.External Links:[Document](https://dx.doi.org/10.64898/2026.06.02.729712)Cited by:[§1](https://arxiv.org/html/2606.31126#S1.p4.1),[§4](https://arxiv.org/html/2606.31126#S4.SS0.SSS0.Px1.p1.1),[§4](https://arxiv.org/html/2606.31126#S4.p1.1)\.
- Descriptor\-based foundation models for molecular property prediction\.Note:Preprint, arXiv:2506\.15792Cited by:[§2](https://arxiv.org/html/2606.31126#S2.p3.1),[§3](https://arxiv.org/html/2606.31126#S3.SS0.SSS0.Px2.p1.1)\.
- T\. Chen and C\. Guestrin \(2016\)XGBoost: a scalable tree boosting system\.InProceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining,KDD ’16,New York, NY, USA,pp\. 785–794\.External Links:ISBN 9781450342322,[Document](https://dx.doi.org/10.1145/2939672.2939785)Cited by:[§5](https://arxiv.org/html/2606.31126#S5.SS0.SSS0.Px2.p1.4)\.
- A\. Elnaggar, M\. Heinzinger, C\. Dallago, G\. Rehawi, Y\. Wang, L\. Jones, T\. Gibbs, T\. Feher, C\. Angerer, M\. Steinegger, D\. Bhowmik, and B\. Rost \(2022\)ProtTrans: toward understanding the language of life through self\-supervised learning\.IEEE Transactions on Pattern Analysis and Machine Intelligence44\(10\),pp\. 7112–7127\.External Links:[Document](https://dx.doi.org/10.1109/TPAMI.2021.3095381)Cited by:[§1](https://arxiv.org/html/2606.31126#S1.p1.1),[§2](https://arxiv.org/html/2606.31126#S2.p2.1)\.
- P\. M\. Granitto, E\. Betta, I\. Khomenko, M\. Pedrotti, A\. Romano, F\. Biasioli,et al\.\(2026\)On the use of TabPFN on mass spectrometry analysis of volatile organic compounds\.Scientific Reports16,pp\. 164\.External Links:[Document](https://dx.doi.org/10.1038/s41598-025-29128-6)Cited by:[§1](https://arxiv.org/html/2606.31126#S1.p2.1)\.
- L\. Grinsztajn, K\. Flöge, O\. Key, F\. Birkel, P\. Jund, B\. Roof, M\. Manium, S\. B\. Hoo, M\. Bühler, A\. Garg, D\. Safaric, J\. Robertson, B\. Jäger, S\. Alessi, A\. Hayler, V\. Moroshan, L\. Purucker, P\. Singer, A\. Arazi, J\. Siems, J\. H\. Metzen, G\. Grab, N\. Erickson, S\. Guo, E\. Kalfon, S\. Bing, D\. Salinas, C\. Cornu, L\. C\. Wehrhahn, D\. Kriuchkova, K\. Kaya, L\. Sidhoum, M\. Salmon, J\. Chen, M\. Hulsebos, Y\. LeCun, S\. Müller, B\. Schölkopf, S\. Gambhir, N\. Hollmann, and F\. Hutter \(2026\)TabPFN\-3: technical report\.External Links:2605\.13986Cited by:[§2](https://arxiv.org/html/2606.31126#S2.p1.1)\.
- T\. Hayes, R\. Rao, H\. Akin, N\. J\. Sofroniew, D\. Oktay, Z\. Lin, R\. Verkuil, V\. Q\. Tran, J\. Deaton, M\. Wiggert,et al\.\(2025\)Simulating 500 million years of evolution with a language model\.Science387\(6736\),pp\. 850–858\.Cited by:[§1](https://arxiv.org/html/2606.31126#S1.p1.1),[§1](https://arxiv.org/html/2606.31126#S1.p4.1),[§2](https://arxiv.org/html/2606.31126#S2.p2.1),[§3](https://arxiv.org/html/2606.31126#S3.SS0.SSS0.Px1.p1.1)\.
- N\. Hollmann, S\. Müller, K\. Eggensperger, and F\. Hutter \(2023\)TabPFN: a transformer that solves small tabular classification problems in a second\.InInternational Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2606.31126#S1.p2.1),[§2](https://arxiv.org/html/2606.31126#S2.p1.1)\.
- N\. Hollmann, S\. Müller, L\. Purucker, A\. Krishnakumar, M\. Körfer, S\. B\. Hoo, R\. T\. Schirrmeister, and F\. Hutter \(2025\)Accurate predictions on small data with a tabular foundation model\.Nature637\(8045\),pp\. 319–326\.External Links:[Document](https://dx.doi.org/10.1038/s41586-024-08328-6)Cited by:[§1](https://arxiv.org/html/2606.31126#S1.p2.1),[§2](https://arxiv.org/html/2606.31126#S2.p1.1)\.
- K\. Huang, T\. Fu, W\. Gao, Y\. Zhao, Y\. Roohani, J\. Leskovec, C\. W\. Coley, C\. Xiao, J\. Sun, and M\. Zitnik \(2021\)Therapeutics data commons: machine learning datasets and tasks for drug discovery and development\.InProceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks,Vol\.1\.Cited by:[§1](https://arxiv.org/html/2606.31126#S1.p4.1),[§2](https://arxiv.org/html/2606.31126#S2.p3.1),[§5](https://arxiv.org/html/2606.31126#S5.SS0.SSS0.Px1.p1.1)\.
- Y\. Ji, L\. Zhang, J\. Wu, B\. Wu, L\. Li, L\. Huang, T\. Xu, Y\. Rong, J\. Ren, D\. Xue, H\. Lai, S\. Xu, J\. Feng, W\. Liu, P\. Luo, S\. Zhou, J\. Huang, P\. Zhao, and Y\. Bian \(2023\)DrugOOD: out\-of\-distribution \(ood\) dataset curator and benchmark for ai\-aided drug discovery\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.37,pp\. 8023–8031\.External Links:[Document](https://dx.doi.org/10.1609/aaai.v37i7.26103)Cited by:[§1](https://arxiv.org/html/2606.31126#S1.p4.1),[§2](https://arxiv.org/html/2606.31126#S2.p3.1),[§5](https://arxiv.org/html/2606.31126#S5.SS0.SSS0.Px1.p1.1)\.
- Q\. Li, R\. Dong, N\. Miklaucic, J\. Hu, S\. S\. Omee, L\. Wei, S\. Dey, M\. Hu, and J\. Hu \(2026\)In context learning foundation models for materials property prediction with small datasets\.npj Computational Materials\.External Links:[Document](https://dx.doi.org/10.1038/s41524-026-02089-8)Cited by:[§1](https://arxiv.org/html/2606.31126#S1.p2.1)\.
- Z\. Lin, H\. Akin, R\. Rao, B\. Hie, Z\. Zhu, W\. Lu, N\. Smetanin, R\. Verkuil, O\. Kabeli, Y\. Shmueli, A\. dos Santos Costa, M\. Fazel\-Zarandi, T\. Sercu, S\. Candido, and A\. Rives \(2023\)Evolutionary\-scale prediction of atomic\-level protein structure with a language model\.Science379\(6637\),pp\. 1123–1130\.External Links:[Document](https://dx.doi.org/10.1126/science.ade2574)Cited by:[§1](https://arxiv.org/html/2606.31126#S1.p1.1),[§2](https://arxiv.org/html/2606.31126#S2.p2.1),[§3](https://arxiv.org/html/2606.31126#S3.SS0.SSS0.Px1.p1.1)\.
- S\. Müller, N\. Hollmann, S\. P\. Arango, J\. Grabocka, and F\. Hutter \(2022\)Transformers can do bayesian inference\.InInternational Conference on Learning Representations,Cited by:[§2](https://arxiv.org/html/2606.31126#S2.p1.1)\.
- P\. Notin, A\. Kollasch, D\. Ritter, L\. Van Niekerk, S\. Paul, H\. Spinner, N\. Rollins, A\. Shaw, R\. Orenbuch, R\. Weitzman,et al\.\(2023\)Proteingym: large\-scale benchmarks for protein fitness prediction and design\.Advances in Neural Information Processing Systems36,pp\. 64331–64379\.Cited by:[§1](https://arxiv.org/html/2606.31126#S1.p1.1),[§2](https://arxiv.org/html/2606.31126#S2.p2.1),[§4](https://arxiv.org/html/2606.31126#S4.SS0.SSS0.Px1.p1.1)\.
- F\. Pedregosa, G\. Varoquaux, A\. Gramfort, V\. Michel, B\. Thirion, O\. Grisel, M\. Blondel, P\. Prettenhofer, R\. Weiss, V\. Dubourg, J\. Vanderplas, A\. Passos, D\. Cournapeau, M\. Brucher, M\. Perrot, and E\. Duchesnay \(2011\)Scikit\-learn: machine learning in Python\.Journal of Machine Learning Research12,pp\. 2825–2830\.Cited by:[§4](https://arxiv.org/html/2606.31126#S4.SS0.SSS0.Px2.p1.1)\.
- J\. Qu, D\. Holzmüller, G\. Varoquaux, and M\. Le Morvan \(2025\)TabICL: a tabular foundation model for in\-context learning on large data\.InProceedings of the 42nd International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.267,pp\. 50817–50847\.Cited by:[§1](https://arxiv.org/html/2606.31126#S1.p2.1),[§2](https://arxiv.org/html/2606.31126#S2.p1.1)\.
- J\. Qu, D\. Holzmüller, G\. Varoquaux, and M\. Le Morvan \(2026\)TabICLv2: a better, faster, scalable, and open tabular foundation model\.Note:Preprint, arXiv:2602\.11139Cited by:[§1](https://arxiv.org/html/2606.31126#S1.p2.1),[§2](https://arxiv.org/html/2606.31126#S2.p1.1)\.
- D\. Rogers and M\. Hahn \(2010\)Extended\-connectivity fingerprints\.Journal of Chemical Information and Modeling50\(5\),pp\. 742–754\.External Links:[Document](https://dx.doi.org/10.1021/ci100050t)Cited by:[§2](https://arxiv.org/html/2606.31126#S2.p3.1),[§3](https://arxiv.org/html/2606.31126#S3.SS0.SSS0.Px2.p1.1)\.
- M\. Stanley, J\. F\. Bronskill, K\. Maziarz, H\. Misztela, J\. Lanini, M\. Segler, N\. Schneider, and M\. Brockschmidt \(2021\)FS\-Mol: a few\-shot learning dataset of molecules\.InProceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks,Vol\.1\.Cited by:[§1](https://arxiv.org/html/2606.31126#S1.p4.1),[§2](https://arxiv.org/html/2606.31126#S2.p3.1),[§5](https://arxiv.org/html/2606.31126#S5.SS0.SSS0.Px1.p1.1)\.
- D\. Wu, J\. Jen, E\. Fajiculay, M\. Hsu, M\. Chang, J\. Yeh, K\. Sargsyan, J\. Kupcinskas, J\. Skieceviciene, R\. Steponaitiene, E\. Morkunas, G\. Gedgaudiene, C\. Hsu, Y\. Chang, C\. Hu,et al\.\(2026\)PanMETAI \- a high performance tabular foundation model for accurate pancreatic cancer diagnosis via NMR metabolomics\.Nature Communications17,pp\. 1595\.External Links:[Document](https://dx.doi.org/10.1038/s41467-026-69426-9)Cited by:[§1](https://arxiv.org/html/2606.31126#S1.p2.1)\.
- Z\. Wu, B\. Ramsundar, E\. N\. Feinberg, J\. Gomes, C\. Geniesse, A\. S\. Pappu, K\. Leswing, and V\. Pande \(2018\)MoleculeNet: a benchmark for molecular machine learning\.Chemical Science9\(2\),pp\. 513–530\.External Links:[Document](https://dx.doi.org/10.1039/C7SC02664A)Cited by:[§1](https://arxiv.org/html/2606.31126#S1.p4.1),[§2](https://arxiv.org/html/2606.31126#S2.p3.1),[§5](https://arxiv.org/html/2606.31126#S5.SS0.SSS0.Px1.p1.1)\.
- K\. Yang, K\. Swanson, W\. Jin, C\. Coley, P\. Eiden, H\. Gao, A\. Guzman\-Perez, T\. Hopper, B\. Kelley, M\. Mathea, A\. Palmer, V\. Settels, T\. Jaakkola, K\. Jensen, and R\. Barzilay \(2019\)Analyzing learned molecular representations for property prediction\.Journal of Chemical Information and Modeling59\(8\),pp\. 3370–3388\.External Links:[Document](https://dx.doi.org/10.1021/acs.jcim.9b00237)Cited by:[§2](https://arxiv.org/html/2606.31126#S2.p3.1),[§3](https://arxiv.org/html/2606.31126#S3.SS0.SSS0.Px2.p1.1)\.

\\closingpage

Similar Articles

Probing Memorization of Tabular In-Context Learning

arXiv cs.LG

This paper investigates parametric memorization in tabular foundation models that use in-context learning, introducing a probing framework (IclMem) to separate context-based predictions from memorization. It finds moderate memorization signals under specific conditions but notes they largely vanish under realistic training scenarios.

The Good, the Bad, and the Ugly of Markov Boundary for Tabular Prediction

Hugging Face Daily Papers

This paper evaluates the practical effectiveness of Markov boundaries for tabular prediction, finding that while theoretically optimal, current causal discovery methods fail to consistently improve predictive performance due to computational limitations and mismatched optimization goals.