Position: Genomic Model Research Must Move Beyond Anecdotal Evaluation of Interpretability Methods
Summary
This position paper argues that genomic model interpretability research must move beyond anecdotal evaluation, proposing a tiered framework for rigorous assessment of consistency, faithfulness, and biological validity, demonstrated through a benchmarking study on transcription factor binding.
View Cached Full Text
Cached at: 06/09/26, 08:50 AM
# Genomic Model Research Must Move Beyond Anecdotal Evaluation of Interpretability Methods
Source: [https://arxiv.org/html/2606.07607](https://arxiv.org/html/2606.07607)
###### Abstract
Advances in machine learning and computational power have unlocked the predictive potential of the human genome, yet biologists now demand that these models also elucidate the underlying biological mechanisms\. While interpretable machine learning \(IML\) techniques have been increasingly applied to bridge this gap, there has been a pervasive reliance on anecdotal validation: the vast majority of research relies on a single IML method and reports only isolated successful instances\. Through a benchmarking study on transcription factor binding, we demonstrate the risks of current practices\. We show that different IML methods can often \(1\) yield contradictory explanations for the same predictions, \(2\) fail to localize known regulatory motifs, and \(3\) fail to faithfully reflect the model’s internal decision process\. In light of this, we argue for a validation framework analogous to clinical trials: just as trials require rigorous design and adverse\-event reporting, genomic interpretability must move beyond cherry\-picked plausibility toward systematic assessment of consistency, faithfulness, and biological validity\. To facilitate this, we propose a tiered framework to guide rigorous evaluation and reporting of genomic IML methods\.
Machine Learning, ICML
## 1Introduction
Understanding the language of the genome, i\.e\., how sequence variations shape molecular phenotypes, is a long\-standing goal in functional genomics, underpinning applications from prioritizing non\-coding variants in human genetics\(Avsecet al\.,[2021a](https://arxiv.org/html/2606.07607#bib.bib44); Zhouet al\.,[2018](https://arxiv.org/html/2606.07607#bib.bib48)\), to rational design in synthetic biology\(Vaishnavet al\.,[2022](https://arxiv.org/html/2606.07607#bib.bib72)\), and uncovering disease mechanisms for therapeutic target discovery\(Gaoet al\.,[2023](https://arxiv.org/html/2606.07607#bib.bib82)\)\. To this end, numerous deep learning approaches have proliferated\(Zouet al\.,[2019](https://arxiv.org/html/2606.07607#bib.bib39); Eraslanet al\.,[2019](https://arxiv.org/html/2606.07607#bib.bib40); Barbadilla\-Martínezet al\.,[2025](https://arxiv.org/html/2606.07607#bib.bib41); Consenset al\.,[2025](https://arxiv.org/html/2606.07607#bib.bib42); Yanget al\.,[2025](https://arxiv.org/html/2606.07607#bib.bib110); Huet al\.,[2026](https://arxiv.org/html/2606.07607#bib.bib111)\)\. With the integration of high\-throughput functional assays\(Melnikovet al\.,[2012](https://arxiv.org/html/2606.07607#bib.bib37); Fowleret al\.,[2010](https://arxiv.org/html/2606.07607#bib.bib38)\)and burgeoning computational power, these deep neural networks \(DNNs\) have enabled potent predictive power across various genomic task, such as gene expression\(Avsecet al\.,[2021a](https://arxiv.org/html/2606.07607#bib.bib44); Zhouet al\.,[2018](https://arxiv.org/html/2606.07607#bib.bib48)\), transcription factor \(TF\) binding\(Avsecet al\.,[2021b](https://arxiv.org/html/2606.07607#bib.bib43); Alipanahiet al\.,[2015](https://arxiv.org/html/2606.07607#bib.bib45); Zhou and Troyanskaya,[2015](https://arxiv.org/html/2606.07607#bib.bib47)\), enhancer\-promoter activity\(de Almeidaet al\.,[2022](https://arxiv.org/html/2606.07607#bib.bib46)\)among many others\.
However, the predictive success of these models has not been matched by a commensurate understanding of their inner workings\. Because of their intrinsic complexity \(e\.g\., by featuring millions to billions of parameters\), such DNNs are often perceived as a “black box”, with no explanation about how a given prediction was made or what biological mechanisms were learned\. To address this challenge, a variety ofpost hocinterpretability methods have been employed for elucidating the predictions of genomic DNNs\(Novakovskyet al\.,[2023](https://arxiv.org/html/2606.07607#bib.bib87); Chenet al\.,[2024](https://arxiv.org/html/2606.07607#bib.bib98)\)\.
The most prevalent are feature attribution methods such as Saliency Maps\(Simonyanet al\.,[2014](https://arxiv.org/html/2606.07607#bib.bib50)\), DeepLIFT\(Shrikumaret al\.,[2017](https://arxiv.org/html/2606.07607#bib.bib51)\), ISM\(Zhou and Troyanskaya,[2015](https://arxiv.org/html/2606.07607#bib.bib47)\), and Integrated Gradients\(Sundararajanet al\.,[2017](https://arxiv.org/html/2606.07607#bib.bib53); Lundberg and Lee,[2017](https://arxiv.org/html/2606.07607#bib.bib54)\), which quantify the position\-specific effect of sequence variations on model predictions\. Beyond such main effects, interaction\-based methods like integrated Hessians\(Janizeket al\.,[2021](https://arxiv.org/html/2606.07607#bib.bib55)\), SQUID\(Seitzet al\.,[2024](https://arxiv.org/html/2606.07607#bib.bib57)\), and DFIM\(Greensideet al\.,[2018](https://arxiv.org/html/2606.07607#bib.bib56)\)probe the epistatic interactions between mutations, while attention\-based interpretation\(Viget al\.,[2021](https://arxiv.org/html/2606.07607#bib.bib11)\)and sparse autoencoders \(SAEs\)\(Brickenet al\.,[2023](https://arxiv.org/html/2606.07607#bib.bib105); Cunninghamet al\.,[2024](https://arxiv.org/html/2606.07607#bib.bib106); Brixiet al\.,[2026](https://arxiv.org/html/2606.07607#bib.bib107)\)instead examine the internal computations of transformer\-based genomic models\. Collectively, these techniques have enabled applications ranging from motif discovery and mechanistic understanding to de novo design and variant prioritization\(Avsecet al\.,[2021b](https://arxiv.org/html/2606.07607#bib.bib43); Alipanahiet al\.,[2015](https://arxiv.org/html/2606.07607#bib.bib45); de Almeidaet al\.,[2022](https://arxiv.org/html/2606.07607#bib.bib46); Avsecet al\.,[2021a](https://arxiv.org/html/2606.07607#bib.bib44)\)\.
Despite this success, the lack of consensus on how to rigorously evaluate and report interpretability results in genomics is a significant concern\. To quantify this gap, we conducted a large\-scale mapping study of3,5753,575papers employing Interpretable Machine Learning \(IML\) in genomics between 2010 and 2025\. We found that the vast majority of studies: \(1\) relied on a single IML method without justification; \(2\) validated interpretations through heterogeneous and often subjective means; and \(3\) reported only “successful” cases while omitting failures \(Section[2](https://arxiv.org/html/2606.07607#S2)\)\.
Furthermore, through a systematic benchmarking study on transcription factor binding prediction, we demonstrate that these practices have tangible consequences\. We show that: \(a\) different IML methods produce drastically inconsistent explanations; \(b\) these explanations often fail to reflect the model’s actual decision\-making process; and \(c\) results frequently misalign with known biological mechanisms—yielding near\-zero recovery of true binding sites on challenging tasks, even when cherry\-picking “successful” examples remains easy \(Section[3](https://arxiv.org/html/2606.07607#S3)\)\.
Because rigorous evaluation is the cornerstone of scientific progress, we argue that genomics researchers must move beyond anecdotal evidence and systematically evaluate IML methods before relying on their outputs\.
Position:The credibility of genomic deep learning is undermined by a reliance on anecdotal evaluation for interpretability\. We demonstrate that current practices—characterized by single\-method usage and cherry\-picked examples—fail to expose that explanations are frequently inconsistent, unfaithful to the model, and biologically misleading\. We argue that the field must move beyond subjective “plausibility” toward systematic probing\. We propose a tiered framework to rigorously evaluate theconsistency,faithfulness, andbiological validityof genomic interpretations\.
## 2Mapping of Existing Practices
Figure 1:Systematic mapping reveals a reliance on anecdotal evaluation practices\.\(a\)Distribution of the number of IML methods employed per study \(n=3,575n=3,575, the same for all panels\)\.\(b\)Breakdown of validation strategies\.\(c\)Frequency of validated interpretation instances per paper\.\(d\)Reporting of failure modes\.To establish a broad understanding of the existing practices in the evaluation and reporting of IML in genomics, we conducted a comprehensive, data\-driven mapping study\.
### 2\.1Methods
We first searched on the Web of Science \(WoS\) database111We use the WoS database because it offers the most comprehensive coverage of literature with highly curated metadata\.for papers that explicitly mentioned the use of IML methods in genomics by applying the following search query\. We applied this query to paper abstracts to optimize the trade\-off between coverage and relevance\.
\(
\("interpretability"OR"explainability"
OR"explainableAI"OR"XAI"
OR"featureimportance"OR"LIME"
OR"featureattribution\*"OR"SHAP"
OR"saliency"OR"attentionmechanism"OR"modelexplanation\*"\)
AND
\("genomics"OR"genome"OR"DNA"OR"RNA"
OR"protein"OR"geneexpression"
OR"sequencing"OR"transcriptomics"
OR"proteomics"\)
AND
\("machinelearning"OR"deeplearning"
OR"neuralnetwork\*"OR"transformer"
OR"foundationmodel\*"
OR"languagemodel\*"OR"predict\*"\)
\)
This search led to3,5753,575papers published between 2010 and 2025\. Full\-text articles were accessed through a combination of institutional subscriptions, open\-access repositories, and preprint servers where available\. As this amount of papers is too large for manual review, inspired byHuanget al\.\([2025a](https://arxiv.org/html/2606.07607#bib.bib80)\), we prompted agemini\-3\-flash\-previewmodel \(to balance costs and accuracy\) to extract structured information from each full\-text article\. Specifically, for each paper, we extracted:
- •Types and number of IML methods employed\. - –Any justifications associated with this choice?
- •Types of validation for IML performed \(e\.g\., visual inspection, querying domain\-specific databases, or via wet\-lab experiments?\)
- •Number of “*successful*” instances, defined as cases the authors*claim*an IML output aligns with established domain knowledge \(e\.g\., an attribution highlighting a known TF binding motif, or identified marker genes matching known cell types\)\. - –Any report on failed instances and their analysis?
To maximize extraction quality, we employed few\-shot in\-context learning \(ICL\) anchored by1010manually curated demonstrations with self\-reflection prompts to avoid hallucinations \(see Appendix[A](https://arxiv.org/html/2606.07607#A1)for details\)\. We also randomly sampled100100papers, and one of the authors manually verified the model outputs, which yielded an overall agreement rate of98%98\\%\(see Table[2](https://arxiv.org/html/2606.07607#A2.T2)in Appendix[B](https://arxiv.org/html/2606.07607#A2)for the per\-category breakdown\)\.
Our analysis reveals several concerning patterns in how IML methods are evaluated and reported in genomics research\. First, the vast majority of studies \(86%86\\%\) employed only a single IML method \(Fig\.[1](https://arxiv.org/html/2606.07607#S2.F1)a\), with no comparison against alternative interpretation techniques, nor justifications on why the chosen approach is more appropriate in this case\. This lack of cross\-method validation makes it difficult to assess the robustness or reliability of the reported interpretations\.
Second, and perhaps most concerning, we observed substantial heterogeneity in validation practices \(Fig\.[1](https://arxiv.org/html/2606.07607#S2.F1)b\)\. Strikingly, nearly half of the studies relied on subjective visual inspection \(e\.g\., whether highlighted regions correspond to known structural contacts\)\. The remaining studies largely relied on indirect proxies like ChIP\-seq peaks or simplified synthetic datasets, while only2\.81%2\.81\\%conducted wet\-lab experiments to validate insights\.
Third, the majority of papers reported only11–33“*successful*” instances where IML outputs aligned with known biology \(Fig\.[1](https://arxiv.org/html/2606.07607#S2.F1)c\), which raises concerns about cherry\-picking\. Furthermore, these papers typically did not disclose whether additional instances were examined, nor whether any interpretations failed to align with expectations \(Fig\.[1](https://arxiv.org/html/2606.07607#S2.F1)d\)\.
Such anecdotal evidence presents a problematic foundation for assessing IML capabilities\. The selective reporting of favorable cases risks overstating the reliability of these methods, potentially fostering unwarranted confidence among practitioners\. In high\-stakes genomics applications—where experimental validation can be prohibitively expensive—over\-reliance on inadequately evaluated IML methods could lead to substantial misallocation of resources and, more critically, erroneous biological conclusions\.
## 3Consequences of Current Practices: An Empirical Audit
The mapping study above reveals that evaluation in genomic IML relies heavily on anecdotal evidence\. To move beyond abstract critique and concretely demonstrate the risks of these practices, we conduct a targeted empirical audit\. We design a controlled case study to illustrate three specific failure modes that arise when rigorous validation is neglected: \(1\)Inconsistency, where methods yield contradictory explanations; \(2\)Unfaithfulness, where explanations decouple from model logic; and \(3\)Biological Misalignment, where “plausible” explanations fail to match causal ground truth\.
Experimental Design\.We anchor our audit on Transcription Factor \(TF\) binding prediction\(Fenget al\.,[2025](https://arxiv.org/html/2606.07607#bib.bib97); Avsecet al\.,[2021b](https://arxiv.org/html/2606.07607#bib.bib43); Vorontsovet al\.,[2025](https://arxiv.org/html/2606.07607#bib.bib90)\), a canonical task in regulatory genomics\. Rather than aggregating broad metrics, we specifically curated five TFs from ENCODE222https://www\.encodeproject\.org/database to serve as distinct stress tests that probe the boundaries of IML validity:
- •Positive Controls \(CTCF, MAX\):We include these factors because they bind long, high\-signal motifs\. They represent the “easy” cases where any valid explanation method*must*succeed\.
- •Compositional Bias Tests \(SP1, TBP\):To test whether methods detect specific regulatory grammar or merely “shortcut learn” background frequencies, we selected SP1 \(which targets GC\-rich islands\) and TBP \(which targets AT\-rich promoters\)\.
- •Resolution Tests \(GATA1\):We selected GATA1 to test the spatial precision of token\-based architectures, as it binds a short \(∼\\sim6 bp\) and often degenerate motif\.
For each transcription factor, we collect high\-confidence narrowPeak files from ENCODE\. Each peak is represented as a fixed\-length DNA sequence centered on the peak summit\. Positive samples correspond to experimentally validated binding regions\. Negative ones are generated from matched genomic regions that do not overlap called peaks, controlling for chromosome distribution and sequence length\.
Table 1:Dataset statistics for each TF\.Ground Truthdenotes the total number of sequences \(across splits\) with UniBind motif annotations used for interpretability evaluation\.To ensure these stress tests are statistically robust, we constructed a large\-scale benchmark comprising approximately289,000289,000training and11,50011,500test sequences \(summarized in Table[1](https://arxiv.org/html/2606.07607#S3.T1)\)\. We strictly split the data by chromosome, holding out Chromosome 9 entirely for testing to prevent information leakage\. Crucially, this scale allows us to move beyond the qualitative inspection of cherry\-picked examples; we utilize over88,00088,000sequences with annotated motif positions to systematically quantify interpretability performance\.
We replicate this audit across three divergent genomic foundation models—DNABERT\-2\(Zhouet al\.,[2022](https://arxiv.org/html/2606.07607#bib.bib99)\)\(Transformer\),HyenaDNA\(Nguyenet al\.,[2023](https://arxiv.org/html/2606.07607#bib.bib101)\)\(convolution/SSM\), andNucleotide Transformer v3\(NTv3\)\(Bosharet al\.,[2025](https://arxiv.org/html/2606.07607#bib.bib100)\)\(hybrid\)—and evaluate five common interpretability methods: DeepLIFT\(Shrikumaret al\.,[2017](https://arxiv.org/html/2606.07607#bib.bib51)\), IG\(Sundararajanet al\.,[2017](https://arxiv.org/html/2606.07607#bib.bib53)\), ISM\(Zhou and Troyanskaya,[2015](https://arxiv.org/html/2606.07607#bib.bib47)\), Local Interpretable Model\-agnostic Explanations \(LIME\)\(Ribeiroet al\.,[2016](https://arxiv.org/html/2606.07607#bib.bib102)\), and the Categorical Jacobian \(CJ\)\(Zhanget al\.,[2024](https://arxiv.org/html/2606.07607#bib.bib108)\)\. The CJ quantifies pairwise dependencies by measuring how the model’s predicted distribution at one position changes under perturbations at every other position; to make it comparable to the four per\-position attribution methods, we reduce its position\-by\-position coupling matrix to a per\-position importance score by averaging absolute couplings over partner positions\. All three models attain non\-trivial predictive performance across the five TFs \(see Table[3](https://arxiv.org/html/2606.07607#A3.T3)in Appendix[C](https://arxiv.org/html/2606.07607#A3)for accuracy and F1 scores\), ensuring that the explanations we audit are derived from models that have learned meaningful patterns from the data\. To establish an objective ground truth, we rely on UniBind333https://unibind\.uio\.no/database\. By retrieving experimentally validated motif coordinates for our held\-out test set, we can quantitatively measure whether explanations align with known causal mechanisms, independent of the model’s labels\.
### 3\.1IML Methods Disagree with Each Other
Figure 2:Different IML methods produce inconsistent explanations\.\(a\) Attribution maps from five IML methods on the same CTCF sequence \(NTv3 model\), showing strikingly different patterns\. \(b\) Mean Spearman rank correlation coefficients between method pairs across all models and datasets\. \(c\) Mean Jaccard similarity of the top\-20 attributed positions between method pairs\.If IML methods reliably recovered the true explanation for a model’s prediction, we would expect different methods to produce consistent results\. We test this assumption by computing explanations from all IML methods on the same trained model and comparing their outputs\.
Methodology\.For each test sequence, we obtain per\-nucleotide importance scores \(absolute values\) from each IML method\. We quantify agreement between method pairs using: \(1\) Spearman rank correlation to measure global consistency across all positions, and \(2\) Jaccard similarity of the top\-2020most important positions to assess agreement on critical features—a particularly relevant metric since regulatory motifs are typically shorter than2020bp\.
Results\.Fig\.[2](https://arxiv.org/html/2606.07607#S3.F2)reveals disagreement among IML methods\. Fig\.[2](https://arxiv.org/html/2606.07607#S3.F2)a visualizes attribution maps from all five IML methods on the same CTCF sequence on NTv3, showing qualitatively different patterns: DeepLIFT and CJ show a dominant peak with broader background activity across the sequence, IG concentrates almost entirely on a single sharp peak, ISM generates a few sparse and isolated peaks, and LIME produces a diffuse, noisy signal throughout\.
Quantitatively, Fig\.[2](https://arxiv.org/html/2606.07607#S3.F2)b shows that the average Spearman rank correlation between method pairs is consistently below0\.40\.4across all models and datasets\. The highest correlation \(ISM–LIME:ρ=0\.373\\rho=0\.373\) still indicates weak agreement, while several pairs show near\-zero or even negative correlations \(IG–CJ:ρ=−0\.021\\rho=\-0\.021\)\. Fig\.[2](https://arxiv.org/html/2606.07607#S3.F2)c demonstrates that even when focusing on the most critical positions, the maximum Jaccard similarity remains below0\.50\.5\(DeepLIFT–IG:0\.4940\.494\), meaning methods agree on fewer than half of the positions deemed most important\. Per\-model and per\-TF breakdowns of these correlation and Jaccard analyses are provided in Fig\.[5](https://arxiv.org/html/2606.07607#A4.F5)and Fig\.[6](https://arxiv.org/html/2606.07607#A4.F6)in Appendix[D](https://arxiv.org/html/2606.07607#A4)\.
These results raise a fundamental concern regarding reliability: since different IML methods produce inconsistent explanations, it is unclear which, if any, accurately captures the model’s reasoning\. If practitioners follow the common practice of using only one IML method, they risk relying on an arbitrary or biased viewpoint that may not reflect the model’s true decision process\.
Figure 3:Faithfulness evaluation via perturbation analysis\.\(a\) Sequential deletion \(MoRF\): prediction probability as top\-ranked positions are progressively masked\. \(b\) Sequential insertion: probability recovery as positions are restored to a neutral baseline\. Curves show means with shaded standard errors on CTCF \(NTv3\)\. \(c, d\) Distribution of mean AUC scores across all sequences for deletion \(c\) and insertion \(d\) experiments, stratified by TF on NTv3\.
### 3\.2IML Explanations Are Not Faithful to Model Decisions
Beyond inter\-method consistency, we ask whether explanations faithfully reflect what the*model*considers important—regardless of biological validity\. We test this through systematic perturbation experiments\(Sameket al\.,[2016](https://arxiv.org/html/2606.07607#bib.bib103)\)\.
Methodology\.If an IML method correctly identifies positions most important to a model’s prediction, then perturbing those positions should maximally affect model output\. We implement two complementary tests:
1. 1\.Sequential deletion \(MoRF\)\(Sameket al\.,[2016](https://arxiv.org/html/2606.07607#bib.bib103)\): Starting from original sequence, we progressively mask positions in descending order of attributed importance by replacing them with neutral tokens\(‘N’\)\. A faithful explanation should produce rapid probability decay\.
2. 2\.Sequential insertion\(Petsiuket al\.,[2018](https://arxiv.org/html/2606.07607#bib.bib104)\): Starting from a fully masked baseline, we progressively restore positions in descending order of importance\. A faithful explanation should produce rapid probability recovery\.
We quantify faithfulness via the area under the probability curve \(AUC\): lower deletion\-AUC and higher insertion\-AUC indicate better faithfulness\.
Results\.Fig\.[3](https://arxiv.org/html/2606.07607#S3.F3)a and b shows representative curves for CTCF on the NTv3 model\. In the deletion experiment, ISM causes the steepest probability drop, followed by LIME and IG, while DeepLIFT and CJ produce the slowest decay\. The insets reveal that meaningful separation between methods emerges within the first 20–50 positions—precisely where motif\-level features reside\. The insertion curves show the inverse pattern: ISM and LIME rapidly recover prediction confidence, while other methods require substantially more positions to achieve comparable recovery\.
Fig\.[3](https://arxiv.org/html/2606.07607#S3.F3)c and d present the distribution of AUC scores across all sequences and TFs on NTv3 \(DNABERT\-2 and HyenaDNA is shown in Appendix[E](https://arxiv.org/html/2606.07607#A5)\)\. For deletion \(Fig\.[3](https://arxiv.org/html/2606.07607#S3.F3)c\), lower AUC indicates better faithfulness; IG consistently achieves the lowest scores across most TFs, though with substantial variance\. Notably, for SP1 and TBP—the compositional bias stress tests—all methods show high AUC values clustered near 1\.0, suggesting that none reliably identifies the features driving model predictions\. The insertion results \(Fig\.[3](https://arxiv.org/html/2606.07607#S3.F3)d\) mirror this pattern: IG generally achieves the highest AUC \(indicating faster recovery\), but performance degrades markedly on SP1 compared to other TFs\.
These results demonstrate that faithfulness varies considerably across both methods and tasks\. While IG and ISM show relatively better faithfulness, no method consistently excels, and performance deteriorates on tasks where compositional biases may confound the explanations\.
### 3\.3IML Explanations Do Not Align with Biological Ground Truth
Figure 4:Alignment between IML explanations and biological ground truth\.\(a\) Distribution of motif overlap scores \(Perception\) across TFs and IML methods on NTv3\. Higher values indicate better alignment with UniBind\-annotated binding sites\. \(b, c\) Representative examples showing attribution profiles \(colored curves\) relative to ground\-truth motif regions \(white background\): \(b\) successful alignment where attribution peaks coincide with the annotated motif; \(c\) failure case where high attributions occur outside the true binding site\.Ultimately, the value of IML in genomics lies in recovering biologically meaningful signals\. We evaluate whether explanations align with experimentally validated TF motifs\.
Methodology\.For each positive test sequence with an annotated UniBind motif, we extract the contiguous region of lengthLL\(matching the motif length\) with the highest summed attribution\. We then compute the overlap ratio between this predicted region and the ground\-truth motif annotation, yielding a Perception score between 0 \(no overlap\) and 1 \(perfect alignment\)\.
Results\.Fig\.[4](https://arxiv.org/html/2606.07607#S3.F4)a reveals stark task\-dependent variation in biological alignment on NTv3\. For CTCF, which is our high\-SNR positive control, all methods achieve reasonable alignment, with median Perception scores exceeding 0\.5 and many sequences approaching perfect overlap\. This confirms that when the biological signal is strong and unambiguous, IML methods can successfully recover it\.
However, performance deteriorates dramatically on other TFs\. For GATA1 \(resolution challenge\) and MAX, median scores drop to 0\.2–0\.4, with substantial probability mass near zero; on GATA1, CJ in particular collapses to a near\-zero median, the lowest of all five methods\. Most strikingly, for the compositional bias stress tests \(SP1 and TBP\), all methods show near\-complete failure: the vast majority of sequences yield Perception scores close to zero, which indicates that IML\-identified important regions rarely coincide with true binding sites\. This suggests that explanations are confounded by background nucleotide composition rather than capturing specific regulatory motifs\. Per\-model Perception distributions are reported in Fig\.[10](https://arxiv.org/html/2606.07607#A6.F10)in Appendix[F](https://arxiv.org/html/2606.07607#A6)\.
Fig\.[4](https://arxiv.org/html/2606.07607#S3.F4)b and c illustrate why anecdotal evidence can be misleading\. Fig\.[4](https://arxiv.org/html/2606.07607#S3.F4)b shows a good example where LIME attributions \(blue\) align precisely with the annotated motif region \(white background\), which is the type of cherry\-picked case often showcased in papers\. Fig\.[4](https://arxiv.org/html/2606.07607#S3.F4)c shows a bad example where high attributions appear entirely outside the true motif location, yet such failures are rarely reported\. Our systematic evaluation reveals that failure cases like Fig\.[4](https://arxiv.org/html/2606.07607#S3.F4)c are far more common than successes like Fig\.[4](https://arxiv.org/html/2606.07607#S3.F4)b for most TF tasks\.
## 4Practical Guidelines for IML Evaluation
With the identified critical limitations in current IML evaluation practices, we now propose a tiered framework of practical guidelines\. In particular, considering that practitioners often operate under varying resource constraints, we organize our recommendations by the level of effort required to accommodate different practical scenarios\.
### 4\.1Low\-Effort Guidelines: Minimum Standards
The following practices require minimal additional effort and should be considered*mandatory*for any study reporting IML results in computational genomics\.
G1: Use Multiple IML Methods\.As demonstrated in the previous section, different IML methods can produce drastically different explanations for the same prediction\. Relying on a single method provides no indication of explanation robustness\. We recommend:
- •Apply at least three IML methods spanning different categories \(e\.g\., gradient\-based, perturbation\-based, and attention\-based\) for generating interpretations\.
- •Report the agreement between methods using quantitative metrics such as rank correlation or top\-kkoverlap, and treat high disagreement as a warning signal that explanations may be unreliable\.
G2: Include Negative Controls\.Explanations should be compared against appropriate baselines to assess whether they provide information beyond chance:
- •Compare IML importance scores against random baselines \(uniform random importance assignment\)\.
- •For sequence data, consider shuffled\-sequence controls where the same IML method is applied to sequences with shuffled nucleotide order\.
G3: Report Quantitative Metrics Over Visual Inspection\.Qualitative assessments \(“the highlighted region appears to overlap with the known binding site”\) are subjective and irreproducible\. Instead, we recommend:
- •Define precise quantitative criteria \(e\.g\., precision, recall, correlation, overlap, etc\., when ground\-truth annotations are available\)*before*examining explanations\.
G4: Report Holistic Statistics Instead of Cherry\-Picked Examples\.The common practice of showing 1 to 5 “successful” examples provides no information about typical explanation quality\. Instead, we recommend:
- •Compute quantitative evaluation metrics across all test samples\.
- •Report distributions \(mean, standard deviation, percentiles\) rather than single exemplars\.
- •If illustrative examples are shown, explicitly state how they were selected and how representative they are of the overall distribution\.
### 4\.2Moderate\-Effort Guidelines: Computational Validation
The following practices require additional computational experiments but no wet\-lab resources\. They provide stronger evidence of explanation quality\.
G5: Perform Faithfulness Tests\.As shown in Section[3](https://arxiv.org/html/2606.07607#S3), explanations often fail to reflect the model’s actual decision process\. Faithfulness can be assessed computationally:
- •Perturbation\-based evaluation: Mask or perturb positions identified as important and measure the prediction drop\. Compare against random masking\.
- •Sufficiency test: Retain*only*the top\-kkimportant positions \(masking everything else\) and verify that the prediction is preserved\.
- •Comprehensiveness test: Mask the top\-kkimportant positions and verify that the prediction degrades substantially\.
- •Report faithfulness metrics relative to a random baseline\. Instead of reporting raw prediction drops, quantify the margin between IML\-guided perturbation and random perturbation \(e\.g\., AUC scores\)\.
G6: Test Across Diverse Conditions\.Explanation quality may vary across different data characteristics\. We recommend systematically evaluating across:
- •Differentprediction confidence levels\(e\.g\., do explanations degrade for uncertain predictions?\)\.
- •Differentsequence contexts\(e\.g\., GC\-rich vs\. AT\-rich regions, as in our experiments\)\.
- •Differentfunctional categories\(e\.g\., promoters vs\. enhancers, coding vs\. non\-coding regions\)\.
### 4\.3High\-Effort Guidelines: Experimental Validation
The following practices require wet\-lab experiments or substantial additional resources\. They provide the strongest evidence but should be guided by computational pre\-screening\.
G7: Design Experiments with Proper Controls\.When wet\-lab validation is pursued, ensure rigorous experimental design:
- •Select validation targets through stratified random sampling across the whole distribution of IML scores, not by cherry\-picking high\-confidence cases\.
- •Include bothnegativecontrols \(e\.g\., instances where the IML method predicts low importance for known functional elements\) andpositivecontrols \(instances where both the model and prior knowledge agree on important regions\)\.
- •Pre\-register the validation targetsbeforeconducting experiments to prevent post\-hoc selection bias\.
- •Report all results, including failures and inconclusive cases, not just successful validations\.
### 4\.4A Recommended Workflow
We synthesize the above guidelines into a practical workflow for IML evaluation:
1. 1\.Baseline Assessment \(Low Effort\): Apply multiple IML methods, compute agreement metrics, and establish random baselines\. If methods strongly disagree or perform near random,*stop*—explanations are likely unreliable\.
2. 2\.Computational Validation \(Moderate Effort\): Conduct faithfulness tests and sanity checks\. Evaluate against available biological databases\. Identify conditions where explanations succeed or fail\.
3. 3\.Targeted Experimental Validation \(High Effort\): Based on computational screening, design rigorous wet\-lab experiments with proper controls\. Validate a representative sample and report all results\.
4. 4\.Iterative Refinement: Use validation results to refine model architecture, training procedure, or IML method selection\. Re\-evaluate after modifications\.
By adopting these tiered guidelines, practitioners can move beyond anecdotal evidence toward rigorous, reproducible evaluation of IML methods\. Even implementing only the low\-effort guidelines would represent a substantial improvement over current practices and help prevent the resource misallocation documented in Section[2](https://arxiv.org/html/2606.07607#S2)\.
## 5Alternative Views
The concerns we raise about IML evaluation practices may seem overly stringent to some practitioners\. Indeed, several alternative viewpoints exist that could justify current practices\. We examine each of these perspectives and explain why they are insufficient to ensure reliable interpretability in genomics applications\.
Theoretical Guarantees\.Several popular IML methods come with theoretical foundations\. For example, IG satisfies axioms like sensitivity and implementation invariance\(Sundararajanet al\.,[2017](https://arxiv.org/html/2606.07607#bib.bib53)\); SHAP values are grounded in cooperative game theory with uniqueness guarantees\(Lundberg and Lee,[2017](https://arxiv.org/html/2606.07607#bib.bib54)\); DeepLIFT provides a principled decomposition of the prediction difference\(Shrikumaret al\.,[2017](https://arxiv.org/html/2606.07607#bib.bib51)\)\. However, theoretical guarantees do not translate to practical reliability\. The axioms satisfied by these methods concern mathematical properties of the attribution \(e\.g\., that attributions sum to the prediction difference\), not whether attributions identify biologically meaningful features\(Bilodeauet al\.,[2022](https://arxiv.org/html/2606.07607#bib.bib93)\)\. A method can satisfy all theoretical axioms while still highlighting spurious correlations learned by the model\. Moreover, different axiom\-satisfying methods can produce drastically different explanations for the same prediction\(Krishnaet al\.,[2024](https://arxiv.org/html/2606.07607#bib.bib89); Kindermanset al\.,[2019](https://arxiv.org/html/2606.07607#bib.bib94); Ghorbaniet al\.,[2019](https://arxiv.org/html/2606.07607#bib.bib2)\), as we demonstrate empirically in Section[3\.1](https://arxiv.org/html/2606.07607#S3.SS1)\.
Visual Inspection\.A common practice in the community is to have biologists visually inspect IML outputs on selected examples and assess whether the highlighted regions “make sense” given prior knowledge\(Alipanahiet al\.,[2015](https://arxiv.org/html/2606.07607#bib.bib45); Zhou and Troyanskaya,[2015](https://arxiv.org/html/2606.07607#bib.bib47); Avsecet al\.,[2021b](https://arxiv.org/html/2606.07607#bib.bib43)\)\. This approach suffers from several critical flaws\. First, confirmation bias is inevitable: practitioners are more likely to notice and report cases where explanations match expectations, while dismissing or ignoring contradictory cases as noise or edge cases\(Kauret al\.,[2020](https://arxiv.org/html/2606.07607#bib.bib83); Lakkaraju and Bastani,[2020](https://arxiv.org/html/2606.07607#bib.bib84); Adebayoet al\.,[2018](https://arxiv.org/html/2606.07607#bib.bib19)\)\. Such selective reporting provides no information about the method’s reliability in the broader scenarios\. Second, visual inspection cannot detect subtle failures—an explanation might highlight a region near but not exactly at the true binding site, which appears correct visually but would fail quantitative evaluation\. Finally, this approach cannot assess faithfulness: an explanation might align with biology by coincidence while not reflecting what the model actually learned\(Lapuschkinet al\.,[2019](https://arxiv.org/html/2606.07607#bib.bib85)\)\.
Accuracy Metrics\.If a model achieves high accuracy on test data, one might assume that its explanations must capture true biological signals—otherwise, how could it predict well? This reasoning suggests that evaluation efforts should focus on model accuracy rather than explanation quality\. Yet, models can achieve high accuracy through spurious correlations that happen to be predictive in the training distribution but do not reflect causal mechanisms\. In genomics, this is particularly concerning: GC content, dinucleotide frequencies, or positional biases can be highly predictive of certain regulatory outcomes without corresponding to specific functional elements\(Ghandiet al\.,[2014](https://arxiv.org/html/2606.07607#bib.bib91); Avsecet al\.,[2021b](https://arxiv.org/html/2606.07607#bib.bib43); Vorontsovet al\.,[2025](https://arxiv.org/html/2606.07607#bib.bib90)\)\. A model exploiting such shortcuts would produce explanations highlighting these confounders rather than true biological signals\(Geirhoset al\.,[2020](https://arxiv.org/html/2606.07607#bib.bib92); Novakovskyet al\.,[2023](https://arxiv.org/html/2606.07607#bib.bib87)\)\. Furthermore, even when a model does learn true signals, IML methods may fail to surface them faithfully, as we demonstrate in Section[3\.2](https://arxiv.org/html/2606.07607#S3.SS2)\. Predictive accuracy and explanation quality are orthogonal properties that must be evaluated independently\(Adebayoet al\.,[2018](https://arxiv.org/html/2606.07607#bib.bib19); Rudin,[2019](https://arxiv.org/html/2606.07607#bib.bib88); Sasseet al\.,[2023](https://arxiv.org/html/2606.07607#bib.bib66)\)\.
IML benchmarks\.The machine learning community has developed various benchmarks for evaluating IML methods in general domains, including synthetic datasets with known ground truth\(Hookeret al\.,[2019](https://arxiv.org/html/2606.07607#bib.bib21)\)and comprehensive evaluation toolkits\(Hedströmet al\.,[2023](https://arxiv.org/html/2606.07607#bib.bib95); Agarwalet al\.,[2022](https://arxiv.org/html/2606.07607#bib.bib30)\)\. While general\-purpose benchmarks are valuable, they do not capture the unique challenges of genomics applications\(Huanget al\.,[2025b](https://arxiv.org/html/2606.07607#bib.bib109)\)\. Biological sequences have distinct properties—combinatorial motif grammars, long\-range dependencies, strand symmetries, and compositional biases—that are absent in natural images or tabular data\(Koo and Eddy,[2019](https://arxiv.org/html/2606.07607#bib.bib96); Greensideet al\.,[2018](https://arxiv.org/html/2606.07607#bib.bib56); Avsecet al\.,[2021b](https://arxiv.org/html/2606.07607#bib.bib43)\)\. IML methods may succeed on standard benchmarks while failing on genomics\-specific challenges\. For instance, a method might correctly identify important pixels in an image but fail to localize a short motif embedded in a longer regulatory sequence\. Domain\-specific evaluation using biologically meaningful ground truth is essential to validate IML methods for genomics\(Sasseet al\.,[2023](https://arxiv.org/html/2606.07607#bib.bib66); Fenget al\.,[2025](https://arxiv.org/html/2606.07607#bib.bib97)\)\.
Wet\-lab validation\.While wet\-lab validation often provides the strongest evidence, italonecannot serve as the primary evaluation strategy for several reasons\. First, experimental validation is expensive and low\-throughput—it is infeasible to validate more than a handful of predictions per study\. Second, selective validation of high\-confidence predictions creates survivorship bias: we only learn about cases where explanations were correct, not the \(potentially larger\) set of failures\. Third, without computational pre\-screening, wet\-lab resources may be wasted on unreliable explanations\. Computational tests serve as essential gatekeepers that identify when explanations are likely unreliable*before*committing experimental resources\. The ideal workflow uses computational validation to filter and prioritize candidates for subsequent experimental confirmation\.
Our position\.None of the above alternatives provides adequate assurance that IML explanations are reliable for guiding biological interpretation or experimental design\. We therefore argue that practitioners should adopt a multi\-layered evaluation strategy as described in Section[4\.4](https://arxiv.org/html/2606.07607#S4.SS4)\. Only through such rigorous evaluation can we move beyond anecdotal evidence and establish genuine confidence in IML\-derived biological insights\.
## 6Conclusion
Our results highlight a growing mismatch between how interpretability methods are used in genomic modeling and how rigorously their outputs are evaluated\. Although post hoc IML techniques are often assumed to reveal meaningful model reasoning, our analysis shows that explanations frequently disagree across methods, only weakly reflect model behavior, and often fail to recover known biological mechanisms\. These issues are largely obscured by current norms that emphasize qualitative plausibility and selectively reported successes\.
This gap matters in practice\. Interpretability results increasingly inform experimental prioritization, mechanistic claims, and downstream biological hypotheses\. When explanations are unreliable, they risk misdirecting experimental effort and overstating model understanding\. We therefore argue that interpretability should be evaluated as an empirical object in its own right\. Systematic and rigorous quantitative validation is necessary for interpretability methods to support reliable scientific inference in genomics rather than anecdotal justification\.
## Acknowledgments
We sincerely thank all the reviewers for their encouraging and constructive feedback\. This work was supported by the UKRI Future Leaders Fellowship under Grant MR/S017062/1 and MR/X011135/1; in part by NSFC under Grant 62376056 and 62076056; in part by the Royal Society Faraday Discovery Fellowship \(FDF/S2/251014\), BBSRC TransformativeResearch Technologies \(UKRI1875\), Royal Society International Exchanges Award \(IES/R3/243136\), Kan Tong Po Fellowship \(KTP/R1/231017\); and the Amazon Research Award and Alan Turing Fellowship\. We also acknowledge the compute support provided by Modal\.
## References
- J\. Adebayo, J\. Gilmer, M\. Muelly, I\. J\. Goodfellow, M\. Hardt, and B\. Kim \(2018\)Sanity checks for saliency maps\.InNeurIPS’18: Proc\. of Advances in Neural Information Processing Systems 31,pp\. 9525–9536\.Cited by:[§5](https://arxiv.org/html/2606.07607#S5.p3.1),[§5](https://arxiv.org/html/2606.07607#S5.p4.1)\.
- C\. Agarwal, S\. Krishna, E\. Saxena, M\. Pawelczyk, N\. Johnson, I\. Puri, M\. Zitnik, and H\. Lakkaraju \(2022\)OpenXAI: towards a transparent evaluation of model explanations\.InNeurIPS’22: Proc\. of Advances in Neural Information Processing Systems 35,Cited by:[§5](https://arxiv.org/html/2606.07607#S5.p5.1)\.
- B\. Alipanahi, A\. Delong, M\. T\. Weirauch, and B\. J\. Frey \(2015\)Predicting the sequence specificities of dna\- and rna\-binding proteins by deep learning\.Nat\. Biotechnol\.33\(8\),pp\. 831–838\.Cited by:[§1](https://arxiv.org/html/2606.07607#S1.p1.1),[§1](https://arxiv.org/html/2606.07607#S1.p3.1),[§5](https://arxiv.org/html/2606.07607#S5.p3.1)\.
- Ž\. Avsec, V\. Agarwal, D\. Visentin, J\. R\. Ledsam, A\. Grabska\-Barwinska, K\. R\. Taylor, Y\. Assael, J\. Jumper, P\. Kohli, and D\. R\. Kelley \(2021a\)Effective gene expression prediction from sequence by integrating long\-range interactions\.Nat\. Methods18\(10\),pp\. 1196–1203\.Cited by:[§1](https://arxiv.org/html/2606.07607#S1.p1.1),[§1](https://arxiv.org/html/2606.07607#S1.p3.1)\.
- Ž\. Avsec, M\. Weilert, A\. Shrikumar, S\. Krueger, A\. Alexandari, K\. Dalal, R\. Fropf, C\. McAnany, J\. Gagneur, A\. Kundaje, and J\. Zeitlinger \(2021b\)Base\-resolution models of transcription\-factor binding reveal soft motif syntax\.Nat\. Genet\.53\(3\),pp\. 354–366\.Cited by:[§1](https://arxiv.org/html/2606.07607#S1.p1.1),[§1](https://arxiv.org/html/2606.07607#S1.p3.1),[§3](https://arxiv.org/html/2606.07607#S3.p2.1),[§5](https://arxiv.org/html/2606.07607#S5.p3.1),[§5](https://arxiv.org/html/2606.07607#S5.p4.1),[§5](https://arxiv.org/html/2606.07607#S5.p5.1)\.
- L\. Barbadilla\-Martínez, N\. Klaassen, B\. van Steensel, and J\. de Ridder \(2025\)Predicting gene expression from DNA sequence using deep learning models\.Nat\. Rev\. Genet\.\.Cited by:[§1](https://arxiv.org/html/2606.07607#S1.p1.1)\.
- B\. L\. Bilodeau, N\. Jaques, P\. W\. Koh, and B\. Kim \(2022\)Impossibility theorems for feature attribution\.abs/2212\.11870\.External Links:2212\.11870Cited by:[§5](https://arxiv.org/html/2606.07607#S5.p2.1)\.
- S\. Boshar, B\. Evans, Z\. Tang, A\. Picard, Y\. Adel, F\. K\. Lorbeer, C\. Rajesh, T\. Karch, S\. Sidbon, D\. Emms,et al\.\(2025\)A foundational model for joint sequence\-function multi\-species modeling at scale for long\-range genomic prediction\.bioRxiv,pp\. 2025–12\.Cited by:[§3](https://arxiv.org/html/2606.07607#S3.p6.1)\.
- T\. Bricken, A\. Templeton, J\. Batson, B\. Chen, A\. Jermyn, T\. Conerly, N\. Turner, C\. Anil, C\. Denison, A\. Askell,et al\.\(2023\)Towards monosemanticity: decomposing language models with dictionary learning\.Transformer Circuits Thread\.Cited by:[§1](https://arxiv.org/html/2606.07607#S1.p3.1)\.
- G\. Brixi, M\. G\. Durrant, J\. Ku, M\. Naghipourfar, M\. Poli, G\. Sun, G\. Brockman, D\. Chang, A\. Fanton, G\. A\. Gonzalez, S\. H\. King, D\. B\. Li, A\. T\. Merchant, E\. Nguyen, C\. Ricci\-Tam, D\. W\. Romero, J\. C\. Schmok, A\. Taghibakhshi, A\. Vorontsov, B\. Yang, M\. Deng, L\. Gorton, N\. Nguyen, N\. K\. Wang, M\. T\. Pearce, E\. Simon, E\. Adams, Z\. J\. Amador, E\. A\. Ashley, S\. A\. Baccus, H\. Dai, S\. Dillmann, S\. Ermon, D\. Guo, M\. H\. Herschl, R\. Ilango, K\. Janik, A\. X\. Lu, R\. Mehta, M\. R\. K\. Mofrad, M\. Y\. Ng, J\. Pannu, C\. Ré, J\. St\. John, J\. Sullivan, J\. Tey, B\. Viggiano, K\. Zhu, G\. Zynda, D\. Balsam, P\. Collison, A\. B\. Costa, T\. Hernandez\-Boussard, E\. Ho, M\. Liu, T\. McGrath, K\. Powell, S\. Pinglay, D\. P\. Burke, H\. Goodarzi, P\. D\. Hsu, and B\. L\. Hie \(2026\)Genome modelling and design across all domains of life with evo 2\.Nature652\(8112\),pp\. 1349–1361\.External Links:ISBN 1476\-4687Cited by:[§1](https://arxiv.org/html/2606.07607#S1.p3.1)\.
- V\. Chen, M\. Yang, W\. Cui, J\. S\. Kim, A\. Talwalkar, and J\. Ma \(2024\)Applying interpretable machine learning in computational biology—pitfalls, recommendations and opportunities for new developments\.Nat\. Methods21\(8\),pp\. 1454–1461\.Cited by:[§1](https://arxiv.org/html/2606.07607#S1.p2.1)\.
- M\. E\. Consens, C\. Dufault, M\. Wainberg, D\. Forster, M\. Karimzadeh, H\. Goodarzi, F\. J\. Theis, A\. Moses, and B\. Wang \(2025\)Transformers and genome language models\.Nat\. Mac\. Intell\.7\(3\),pp\. 346–362\.Cited by:[§1](https://arxiv.org/html/2606.07607#S1.p1.1)\.
- H\. Cunningham, A\. Ewart, L\. Riggs, R\. Huben, and L\. Sharkey \(2024\)Sparse autoencoders find highly interpretable features in language models\.InICLR’24: Proc\. of the 12th International Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2606.07607#S1.p3.1)\.
- B\. P\. de Almeida, F\. Reiter, M\. Pagani, and A\. Stark \(2022\)DeepSTARR predicts enhancer activity from dna sequence and enables the de novo design of synthetic enhancers\.Nat\. Genet\.54\(5\),pp\. 613–624\.Cited by:[§1](https://arxiv.org/html/2606.07607#S1.p1.1),[§1](https://arxiv.org/html/2606.07607#S1.p3.1)\.
- G\. Eraslan, Ž\. Avsec, J\. Gagneur, and F\. J\. Theis \(2019\)Deep learning: new computational modelling techniques for genomics\.Nat\. Rev\. Genet\.20\(7\),pp\. 389–403\.Cited by:[§1](https://arxiv.org/html/2606.07607#S1.p1.1)\.
- H\. Feng, L\. Wu, B\. Zhao, C\. Huff, J\. Zhang, J\. Wu, L\. Lin, P\. Wei, and C\. Wu \(2025\)Benchmarking dna foundation models for genomic and genetic tasks\.Nature Commun\.16\(1\),pp\. 10780\.Cited by:[§3](https://arxiv.org/html/2606.07607#S3.p2.1),[§5](https://arxiv.org/html/2606.07607#S5.p5.1)\.
- D\. M\. Fowler, C\. L\. Araya, S\. J\. Fleishman, E\. H\. Kellogg, J\. J\. Stephany, D\. Baker, and S\. Fields \(2010\)High\-resolution mapping of protein sequence\-function relationships\.Nat\. Methods7\(9\),pp\. 741–746\.Cited by:[§1](https://arxiv.org/html/2606.07607#S1.p1.1)\.
- H\. Gao, T\. Hamp, J\. Ede, J\. G\. Schraiber, J\. McRae, M\. Singer\-Berk, Y\. Yang, A\. S\. D\. Dietrich, P\. P\. Fiziev, L\. F\. K\. Kuderna, L\. Sundaram, Y\. Wu, A\. Adhikari, Y\. Field, C\. Chen, S\. Batzoglou, F\. Aguet, G\. Lemire, R\. Reimers, D\. Balick, M\. C\. Janiak, M\. Kuhlwilm, J\. D\. Orkin, S\. Manu, A\. Valenzuela, J\. Bergman, M\. Rousselle, F\. E\. Silva, L\. Agueda, J\. Blanc, M\. Gut, D\. de Vries, I\. Goodhead, R\. A\. Harris, M\. Raveendran, A\. Jensen, I\. S\. Chuma, J\. E\. Horvath, C\. Hvilsom, D\. Juan, P\. Frandsen, F\. R\. de Melo, F\. Bertuol, H\. Byrne, I\. Sampaio, I\. Farias, J\. V\. do Amaral, M\. Messias, M\. N\. F\. da Silva, M\. Trivedi, R\. Rossi, T\. Hrbek, N\. Andriaholinirina, C\. J\. Rabarivola, A\. Zaramody, C\. J\. Jolly, J\. Phillips\-Conroy, G\. Wilkerson, C\. Abee, J\. H\. Simmons, E\. Fernandez\-Duque, S\. Kanthaswamy, F\. Shiferaw, D\. Wu, L\. Zhou, Y\. Shao, G\. Zhang, J\. D\. Keyyu, S\. Knauf, M\. D\. Le, E\. Lizano, S\. Merker, A\. Navarro, T\. Bataillon, T\. Nadler, C\. C\. Khor, J\. Lee, P\. Tan, W\. K\. Lim, A\. C\. Kitchener, D\. Zinner, I\. Gut, A\. Melin, K\. Guschanski, M\. H\. Schierup, R\. M\. D\. Beck, G\. Umapathy, C\. Roos, J\. P\. Boubli, M\. Lek, S\. Sunyaev, A\. O’Donnell\-Luria, H\. L\. Rehm, J\. Xu, J\. Rogers, T\. Marques\-Bonet, and K\. K\. Farh \(2023\)The landscape of tolerated genetic variation in humans and primates\.Science380\(6648\),pp\. eabn8153\.Cited by:[§1](https://arxiv.org/html/2606.07607#S1.p1.1)\.
- R\. Geirhos, J\. Jacobsen, C\. Michaelis, R\. Zemel, W\. Brendel, M\. Bethge, and F\. A\. Wichmann \(2020\)Shortcut learning in deep neural networks\.Nat\. Mach\. Intell\.2\(11\),pp\. 665–673\.Cited by:[§5](https://arxiv.org/html/2606.07607#S5.p4.1)\.
- M\. Ghandi, D\. Lee, M\. Mohammad\-Noori, and M\. A\. Beer \(2014\)Enhanced regulatory sequence prediction using gapped k\-mer features\.PLoS Comput\. Biol\.10\(7\),pp\. e1003711\.Cited by:[§5](https://arxiv.org/html/2606.07607#S5.p4.1)\.
- A\. Ghorbani, A\. Abid, and J\. Y\. Zou \(2019\)Interpretation of neural networks is fragile\.InAAAI’19: Proc\. of the 33th AAAI Conference on Artificial Intelligence,pp\. 3681–3688\.Cited by:[§5](https://arxiv.org/html/2606.07607#S5.p2.1)\.
- P\. Greenside, T\. Shimko, P\. Fordyce, and A\. Kundaje \(2018\)Discovering epistatic feature interactions from neural network models of regulatory dna sequences\.Bioinformatics34\(17\),pp\. i629–i637\.External Links:ISSN 1367\-4803Cited by:[§1](https://arxiv.org/html/2606.07607#S1.p3.1),[§5](https://arxiv.org/html/2606.07607#S5.p5.1)\.
- A\. Hedström, L\. Weber, D\. Krakowczyk, D\. Bareeva, F\. Motzkus, W\. Samek, S\. Lapuschkin, and M\. M\.\-C\. Höhne \(2023\)Quantus: an explainable AI toolkit for responsible evaluation of neural network explanations and beyond\.J\. Mach\. Learn\. Res\.24,pp\. 34:1–34:11\.Cited by:[§5](https://arxiv.org/html/2606.07607#S5.p5.1)\.
- S\. Hooker, D\. Erhan, P\. Kindermans, and B\. Kim \(2019\)A benchmark for interpretability methods in deep neural networks\.InNeurIPS’19: Proc\. of Advances in Neural Information Processing Systems 32,pp\. 9734–9745\.Cited by:[§5](https://arxiv.org/html/2606.07607#S5.p5.1)\.
- T\. Hu, Y\. Cui, B\. Luo, and K\. Li \(2026\)RIDER: 3d RNA inverse design with reinforcement learning\-guided diffusion\.ICLR’26: Proc\. of the 14th International Conference on Learning Representations\.Cited by:[§1](https://arxiv.org/html/2606.07607#S1.p1.1)\.
- M\. Huang, S\. Zhou, Y\. Chen, and K\. Li \(2025a\)Conversational exploration of literature landscape with litchat\.InIJCAI’25: Proc, of the 34th International Joint Conference on Artificial Intelligence,,pp\. 11058–11061\.Cited by:[§2\.1](https://arxiv.org/html/2606.07607#S2.SS1.p3.1)\.
- M\. Huang, S\. Zhou, and K\. Li \(2025b\)Augmenting biological fitness prediction benchmarks with landscapes features from graphfla\.NeurIPS’25: Proc\. of the Conference on Neural Information Processing Systems\.Cited by:[§5](https://arxiv.org/html/2606.07607#S5.p5.1)\.
- J\. D\. Janizek, P\. Sturmfels, and S\. Lee \(2021\)Explaining explanations: axiomatic feature interactions for deep networks\.J\. Mach\. Learn\. Res\.22,pp\. 104:1–104:54\.Cited by:[§1](https://arxiv.org/html/2606.07607#S1.p3.1)\.
- H\. Kaur, H\. Nori, S\. Jenkins, R\. Caruana, H\. M\. Wallach, and J\. W\. Vaughan \(2020\)Interpreting interpretability: understanding data scientists’ use of interpretability tools for machine learning\.InCHI ’20: Proc\. of the CHI Conference on Human Factors in Computing Systems,pp\. 1–14\.Cited by:[§5](https://arxiv.org/html/2606.07607#S5.p3.1)\.
- P\. Kindermans, S\. Hooker, J\. Adebayo, M\. Alber, K\. T\. Schütt, S\. Dähne, D\. Erhan, and B\. Kim \(2019\)The \(un\)reliability of saliency methods\.InExplainable AI: Interpreting, Explaining and Visualizing Deep Learning,Lecture Notes in Computer Science, Vol\.11700,pp\. 267–280\.Cited by:[§5](https://arxiv.org/html/2606.07607#S5.p2.1)\.
- P\. K\. Koo and S\. R\. Eddy \(2019\)Representation learning of genomic sequence motifs with convolutional neural networks\.PLoS Comput\. Biol\.15\(12\),pp\. e1007560\.Cited by:[§5](https://arxiv.org/html/2606.07607#S5.p5.1)\.
- S\. Krishna, T\. Han, A\. Gu, S\. Wu, S\. Jabbari, and H\. Lakkaraju \(2024\)The disagreement problem in explainable machine learning: A practitioner’s perspective\.Trans\. Mach\. Learn\. Res\.2024\.Cited by:[§5](https://arxiv.org/html/2606.07607#S5.p2.1)\.
- H\. Lakkaraju and O\. Bastani \(2020\)”How do I fool you?”: manipulating user trust via misleading black box explanations\.InAIES ’20: Proc\. of the AAAI/ACM Conference on AI, Ethics, and Society,pp\. 79–85\.Cited by:[§5](https://arxiv.org/html/2606.07607#S5.p3.1)\.
- S\. Lapuschkin, S\. Wäldchen, A\. Binder, G\. Montavon, W\. Samek, and K\. Müller \(2019\)Unmasking clever hans predictors and assessing what machines really learn\.Nat\. Commun\.10\(1\),pp\. 1096\.Cited by:[§5](https://arxiv.org/html/2606.07607#S5.p3.1)\.
- S\. M\. Lundberg and S\. Lee \(2017\)A unified approach to interpreting model predictions\.InNIPS’17: Proc\. of Advances in Neural Information Processing Systems 30,pp\. 4765–4774\.Cited by:[§1](https://arxiv.org/html/2606.07607#S1.p3.1),[§5](https://arxiv.org/html/2606.07607#S5.p2.1)\.
- A\. Melnikov, A\. Murugan, X\. Zhang, T\. Tesileanu, L\. Wang, P\. Rogov, S\. Feizi, A\. Gnirke, C\. G\. Callan, J\. B\. Kinney, M\. Kellis, E\. S\. Lander, and T\. S\. Mikkelsen \(2012\)Systematic dissection and optimization of inducible enhancers in human cells using a massively parallel reporter assay\.Nat\. Biotechnol\.30\(3\),pp\. 271–277\.Cited by:[§1](https://arxiv.org/html/2606.07607#S1.p1.1)\.
- E\. Nguyen, M\. Poli, M\. Faizi, A\. Thomas, M\. Wornow, C\. Birch\-Sykes, S\. Massaroli, A\. Patel, C\. Rabideau, Y\. Bengio,et al\.\(2023\)Hyenadna: long\-range genomic sequence modeling at single nucleotide resolution\.Cited by:[§3](https://arxiv.org/html/2606.07607#S3.p6.1)\.
- G\. Novakovsky, N\. Dexter, M\. W\. Libbrecht, W\. W\. Wasserman, and S\. Mostafavi \(2023\)Obtaining genetics insights from deep learning via explainable artificial intelligence\.Nat\. Rev\. Genet\.24\(2\),pp\. 125–137\.External Links:ISBN 1471\-0064Cited by:[§1](https://arxiv.org/html/2606.07607#S1.p2.1),[§5](https://arxiv.org/html/2606.07607#S5.p4.1)\.
- V\. Petsiuk, A\. Das, and K\. Saenko \(2018\)Rise: randomized input sampling for explanation of black\-box models\.arXiv preprint arXiv:1806\.07421\.Cited by:[item 2](https://arxiv.org/html/2606.07607#S3.I2.i2.p1.1)\.
- M\. T\. Ribeiro, S\. Singh, and C\. Guestrin \(2016\)” Why should i trust you?” explaining the predictions of any classifier\.InKDD’16: Proc\. of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining,pp\. 1135–1144\.Cited by:[§3](https://arxiv.org/html/2606.07607#S3.p6.1)\.
- C\. Rudin \(2019\)Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead\.Nat\. Mach\. Intell\.1\(5\),pp\. 206–215\.Cited by:[§5](https://arxiv.org/html/2606.07607#S5.p4.1)\.
- W\. Samek, A\. Binder, G\. Montavon, S\. Lapuschkin, and K\. Müller \(2016\)Evaluating the visualization of what a deep neural network has learned\.IEEE Trans\. Neural Netw\. Learn\. Syst\.28\(11\),pp\. 2660–2673\.Cited by:[item 1](https://arxiv.org/html/2606.07607#S3.I2.i1.p1.1),[§3\.2](https://arxiv.org/html/2606.07607#S3.SS2.p1.1)\.
- A\. Sasse, B\. Ng, A\. E\. Spiro, S\. Tasaki, D\. A\. Bennett, C\. Gaiteri, P\. L\. De Jager, M\. Chikina, and S\. Mostafavi \(2023\)Benchmarking of deep neural networks for predicting personal gene expression from dna sequence highlights shortcomings\.Nat\. Genet\.55\(12\),pp\. 2060–2064\.Cited by:[§5](https://arxiv.org/html/2606.07607#S5.p4.1),[§5](https://arxiv.org/html/2606.07607#S5.p5.1)\.
- E\. E\. Seitz, D\. M\. McCandlish, J\. B\. Kinney, and P\. K\. Koo \(2024\)Interpreting cis\-regulatory mechanisms from genomic deep neural networks using surrogate models\.Nat\. Mach\. Intell\.6\(6\),pp\. 701–713\.Cited by:[§1](https://arxiv.org/html/2606.07607#S1.p3.1)\.
- A\. Shrikumar, P\. Greenside, and A\. Kundaje \(2017\)Learning important features through propagating activation differences\.InICML’17: Proc\. of the 34th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.70,pp\. 3145–3153\.Cited by:[§1](https://arxiv.org/html/2606.07607#S1.p3.1),[§3](https://arxiv.org/html/2606.07607#S3.p6.1),[§5](https://arxiv.org/html/2606.07607#S5.p2.1)\.
- K\. Simonyan, A\. Vedaldi, and A\. Zisserman \(2014\)Deep inside convolutional networks: visualising image classification models and saliency maps\.InICLR’14 Workshop: Workshop Track Proc\. of the 2nd International Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2606.07607#S1.p3.1)\.
- M\. Sundararajan, A\. Taly, and Q\. Yan \(2017\)Axiomatic attribution for deep networks\.InICML’17: Proc\. of the 34th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.70,pp\. 3319–3328\.Cited by:[§1](https://arxiv.org/html/2606.07607#S1.p3.1),[§3](https://arxiv.org/html/2606.07607#S3.p6.1),[§5](https://arxiv.org/html/2606.07607#S5.p2.1)\.
- E\. D\. Vaishnav, C\. G\. de Boer, J\. Molinet, M\. Yassour, L\. Fan, X\. Adiconis, D\. A\. Thompson, J\. Z\. Levin, F\. A\. Cubillos, and A\. Regev \(2022\)The evolution, evolvability and engineering of gene regulatory dna\.Nature603\(7901\),pp\. 455–463\.External Links:ISBN 1476\-4687Cited by:[§1](https://arxiv.org/html/2606.07607#S1.p1.1)\.
- J\. Vig, A\. Madani, L\. R\. Varshney, C\. Xiong, R\. Socher, and N\. F\. Rajani \(2021\)BERTology meets biology: interpreting attention in protein language models\.InICLR’21: Proc\. of the 9th International Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2606.07607#S1.p3.1)\.
- I\. E\. Vorontsov, I\. Kozin, S\. Abramov, A\. Boytsov, A\. Jolma, M\. Albu, G\. Ambrosini, K\. Faltejskova, A\. J\. Gralak, N\. Gryzunov, S\. Inukai, S\. Kolmykov, P\. Kravchenko, J\. F\. Kribelbauer\-Swietek, K\. U\. Laverty, V\. Nozdrin, Z\. M\. Patel, D\. Penzar, M\. Plescher, S\. E\. Pour, R\. Razavi, A\. W\. H\. Yang, I\. Yevshin, A\. Zinkevich, M\. T\. Weirauch, P\. Bucher, B\. Deplancke, O\. Fornes, J\. Grau, I\. Grosse, F\. A\. Kolpakov, M\. Barazandeh, A\. Brechalov, Z\. Deng, A\. Fathi, C\. Hu, S\. A\. Lambert, M\. Salnikov, I\. Yellan, H\. Zheng, G\. Meshcheryakov, M\. Nikonov, V\. Kamenets, A\. Vlasov, A\. Hernandez\-Corchado, H\. S\. Najafabadi, Q\. Morris, X\. Chen, V\. J\. Makeev, T\. R\. Hughes, I\. V\. Kulakovskiy, and T\. C\. Consortium \(2025\)Cross\-platform motif discovery and benchmarking to explore binding specificities of poorly studied human transcription factors\.Commun\. Biol\.8\(1\),pp\. 1545\.Cited by:[§3](https://arxiv.org/html/2606.07607#S3.p2.1),[§5](https://arxiv.org/html/2606.07607#S5.p4.1)\.
- H\. Yang, R\. Chen, and K\. Li \(2025\)Bridging sequence\-structure alignment in RNA foundation models\.InAAAI’25: Proc\. of the Thirty\-Ninth AAAI Conference on Artificial Intelligence,T\. Walsh, J\. Shah, and Z\. Kolter \(Eds\.\),pp\. 21929–21937\.Cited by:[§1](https://arxiv.org/html/2606.07607#S1.p1.1)\.
- Z\. Zhang, H\. K\. Wayment\-Steele, G\. Brixi, H\. Wang, D\. Kern, and S\. Ovchinnikov \(2024\)Protein language models learn evolutionary statistics of interacting sequence motifs\.Proceedings of the National Academy of Sciences121\(45\),pp\. e2406285121\.Cited by:[§3](https://arxiv.org/html/2606.07607#S3.p6.1)\.
- J\. Zhou, C\. L\. Theesfeld, K\. Yao, K\. M\. Chen, A\. K\. Wong, and O\. G\. Troyanskaya \(2018\)Deep learning sequence\-based ab initio prediction of variant effects on expression and disease risk\.Nat\. Genet\.50\(8\),pp\. 1171–1179\.Cited by:[§1](https://arxiv.org/html/2606.07607#S1.p1.1)\.
- J\. Zhou and O\. G\. Troyanskaya \(2015\)Predicting effects of noncoding variants with deep learning–based sequence model\.Nat\. Methods12\(10\),pp\. 931–934\.Cited by:[§1](https://arxiv.org/html/2606.07607#S1.p1.1),[§1](https://arxiv.org/html/2606.07607#S1.p3.1),[§3](https://arxiv.org/html/2606.07607#S3.p6.1),[§5](https://arxiv.org/html/2606.07607#S5.p3.1)\.
- Z\. Zhou, Y\. Ji, W\. Li, P\. Dutta, R\. Davuluri, and H\. Liu \(2022\)Dnabert\-2: efficient foundation model and benchmark for multi\-species genome\.Cited by:[§3](https://arxiv.org/html/2606.07607#S3.p6.1)\.
- J\. Zou, M\. Huss, A\. Abid, P\. Mohammadi, A\. Torkamani, and A\. Telenti \(2019\)A primer on deep learning in genomics\.Nat\. Genet\.51\(1\),pp\. 12–18\.Cited by:[§1](https://arxiv.org/html/2606.07607#S1.p1.1)\.
## Appendix ALLM\-based Structure Extraction Prompt for the Mapping Study
We usegemini\-3\-flash\-previewto extract structured information about interpretable machine learning \(IML\) evaluation and reporting practices from full\-text genomics papers\. The whole pipeline can be divided into two stages\. At the first stage, we use a few\-shot in\-context learning approach to extract the structured information from the paper\. At the second stage, we use a self\-reflection pass to verify the extraction against the actual paper text and correct any errors\.
### A\.1Stage 1: Few\-shot In\-context Learning for Structured Extraction
#### A\.1\.1System Instruction
``
`A\.1\.2 Task Prompt A\.1\.3 Few\-Shot Demonstrations Ten manually curated demonstrations are provided below\. Each consists of an abbreviated paper description and the corresponding gold\-standard extraction \(annotated by an author\)\. A\.1\.4 Input Format The full text of the paper is provided below the prompt, enclosed in delimiters: A\.2 Stage 2: Self\-Reflection Prompt After Stage 1 returns a JSON extraction for a paper, we will input the paper text and the JSON output back to the LLM to perform the self\-reflection\. The LLM will re\-read the paper and verify each field in the extraction against the actual paper text\. It will then correct any hallucinations, miscounts, and misclassifications that slip through Stage 1\. A\.2\.1 System Instruction A\.2\.2 Task Prompt Appendix B Per\-Category Agreement Rates of LLM\-based Extraction To assess the reliability of the LLM\-based extraction pipeline described in Appendix A, one of the authors manually annotated a random sample of 100100 papers and compared the manual labels against the LLM outputs across the four extraction categories used in our mapping study\. The per\-category agreement rates are reported in Table 2\. The overall agreement rate, averaged across all categories, is 98%98\\%\. Table 2: Agreement rates between manual annotation and LLM extraction across evaluated categories\. Extraction Category Agreement Rate \(%\) Number of IML methods 100 Validation strategy type 96 Number of reported instances 97 Failure reporting \(yes/no\) 99 Appendix C Predictive Performance of Foundation Models on TF Binding Prediction Table 3 reports the test\-set predictive performance \(accuracy and F1 score\) of the three genomic foundation models—DNABERT\-2, HyenaDNA, and NTv3—fine\-tuned on the TF binding prediction task across the five transcription factors \(CTCF, MAX, SP1, TBP, and GATA1\) introduced in Section 3\. These results establish that all three models attain non\-trivial predictive performance, which is a prerequisite for the subsequent interpretability audit: the explanations evaluated in our study are derived from models that have learned meaningful patterns from the data\. Table 3: Predictive performance \(Accuracy and F1 score\) of the three foundation models on TF binding prediction across five transcription factors\. Model CTCF MAX SP1 TBP GATA1 Acc F1 Acc F1 Acc F1 Acc F1 Acc F1 DNABERT\-2 0\.804 0\.823 0\.587 0\.684 0\.774 0\.772 0\.797 0\.770 0\.832 0\.821 HyenaDNA 0\.916 0\.916 0\.876 0\.875 0\.790 0\.794 0\.797 0\.780 0\.915 0\.916 NTv3 0\.950 0\.951 0\.885 0\.888 0\.872 0\.876 0\.853 0\.849 0\.941 0\.942 Appendix D Per\-Model and Per\-TF Breakdown of IML Disagreement Fig\. 5 and Fig\. 6 report the inter\-method Spearman rank correlation and top\-2020 Jaccard similarity, broken down by foundation model \(DNABERT\-2, HyenaDNA, NTv3\) and by transcription factor \(CTCF, MAX, SP1, TBP, GATA1\)\. These breakdowns complement the model\- and TF\-aggregated heatmaps in Fig\. 2b and c, and show that the low pairwise agreement reported in 3\.1 is not driven by a particular model or TF\. Figure 5: Per\-model and per\-TF Spearman rank correlation between IML method pairs\. Each subplot reports the pairwise Spearman correlation matrix over the five interpretability methods \(DeepLIFT, IG, ISM, LIME, CJ\) for one combination of foundation model \(DNABERT\-2, HyenaDNA, NTv3\) and transcription factor \(CTCF, MAX, SP1, TBP, GATA1\)\. Figure 6: Per\-model and per\-TF top\-2020 Jaccard similarity between IML method pairs\. Each subplot reports the pairwise Jaccard similarity of the top\-2020 attributed positions over the five interpretability methods for one combination of foundation model \(DNABERT\-2, HyenaDNA, NTv3\) and transcription factor \(CTCF, MAX, SP1, TBP, GATA1\)\. Appendix E Per\-Model Faithfulness Evaluation Fig\. 7, Fig\. 8, and Fig\. 9 report the deletion\-AUC and insertion\-AUC distributions on DNABERT\-2, HyenaDNA, and NTv3, respectively\. Figure 7: Faithfulness AUC distributions on DNABERT\-2\. Distribution of mean AUC scores across all test sequences, stratified by transcription factor \(CTCF, MAX, SP1, TBP, GATA1\) and IML method, for \(a\) the deletion experiment \(lower AUC indicates better faithfulness\) and \(b\) the insertion experiment \(higher AUC indicates better faithfulness\)\. Figure 8: Faithfulness AUC distributions on HyenaDNA\. Distribution of mean AUC scores across all test sequences, stratified by transcription factor \(CTCF, MAX, SP1, TBP, GATA1\) and IML method, for \(a\) the deletion experiment \(lower AUC indicates better faithfulness\) and \(b\) the insertion experiment \(higher AUC indicates better faithfulness\)\. Figure 9: Faithfulness AUC distributions on NTv3\. Distribution of mean AUC scores across all test sequences, stratified by transcription factor \(CTCF, MAX, SP1, TBP, GATA1\) and IML method, for \(a\) the deletion experiment \(lower AUC indicates better faithfulness\) and \(b\) the insertion experiment \(higher AUC indicates better faithfulness\)\. Appendix F Per\-Model Biological Alignment Fig\. 10 reports the per\-TF and per\-IML method Perception scores on DNABERT\-2, HyenaDNA, and NTv3\. Figure 10: Per\-model alignment between IML explanations and biological ground truth\. Distribution of motif overlap \(Perception\) scores across the five transcription factors \(CTCF, MAX, SP1, TBP, GATA1\) and the five IML methods \(DeepLIFT, IG, ISM, LIME, CJ\), shown separately for \(a\) DNABERT\-2, \(b\) HyenaDNA, and \(c\) NTv3\. Higher values indicate better alignment with UniBind\-annotated binding sites\.`Similar Articles
GENEB: Why Genomic Models Are Hard to Compare
GENEB is a large-scale diagnostic benchmark that evaluates 40 genomic foundation models across 100 tasks in 13 functional categories under a unified probing protocol, exposing that aggregate leaderboards are unstable and that architectural alignment often outweighs model scale. The work addresses the fragmented evaluation landscape in genomic machine learning, analogous to what MTEB did for NLP.
The Dark Regulome: Disentangling Predictability from Regulation in Genomic Foundation Models
This paper introduces a residualization-and-permutation diagnostic to separate predictability-driven from regulation-driven variance in regulatory importance scores from genomic foundation models, applied to dark genome elements at glioma-relevant loci.
Interpretability Can Be Actionable
This position paper argues that interpretability research should be evaluated based on actionability—the extent to which insights enable concrete decisions and interventions. The authors propose a framework with evaluation criteria aligned with practical outcomes to address the lack of real-world impact in current interpretability work.
@DivyanshT91162: Microsoft Research just dropped a paper that completely flips interpretability on its head. (bookmark this) For years, …
Microsoft Research introduced Agentic-iModels, a framework where coding agents evolve scikit-learn regressors optimized for LLM interpretability rather than human readability, outperforming traditional interpretable ML methods across 65 datasets.
Probing, Fusion, and Trustworthiness: A Systematic Evaluation of Foundation Model Representations for Multimodal Cancer Analysis
This paper systematically evaluates foundation model representations for multimodal cancer analysis, benchmarking unimodal and multimodal fusion strategies on real-world cohorts, and assessing trustworthiness via conformal prediction.