CASCADE: An Agentic Regulatory Network Framework for Patient-Data-Validated Downstream Perturbation Prediction

arXiv cs.AI Papers

Summary

CASCADE is an agentic framework that predicts downstream transcriptional effects of gene perturbations using regulatory networks, with patient-data validation across TCGA cancer types. The paper demonstrates directional concordance for MYC and other regulators, though it does not exceed existing curated knowledge.

arXiv:2608.05359v1 Announce Type: new Abstract: CASCADE is an agentic framework that predicts downstream transcriptional effects of gene perturbation from precomputed ARACNe regulatory networks, exposed via MCP. Prior work validates such tools by checking whether predicted genes are known cancer genes (membership); we instead test whether the predicted direction of change matches reality, using focal-gene copy-number amplification as a dosage-based proxy for the inverse of knockdown against real TCGA patient tumor data. For MYC, CASCADE's predicted knockdown targets show strong concordance with real amplified-vs-non-amplified tumor expression across three cancer types (BRCA: 90.0%, COAD: 72.0%, STAD: 85.7%; all p<0.0013), well above permutation baselines, surviving a PAM50 subtype control and replicating in an independent cohort (METABRIC, 87.2%). Compared against curated MSigDB gene-set baselines via Fisher's exact test, CASCADE's accuracy is not shown to exceed existing public knowledge of MYC- or E2F-driven biology, though its gene-specific direction-calling clearly outperforms a naive uniform guess. Extending to fifteen additional genes, validation proves gene-specific rather than universal: proliferation-machinery regulators mostly replicate, while lineage-identity transcription factors and one cyclin-D paralog (CCND2) consistently fail, a pattern we discuss as a hedged, post-hoc hypothesis. We separately benchmark whether an LLM-based agent correctly grounds natural-language requests into CASCADE's real MCP tool calls. Across 35 queries, a documented local model reaches 71.4% exact match (85.7% for a larger model); schema and gene-alias failures are resolved by scale or server-side correction, but both models confidently default to the wrong perturbation type on ambiguous queries, a failure a targeted fix could not resolve because its trigger condition never occurs.
Original Article
View Cached Full Text

Cached at: 08/07/26, 07:46 AM

# CASCADE: An Agentic Regulatory Network Framework for Patient-Data-Validated Downstream Perturbation Prediction
Source: [https://arxiv.org/html/2608.05359](https://arxiv.org/html/2608.05359)
Jose A\. Bird ORCID: 0009\-0006\-2744\-0606 Affiliation: Independent Researcher Correspondence: jbird@birdaisolutions\.com

###### Abstract

CASCADE is an agentic framework that predicts downstream transcriptional effects of gene perturbation from precomputed ARACNe regulatory networks, exposed through a LangGraph\-orchestrated workflow and the Model Context Protocol \(MCP\)\. Because this perturbation\-prediction capability is CASCADE’s central function – the specific claim an MCP\-exposed tool call would surface to a downstream agent or user – we introduce a patient\-data validation methodology that tests whether that claim is trustworthy, not merely plausible\. Every experiment in this paper invokesCascadeWorkflow\.run\(\)directly withanalysis\_depth="focused", CASCADE’s real, public agentic entry point\. We confirm empirically, not merely by reading the routing code, exactly which internal analyses this triggers for the gene roles tested here \(Section[2\.1](https://arxiv.org/html/2608.05359#S2.SS1)\); only the resulting perturbation\-effects output is used in this paper’s statistical tests\.

CASCADE extends the multi\-agent orchestration pattern introduced by RegNetAgents\(Bird,[2026](https://arxiv.org/html/2608.05359#bib.bib1)\)to a complementary analytical direction\. Rather than classifying a focal gene’s*upstream*regulators by cross\-network source, CASCADE predicts the*downstream*transcriptional consequences of perturbing a focal gene via directed propagation through tumor\-specific ARACNe networks\. Validating downstream predictions is harder than validating regulator\-candidate lists, because it requires either sparse experimental perturbation data or a genotype\-driven natural\-experiment proxy\. We use the latter – focal\-gene copy\-number amplification as a dosage\-based proxy for the inverse of knockdown – tested against real patient tumor expression data rather than curated cancer\-gene annotation lists\.

For MYC, CASCADE’s predicted knockdown targets show strong directional concordance with real patient expression differences between amplified and non\-amplified tumors, across three TCGA cancer types \(BRCA: 90\.0%,p<10−8p<10^\{\-8\}; COAD: 72\.0%,p=0\.0013p=0\.0013; STAD: 85\.7%,p<0\.0001p<0\.0001\)\. Permutation baselines confirm these results are not artifacts of gene\-set structure, landing within four points of the theoretical 50% chance level in every case\. The result survives a PAM50 breast\-cancer\-subtype control \(92% concordance within Luminal B alone\), and it replicates in METABRIC, an independent cohort with no data overlap with CASCADE’s network\-construction data \(87\.2%,p<10−6p<10^\{\-6\}\)\. We also compared CASCADE against two curated, identity\-matched gene\-set baselines using Fisher’s exact test \(Section[3\.2](https://arxiv.org/html/2608.05359#S3.SS2)\)\. An exhaustive curated MYC\-target gene set scores within 1–4 percentage points of CASCADE across all three cancer types \(ahead in BRCA and STAD, behind in COAD\), and CASCADE’s E2F3 rate is nominally higher than its curated comparison \(96\.0% vs\. 90\.7%\)\. Neither difference reaches statistical significance \(p=0\.34p=0\.34–0\.860\.86andp=0\.384p=0\.384respectively\)\. CASCADE’s own direction\-calling clearly matters – forcing its predictions to a uniform guess drops concordance by 14–29 points in every panel tested – but its overall accuracy is not shown to exceed existing public knowledge of MYC\- or E2F\-driven biology \(Section[3\.2](https://arxiv.org/html/2608.05359#S3.SS2)\)\.

Beyond MYC, validation is gene\-specific rather than universal\. Seven further proliferation\-machinery\-associated regulators \(E2F3, CCND1, AURKA, CCNE1, FOXM1, TOP2A, RPS6KB1\) replicate the pattern cleanly in BRCA \(89\.6–100% each\), and two more \(MDM2: 68%, CCND3: 76%\) replicate more weakly\. Four lineage\-identity transcription factors \(SOX9, FOXA1, GATA3, ESR1: 38\.8%, 34\.0%, 0\.0%, 4\.0%\), a receptor tyrosine kinase \(ERBB2, 18\.0% in BRCA, further inconsistent across cancer types below\), and CCND2 do not\. CCND2 is a cyclin\-D paralog of the two CCND genes that do validate \(CCND1, CCND3\), and fails outright in the same cancer type \(BRCA\) where both paralogs succeed – the clearest single\-gene counterexample to a clean proliferation\-machinery/lineage\-identity rule\. ESR1 was selected specifically because the proliferation\-machinery/lineage\-identity hypothesis below predicted it would fail, which it did\. Extending seven of these proliferation\-machinery genes to a second cancer type, where amplification frequency allowed it, shows the split is not purely gene\-class\-determined\. AURKA \(COAD, 92\.0%\), CCNE1 \(STAD, 98\.0%\), TOP2A \(STAD, 100\.0%\), and CCND3 \(STAD, 79\.6%\) replicate outside BRCA\. CCND1 \(STAD, 57\.1%\) and MDM2 \(STAD, 42\.0%\) do not, unlike in BRCA where both replicate; ERBB2 shows the same cancer\-type\-dependent pattern in the opposite direction – anti\-concordant in BRCA and COAD \(18\.0%, 14\.3%\) but concordant in STAD \(93\.6%\)\. CCND2 \(COAD, 54\.2%\) also fails to replicate, but unlike CCND1 and MDM2 it never validated in BRCA either – the only proliferation\-machinery gene in the panel that fails in every cancer type tested\. Section[4](https://arxiv.org/html/2608.05359#S4)discusses this split as a hedged, post\-hoc hypothesis, not an established mechanism\.

The concordance result above validates CASCADE’s predictions once correctly invoked; we separately test the step upstream of that, whether an LLM\-based agent correctly grounds a natural\-language request into CASCADE’s real MCP tool\-call parameters\. Across 35 hand\-labeled queries spanning gene aliases, informal cancer\-type names, non\-canonical perturbation phrasing, implicit/default parameters, and multi\-entity distractors, CASCADE’s own documented local\-deployment default \(llama3\.1:8bvia Ollama\) reaches 71\.4% exact parameter\-match on its raw output, and a larger local model \(qwen2\.5:72b\) reaches 85\.7%\. The gap is not uniform: schema\-adherence failures \(a required parameter left unset or misrouted\) occur almost entirely with the smaller model and are fully resolved by the larger one, and gene\-alias failures – though model\-size\-invariant in each model’s raw output – resolve correctly for both models once CASCADE’s ownresolve\_alias\(\)step is applied server\-side \(the same step CASCADE’s real MCP server runs before any network lookup\)\. A third failure mode does not resolve, however: both models confidently default to the wrong perturbation type on maximally ambiguous queries, and a targeted server\-side fix cannot help, because its trigger condition – the field being left unset – never actually occurs \(Section[3\.6](https://arxiv.org/html/2608.05359#S3.SS6)\)\.

## 1Introduction

Predicting the downstream consequences of perturbing a gene is a central task in cancer systems biology, informing both mechanistic hypotheses and therapeutic target prioritization\(Califano & Alvarez,[2017](https://arxiv.org/html/2608.05359#bib.bib2)\)\. Tools built on precomputed regulatory networks can generate such predictions cheaply, without new experimental data or model training – but this efficiency comes at a validation cost\. The strongest evidence would be a held\-out experimental perturbation screen; lacking that, network\-propagation tools are typically validated instead against curated gene\-annotation databases, the approach RegNetAgents\(Bird,[2026](https://arxiv.org/html/2608.05359#bib.bib1)\)itself takes, via enrichment for cancer\-gene annotation among candidate genes\. That check only asks whether a predicted candidate happens to already be a known cancer gene \(*membership*validation\) – it does not ask whether the prediction’s specific directional claim, this gene goes up or that one goes down, matches what actually happens in real disease\. This distinction has real consequences: when such a tool is exposed as a callable function to an LLM\-based agent, as CASCADE is via MCP, the agent’s report to a user inherits whatever validation gap the underlying tool carries, and membership validation cannot close it, since it never checks the specific directional claim the agent goes on to repeat\.

RegNetAgents\(Bird,[2026](https://arxiv.org/html/2608.05359#bib.bib1)\)demonstrated one rigorous path to database\-level validation: querying two independently\-inferred ARACNe network sources for a focal gene’s*upstream*regulators, classifying candidates by network origin, and showing via Fisher’s exact enrichment \(with permutation controls and negative\-control gene panels\) that the resulting candidate lists are significantly enriched for OncoKB\-annotated\(Chakravarty et al\.,[2017](https://arxiv.org/html/2608.05359#bib.bib12)\)cancer genes across two cancer types\. That validation strategy is well suited to a regulator\-*identification*task, where the object of interest is a ranked list of candidate genes to be cross\-referenced against curated annotation\.

CASCADE addresses a different, complementary task: given a focal gene and a perturbation type \(knockdown or overexpression\), predict the*downstream*transcriptional consequences by propagating a signed perturbation signal through a precomputed regulatory network\. This is architecturally similar to RegNetAgents – both are implemented as LangGraph\-orchestrated workflows over precomputed ARACNe networks, exposed via MCP for natural\-language queries – but the object being validated is fundamentally different: not ”is this candidate a known cancer gene” but ”does this specific predicted direction of change actually happen\.” Curated\-annotation enrichment, the validation strategy that worked cleanly for RegNetAgents, is a weaker test for this kind of claim, because it can only check whether predicted genes are plausible cancer genes in general, not whether the predicted*direction*of change is correct\. We instead validate against real patient tumor data directly, so that the specific claim an agent surfaces when it calls this tool is one we have directly checked against reality, not merely inferred to be plausible from network topology or gene\-set membership\.

CASCADE’s downstream perturbation\-prediction task relates to a broader literature on computational perturbation\-effect prediction\. VIPER\(Alvarez et al\.,[2016](https://arxiv.org/html/2608.05359#bib.bib15)\), built on the same ARACNe/regulatory\-network lineage as CASCADE’s networks, infers protein activity from regulon enrichment in observed gene expression – the reverse of CASCADE’s task, which predicts the transcriptional consequence of a hypothetical perturbation rather than inferring activity from data already observed\. GEARS\(Roohani et al\.,[2023](https://arxiv.org/html/2608.05359#bib.bib17)\)and CPA\(Lotfollahi et al\.,[2023](https://arxiv.org/html/2608.05359#bib.bib18)\)address a task closer to CASCADE’s: predicting transcriptional responses to gene perturbations, including unseen combinations\. Both are learned models trained on large single\-cell perturbational screens, giving them access to combinatorial and dose\-dependent effects that CASCADE’s static network propagation does not model, at the cost of requiring substantial training data and infrastructure that CASCADE’s precomputed\-network approach avoids\. We are not aware of a prior application of any of these approaches to real patient tumor genotype\-expression data in the manner reported here; this paper’s contribution is a validation methodology for CASCADE specifically, not a claim that its propagation\-based predictions are more accurate than learned alternatives – a head\-to\-head comparison we did not attempt and consider open future work\.

### 1\.1Contribution

This paper makes two related but distinct contributions: a patient\-data concordance methodology for validating CASCADE’s downstream perturbation predictions, and a benchmark of whether an LLM\-based agent correctly grounds natural\-language requests into CASCADE’s real MCP tool\-call parameters:

1. 1\.A directional concordance methodology – focal\-gene copy\-number amplification as a proxy for the inverse of knockdown, tested against real patient genotype\-expression data with permutation\-based significance – that validates downstream perturbation predictions independent of curated cancer\-gene annotation, and generalizes in principle to any regulatory\-network\-based perturbation tool\.
2. 2\.For MYC: a strong, permutation\-controlled concordance signal that survives a PAM50 subtype control and replicates in an independent cohort \(METABRIC\) with no data\-provenance overlap with CASCADE’s networks\.
3. 3\.A fifteen\-gene generalization test showing this validation is gene\-specific: nine proliferation\-machinery regulators replicate \(seven cleanly, two weakly\) while six do not – four lineage\-identity transcription factors, a receptor tyrosine kinase, and CCND2, a cyclin\-D paralog of two validating genes\. Section[4](https://arxiv.org/html/2608.05359#S4)discusses the resulting split as a post\-hoc hypothesis, not an established mechanism\.
4. 4\.A benchmark of whether an LLM\-based agent correctly grounds natural\-language requests into CASCADE’s real MCP tool\-call parameters, distinct from whether CASCADE’s predictions are accurate once invoked: schema\-adherence failures are model\-size\-dependent, gene\-alias failures are resolved server\-side by CASCADE’s alias table, and a third failure mode – confidently defaulting to the wrong perturbation type on maximally ambiguous queries – persisted through a targeted fix, because both models tested always supply a guess rather than leaving the field unset \(Section[3\.6](https://arxiv.org/html/2608.05359#S3.SS6)\)\.

Section[2\.5](https://arxiv.org/html/2608.05359#S2.SS5)\(methodology\) and Section[3\.6](https://arxiv.org/html/2608.05359#S3.SS6)\(results\) cover the agentic tool\-call grounding contribution; the remaining Methods and Results subsections cover the patient\-data validation methodology and results – ARACNe network construction, TCGA/METABRIC cohort details, and gene\-panel selection – and assume more domain background\.

## 2Methods

### 2\.1CASCADE architecture

CASCADE is a Python package exposing gene perturbation analysis through a LangGraph\-orchestrated workflow and an MCP server, following the same general orchestration pattern as RegNetAgents\(Bird,[2026](https://arxiv.org/html/2608.05359#bib.bib1); LangChain AI,[2024](https://arxiv.org/html/2608.05359#bib.bib3)\): a directed workflow classifies a focal gene’s regulatory role from precomputed network topology, routes to appropriate analyses, and executes independent sub\-analyses concurrently before assembling a structured report\. Every result in this paper uses CASCADE’s actual default perturbation\-propagation behavior for TCGA networks: network propagation blended with GREmLN embedding\-based cosine similarity\(Zhang et al\.,[2025](https://arxiv.org/html/2608.05359#bib.bib14)\)via a weighted parameter \(α=0\.7\\alpha=0\.7\), with automatic fallback to network\-only propagation only if the embedding model is unavailable \(not the case for any result reported here\)\. Every experiment in this paper invokesCascadeWorkflow\.run\(\)\(below\) directly\. The network component is a deterministic breadth\-first algorithm over the ARACNe network’s directed regulator→\\totarget edges: for a knockdown, the focal gene receives an initial effect of−1\.0\-1\.0, and at each hop the effect propagated to a target gene is the parent gene’s current effect multiplied by the edge weight and a decay factor of0\.50\.5, accumulating additively across paths\. For TCGA networks, the edge weight is the ARACNe mutual\-information magnitude signed by the edge’s mode\-of\-action annotation \(activating edges preserve sign, repressive edges flip it\) – this sign\-flipping is the primary mechanism producing the up/down direction predictions the concordance test in Section[2\.2](https://arxiv.org/html/2608.05359#S2.SS2)evaluates\.

The embedding component adds cosine\-similarity\-weighted contributions from CASCADE’s pre\-trained gene embeddings on top of this network signal, and can surface additional candidate genes reachable via embedding similarity but not network propagation; becauseα=0\.7\\alpha=0\.7keeps the network\-derived signal dominant, embedding blending rescales magnitude and can add new candidates but does not flip the direction of genes already reachable via the network\. All analyses in this paper use CASCADE’s own default propagation depth \(2 hops\) and default result size \(top 25 by\|effect\|\|\\text\{effect\}\|for the internal role\-classification pathway; the validation experiments described below instead use the top 50 by\|effect\|\|\\text\{effect\}\|, as specified per experiment\)\.

CASCADE ships with ARACNe regulatory networks for 14 TCGA cancer types, sourced from the Bioconductoraracne\.networkspackage\(Lim & Califano,[2018](https://arxiv.org/html/2608.05359#bib.bib4)\), itself built from TCGA RNA\-seq data downloaded April 2015 using ARACNe\-AP\(Lachmann et al\.,[2016](https://arxiv.org/html/2608.05359#bib.bib5)\)\. Networks are symbol\-native \(no Ensembl mapping required for TCGA network operations\) and carry per\-edge mode\-of\-action annotation \(activating/repressive\) derived from the same source\.

Before propagation, CASCADE’s workflow classifies the focal gene’s regulatory role from network topology \(Table[1](https://arxiv.org/html/2608.05359#S2.T1)\), which determines downstream routing within the full CASCADE system \(e\.g\. whether regulator analysis or protein\-interaction evidence is prioritized\)\. This classification is deterministic and threshold\-based, not learned; it is reported here for completeness of the architecture description, not as a claim evaluated in this paper\. Because this routing logic is deterministic code rather than a learned or LLM\-driven decision, we verify its correctness via conventional unit tests \(tests/test\_workflow\.py, covering role classification and batch\-dispatch routing across gene\-role and depth combinations\) rather than an empirical benchmark; the agent tool\-call grounding benchmark \(Section[2\.5](https://arxiv.org/html/2608.05359#S2.SS5)\) is reserved for the one step in CASCADE’s pipeline that does involve genuine model judgment under ambiguity\.

Table 1:CASCADE gene\-role classification rules \(cascade\_langgraph\_workflow\.py\)\.Figure[1](https://arxiv.org/html/2608.05359#S2.F1)summarizes this pipeline end to end\.\_decide\_next\_stepscomputes a role\- and depth\-conditional required\-analysis set and routes the workflow’s concurrent batch dispatch to cover exactly that set \(detailed below\); an optional LLM\-based node can synthesize a narrative interpretation of a completed report without altering any of its underlying scores, but this LLM node \(include\_llm\_insights\) is not enabled in any experiment in this paper\.

![Refer to caption](https://arxiv.org/html/2608.05359v1/x1.png)Figure 1:Pipeline exercised by this paper’s experiments: focal\-gene knockdown propagation via CASCADE’s agentic entry point,CascadeWorkflow\.run\(\), producing direction\-labeled predicted targets, tested for directional concordance against real patient copy\-number/expression data, with binomial and permutation\-based statistical validation\. CASCADE’s own routing logic concurrently computes additional non\-LLM analyses as a byproduct of this call \(grey panel\); these are not used in this paper’s statistics\.CASCADE’s public agentic API, and the exact call used directly by every experiment script in this paper, is:

```
from cascade_langgraph_workflow import CascadeWorkflow

workflow = CascadeWorkflow()
report = await workflow.run(
    gene="MYC",
    perturbation_type="knockdown",
    analysis_depth="focused",
    network_source="tcga",
    tcga_network="brca",
    top_k=50,
)
predicted_targets = report["perturbation_effects"]["top_affected_genes"]
```

top\_kis discussed further below; every argument shown here is part of CASCADE’s public API\.

We verified empirically, not merely by reading the routing code, exactly which of CASCADE’s internal analysesanalysis\_depth="focused"triggers for the gene roles tested in this paper\.\_decide\_next\_stepsclassifies the focal gene’s regulatory role and, for"focused", requiresperturbationplus eithertargets\(master\-regulator and transcription\-factor genes\) orppi\(other roles\); every gene\-cancer\-type combination tested in this paper is classified as master regulator or transcription factor \(Table[5](https://arxiv.org/html/2608.05359#S3.T5)\), so every call here requires exactly \{perturbation,targets\}, both members of the same core\-analysis batch group, and routing sends the call to that single batch node rather than the full three\-way concurrent dispatch\. Only theperturbation\_effectsfield of the returned report is extracted and used in this paper’s statistical tests; the concurrently\-computedtargetsanalysis is not evaluated further and is not reported as a result in this paper\.

CascadeWorkflow\.run\(\)accepts atop\_kparameter \(default 25, CASCADE’s single\-query default\), threaded through to every internal propagation call site; this paper’sN=50N=50concordance panels passtop\_k=50explicitly\. It likewise accepts apropagation\_depthparameter \(default 2, CASCADE’s default BFS hop count\), threaded through the same call sites; every result in this paper uses this default depth of 2\. Every experiment in this paper – the MYC primary panels, the PAM50 and METABRIC checks, and the fifteen\-gene generalization panel \(Section[3\.5](https://arxiv.org/html/2608.05359#S3.SS5)\) – usesanalysis\_depth="focused"; none uses"comprehensive", so the additional analyses"comprehensive"would require for master\-regulator/transcription\-factor genes \(vulnerability,lincs,regulators,similar,dorothea,depmap,cbioportal\) are not computed by any result reported here\. Table[5](https://arxiv.org/html/2608.05359#S3.T5)’s role classifications are also examined descriptively in Results \(Section[3\.5](https://arxiv.org/html/2608.05359#S3.SS5)\) against this paper’s biological concordance findings, beyond their role in determining routing here – though not as a systematically tested variable, as that section itself notes\.

### 2\.2Directional concordance methodology

For a focal geneggin a given TCGA cancer\-type network, CASCADE’s knockdown propagation yields a signed effect for every reachable downstream gene\. We take the topN=50N=50genes by\|effect\|\|\\text\{effect\}\|and record each gene’s predicted direction \(down if effect<0<0, up if effect\>0\>0\)\.

We treat focal\-gene copy\-number amplification \(GISTIC discrete value=2=2\) as an approximate dosage\-based proxy for increased focal\-gene activity – the inverse, in direction, of a knockdown\. For each predicted target gene, we compare its mean mRNA expressionzz\-score between amplified and non\-amplified \(GISTIC=0=0\) patient samples\. A gene predicted to go*down*upon knockdown \(i\.e\. positively regulated by the focal gene\) is scored*concordant*if its mean expression is*higher*in amplified samples; a gene predicted to go*up*\(negatively regulated\) is concordant if its mean expression is*lower*in amplified samples\. Genes with fewer than 10 samples in either group are excluded\.

Significance is assessed with a one\-tailed binomial test \(null: concordance rate=0\.5=0\.5\) and an independent permutation control: 1,000 draws ofNNgenes from a background pool of 200 randomly sampled network genes, excluding the focal gene itself and every gene in CASCADE’s own predicted\-target set for that run \(so the null distribution cannot be contaminated by the real signal it is meant to be compared against\), with real amplified/non\-amplified expression differences precomputed, each draw randomly assigned a predicted\-direction label matching the true up/down proportion of the real predicted set\. The empiricalpp\-value is the fraction of permuted concordance rates meeting or exceeding the observed rate\. This directly tests whether the observed concordance rate exceeds what arbitrary genes, given the same class balance, would produce by chance\.

### 2\.3Focal gene panel selection for generalization testing

Beyond MYC, we sought additional focal genes to test whether the directional concordance methodology generalizes\. Because the methodology requires focal\-gene amplification as the dosage proxy, candidate genes were required to satisfy two independent criteria in a given TCGA network: \(i\) network out\-degree≥25\\geq 25\(so the predicted\-target panel reaches full size\), and \(ii\) at least∼\\sim15–20 GISTIC\-amplified patient samples in the corresponding cBioPortal PanCancer Atlas study \(so amplified/non\-amplified group means are estimable\)\. We screened MYC, ERBB2, and seven additional transcription factors with documented amplification\-driven oncogenic roles in at least one tissue \(CTNNB1, GATA3, FOXA1, KLF5, SOX9, E2F3, MYB\) across BRCA, COAD, and STAD\.

Most lineage\-restricted transcription factors failed criterion \(ii\) in at least two of the three cancer types – e\.g\. CTNNB1 is activated predominantly by point mutation rather than amplification and had at most one amplified sample in any tested cancer type; GATA3 and FOXA1 similarly lacked sufficient amplified samples outside BRCA\. Only MYC and ERBB2 satisfied both criteria in all three cancer types; E2F3, SOX9, GATA3, and FOXA1 satisfied both criteria in BRCA only\. This selection process is itself a finding: strong, cross\-tissue, amplification\-driven dosage effects comparable to MYC are genuinely rare among transcription factors, which constrains the achievable size of any generalization panel built on this methodology\.

### 2\.4Data sources

Real\-patient expression and copy\-number data are drawn from the TCGA PanCancer Atlas 2018\(Hoadley et al\.,[2018](https://arxiv.org/html/2608.05359#bib.bib6)\), with copy\-number calls generated by GISTIC2\.0\(Mermel et al\.,[2011](https://arxiv.org/html/2608.05359#bib.bib7)\), accessed via the cBioPortal REST API\(Cerami et al\.,[2012](https://arxiv.org/html/2608.05359#bib.bib8); Gao et al\.,[2013](https://arxiv.org/html/2608.05359#bib.bib9)\)\. This API’s data is GDC\-harmonized against GRCh38 – a different reference genome and processing pipeline than the 2015 raw\-download data underlying CASCADE’s ARACNe networks, though very likely drawing on an overlapping TCGA patient cohort\. PAM50 molecular subtype annotations were obtained from the same cBioPortal study’s patient\-level clinical data\. Independent\-cohort replication uses METABRIC\(Curtis et al\.,[2012](https://arxiv.org/html/2608.05359#bib.bib10); Pereira et al\.,[2016](https://arxiv.org/html/2608.05359#bib.bib11)\), a breast cancer cohort of 2,509 tumors profiled by microarray at UK and Canadian institutions, with no data\-provenance relationship to TCGA\.

All fetches use cBioPortal’s batch endpoints \(genes/fetchfor symbol→\\toEntrez resolution,molecular\-data/fetchfor multi\-gene expression retrieval\), making each experiment a small, fixed number of API calls regardless of gene panel size\. Analyses are otherwise read\-only scripts against local network files and public APIs\.

### 2\.5Agent tool\-call grounding methodology

Every experiment above callsCascadeWorkflow\.run\(\)directly with hand\-specified parameters, which validates CASCADE’s predictions but not the step upstream of them: whether an LLM\-based agent, given only a natural\-language request and CASCADE’s real MCP tool schema, correctly fills in the parameters that call actually needs\.

This complements, at much smaller scale, an existing benchmark literature on LLM tool\-use and function\-calling generally: Gorilla\(Patil et al\.,[2023](https://arxiv.org/html/2608.05359#bib.bib21)\)and the Berkeley Function\-Calling Leaderboard\(Patil et al\.,[2025](https://arxiv.org/html/2608.05359#bib.bib22)\)evaluate schema\-adherence and API\-selection accuracy across large synthetic or aggregated API collections\. The benchmark below instead tests end\-to\-end natural\-language grounding against one real, deployed tool’s actual MCP schema, on 35 queries constructed specifically to probe CASCADE’s own failure surface \(gene aliases, informal cancer\-type names, ambiguous perturbation direction\) rather than to sample broadly across unrelated APIs\.

We test this separately and explicitly, using CASCADE’s actualcomprehensive\_perturbation\_analysistool definition \(verbatim fromcascade\_langgraph\_mcp\_server\.py\) against 35 hand\-labeled natural\-language queries spanning seven categories \(five queries each\): baseline TCGA\-cancer\-type requests, baseline immune\-cell\-type requests, gene aliases \(a deliberate mix of aliases resolvable via CASCADE’s alias table and one resolvable only via general reasoning, to test both cases, Section[3\.6](https://arxiv.org/html/2608.05359#S3.SS6)\), informal/full cancer\-type names, non\-canonical perturbation\-verb phrasing \(“silence,” “boost,” “suppress,” “delete function”\), queries that omit parameters entirely \(testing whether the agent falls back to CASCADE’s documented defaults\), and queries containing a distractor gene or cancer type not actually requested\. Ground truth for each query is the parameter set CASCADE’s tool would need to receive to behave correctly, with an omitted field scored against the default CASCADE’s own schema documents for it \(e\.g\.perturbation\_typedefaults toknockdown\), not against a bare absence\.

Two models are tested via a local Ollama server:llama3\.1:8b, CASCADE’s own documented default for its separate LLM\-insights feature \(Section[2\.1](https://arxiv.org/html/2608.05359#S2.SS1)\) and the more representative test of what a fully local/offline CASCADE deployment would use, andqwen2\.5:72b\-instruct\-q4\_0, a substantially larger local model, reported as a secondary comparison rather than the primary result to avoid presenting only the model that performs best\. Each query is sent once per model with the real tool schema attached; the model’s emitted tool\-call arguments are compared field\-by\-field against ground truth\.

CASCADE resolves a small, fixed set of common informal gene names server\-side \(e\.g\.HER2→\\toERBB2\) viatools/gene\_id\_mapper\.py’sresolve\_alias\(\)function, before any network lookup\. Because a model can still emit an alias’s literal informal name even though CASCADE’s table can resolve it server\-side, we score each gene\-alias query two ways: an exact match against the model’s raw tool\-call output, and a second match after running the model’sgenevalue through CASCADE’s realresolve\_alias\(\)function, the same step its MCP server runs before any network lookup \(Section[3\.6](https://arxiv.org/html/2608.05359#S3.SS6)\)\.

We also implemented a targeted fix for the perturbation\-type default\-handling failure this benchmark surfaced \(Section[3\.6](https://arxiv.org/html/2608.05359#S3.SS6)\): an optionalqueryparameter added tocomprehensive\_perturbation\_analysis’s schema, carrying the caller’s original natural\-language request, and a check in the server’s handler that scans this text for directional cues \(e\.g\. “knock down”/“silence” vs\. “overexpress”/“boost”, matched via a small regex list\) wheneverperturbation\_typeis omitted from the tool call, returning aclarification\_neededresponse instead of applying theknockdowndefault when no single direction is detected; an explicitly\-suppliedperturbation\_typevalue is never overridden\. This detection is a simple keyword scan and has a known blind spot we did not exercise in this benchmark: it cannot attribute a directional cue to a specific gene, so a multi\-entity query naming two genes with different directions \(e\.g\. “knocked down GATA3, now overexpress FOXA1”\) would read as conflicting cues even though the request is unambiguous once each verb is attributed to its gene\. None of the 35 queries tested exercises this case, so we report it here as an acknowledged gap in the detection logic rather than as a tested limitation of the results below\.

## 3Results

### 3\.1MYC knockdown predictions show strong, multi\-cancer\-type concordance with real patient data

Table[2](https://arxiv.org/html/2608.05359#S3.T2)summarizes the primary result, using CASCADE’s actual default embedding\-enhanced propagation \(α=0\.7\\alpha=0\.7; Section[2\.1](https://arxiv.org/html/2608.05359#S2.SS1)\)\. Across all three tested TCGA cancer types, CASCADE’s top\-50 predicted MYC\-knockdown targets show concordance rates 24–36 percentage points above their respective permutation baselines, each with permutation\-empiricalp=0\.0000p=0\.0000\(no permutation among 1,000 draws matched or exceeded the observed rate\)\. Zero candidates in any cancer type were embedding\-only additions \(genes reachable via embedding similarity but not network propagation\), consistent with MYC’s broad network connectivity already spanning the genes embedding similarity would otherwise surface\.

Table 2:MYC knockdown concordance with real patient MYC\-amplification status, by TCGA cancer type\.The concordant gene set in BRCA is biologically coherent, not merely statistically significant: top concordant genes include NOP14, NOB1, RIOK1, NOLC1, WDR43, TAF1D, SNRPD1, and KAT2A – ribosome biogenesis and nucleolar machinery components, MYC’s best\-established transcriptional program in cancer biology\(van Riggelen et al\.,[2010](https://arxiv.org/html/2608.05359#bib.bib16)\)\. This independent biological plausibility check is consistent with the statistical result reflecting genuine regulatory signal rather than an artifact of the test design\.

The result is not sensitive to the arbitrary choice of panel sizeN=50N=50: repeating the identical BRCA/MYC test atN=25N=25andN=100N=100yields 22/25 \(88\.0%, binomialp=0\.0001p=0\.0001\) and 92/100 \(92\.0%, binomialp<0\.0001p<0\.0001\) concordant respectively, both with permutation\-empiricalp≤0\.001p\\leq 0\.001– concordance rates of 88\.0%, 90\.0%, and 92\.0% atN=25N=25,5050, and100100respectively, essentially flat across a fourfold range in panel size\.

Table[3](https://arxiv.org/html/2608.05359#S3.T3)shows the fifteen highest\-magnitude predicted targets from a single MYC/BRCA query, illustrating the practical output the concordance test evaluates: for each predicted target, its CASCADE\-predicted direction and its actual mean expressionzz\-score in MYC\-amplified versus non\-amplified patients\. This top\-15 slice’s 13/15 \(86\.7%\) “down” share is not representative of the full top\-50 panel, which is considerably more balanced \(Appendix[A](https://arxiv.org/html/2608.05359#A1)\)\.

Table 3:Top 15 CASCADE\-predicted MYC knockdown targets \(BRCA\), by\|effect\|\|\\text\{effect\}\|, with real patient expression comparison\.Illustrative top\-15\-by\-magnitude slice; its 86\.7% “down” share is not representative of the full 50\-gene panel’s up/down balance, which is considerably more balanced \(Appendix[A](https://arxiv.org/html/2608.05359#A1)\)\.

Figure[2](https://arxiv.org/html/2608.05359#S3.F2)summarizes observed concordance rates against permutation baselines for every gene\-cancer\-type combination tested in this paper, including the additional\-gene generalization tests reported in Section[3\.5](https://arxiv.org/html/2608.05359#S3.SS5)\.

![Refer to caption](https://arxiv.org/html/2608.05359v1/x2.png)Figure 2:Observed concordance rate \(colored bars\) versus permutation\-derived chance baseline \(grey bars\), for every gene\-cancer\-type combination tested, including the cross\-cancer\-type extension for AURKA, CCND1, CCND2, CCND3, CCNE1, MDM2, and TOP2A\. Green bars: proliferation\-machinery\-associated genes \(MYC, E2F3, CCND1, CCND2, CCND3, AURKA, CCNE1, MDM2, FOXM1, TOP2A, RPS6KB1\)\. Red bars: lineage\-identity transcription factors \(SOX9, FOXA1, GATA3, ESR1\) or the receptor tyrosine kinase ERBB2\. Bar color reflects gene category, not whether that individual result was concordant – e\.g\. ERBB2/STAD is colored red \(category\) despite exceeding its own permutation baseline \(result\), and conversely CCND1/STAD, CCND2/BRCA, CCND2/COAD, and MDM2/STAD are colored green \(category\) despite not being significantly concordant \(result\); all instances of this inconsistency are discussed in Section[3\.5](https://arxiv.org/html/2608.05359#S3.SS5)\. MYC is concordant in all four contexts tested \(three TCGA cancer types plus the independent METABRIC cohort\); the proliferation\-machinery group validates in most but not all gene\-cancer\-type combinations \(AURKA, CCNE1, TOP2A, and CCND3 in a second cancer type; CCND1 and MDM2 only in BRCA; CCND2 not at all, in either cancer type tested, despite its paralogs CCND1 and CCND3 both validating\) while the lineage\-identity/RTK group shows mixed to strongly anti\-concordant results\. ESR1 was selected after, and specifically to test, the proliferation\-machinery/lineage\-identity hypothesis discussed in Section[4](https://arxiv.org/html/2608.05359#S4)– the only gene in this figure chosen because it was predicted to fail rather than to validate\.
### 3\.2A baseline comparison against curated, identity\-matched gene sets

A positive concordance result does not by itself show that CASCADE’s specific algorithm adds value beyond a public, non\-CASCADE gene list, so we checked this directly against two curated, identity\-matched gene lists – genes with well\-established roles as transcriptional activators \(MYC, and the activating E2Fs including E2F3\), such that predicting their target\-set members uniformly “down” upon knockdown is a biologically grounded assumption rather than an artificial convention invented for comparability – using each list’s full resolvable gene set and comparing its concordance rate against CASCADE’s own via Fisher’s exact test, which accounts for the two groups’ different sample sizes directly rather than treating both rates as equally precise \(full methodology and results in Appendix[A](https://arxiv.org/html/2608.05359#A1), Table[8](https://arxiv.org/html/2608.05359#A1.T8)\)\.

Against an exhaustive curated MYC\-target gene list, the gap is narrow and inconsistent in direction – 1–4 percentage points across all three cancer types, with the curated list ahead in BRCA and STAD and CASCADE ahead in COAD – and not statistically significant in any of the three \(Fisher’s exactp=0\.34p=0\.34–0\.860\.86\): CASCADE’s overall accuracy is not established to exceed an independent expert\-curated database, but equally not established to fall short of it\. For E2F3, CASCADE’s rate is nominally higher than its curated comparison list \(96\.0% vs\. 90\.7%\), but this difference is not statistically significant either \(p=0\.384p=0\.384\)\. None of the four comparisons survives Benjamini\-Hochberg correction across all four \(q=0\.767q=0\.767–0\.8630\.863, Table[8](https://arxiv.org/html/2608.05359#A1.T8)\)\. This correction is applied within its own four\-test family, separate from the twenty\-seven\-combination correction in Table[4](https://arxiv.org/html/2608.05359#S3.T4): the two test different claims – baseline superiority over a curated gene list here, versus cross\-gene generalization there – so we do not pool them into a single correction family\.

This is a separate question from whether CASCADE’s own direction\-calling matters: forcing every predicted gene to “down” lowers concordance by 14–29 percentage points in every panel tested, so CASCADE’s gene\-specific reasoning is doing substantial, measurable work relative to a naive version of itself\. Taken together, this comparison shows CASCADE’s MYC and E2F3 predictions are correct and non\-trivially derived, not that CASCADE’s overall accuracy is established to exceed existing public knowledge of MYC\- or E2F\-driven biology in either of the two gene\-identity comparisons tested\.

### 3\.3The result survives a PAM50 subtype control

MYC amplification is not uniformly distributed across PAM50 molecular subtypes in our BRCA cohort \(Basal 61 amplified/13 non\-amplified; LumA 36/251; LumB 32/33; Her2 18/20; Normal 6/21\), raising the possibility that predicted\-target concordance reflects subtype composition differences rather than MYC dosage itself\. Restricting the identical test to BRCA\_LumB alone – the subtype with the most balanced amplified/non\-amplified split, and therefore the best\-powered within\-subtype test available – yielded 46/50 concordant \(92\.0%, binomialp=2\.2×10−10p=2\.2\\times 10^\{\-10\}\), a result that*strengthened*rather than weakened relative to the full unstratified cohort\. This is inconsistent with a pure subtype\-confound explanation\.

### 3\.4Independent\-cohort replication rules out a shared\-data artifact

CASCADE’s TCGA ARACNe networks and the TCGA PanCancer Atlas expression/copy\-number data used above are both ultimately derived from the TCGA\-BRCA cohort, processed through different pipelines at different times \(2015 raw download vs\. 2018 GDC\-harmonized GRCh38\) but very likely overlapping in the underlying patients\. To rule out the possibility that real patient\-specific biology shared between these two data releases was inflating apparent validation, we repeated the identical concordance test against METABRIC\(Curtis et al\.,[2012](https://arxiv.org/html/2608.05359#bib.bib10); Pereira et al\.,[2016](https://arxiv.org/html/2608.05359#bib.bib11)\)– a cohort with no data\-provenance relationship to TCGA whatsoever \(different patients, different countries, microarray rather than RNA\-seq quantification\)\. CASCADE’s TCGA\-network\-derived MYC predictions were tested, unmodified, against METABRIC’s amplified \(n=554n=554\) and non\-amplified \(n=1,110n=1\{,\}110\) patient expression data: 41/47 testable genes were concordant \(87\.2%, binomialp<10−6p<10^\{\-6\}, permutation\-empiricalp=0\.0000p=0\.0000, permutation mean 49\.2%\) – a result essentially identical in magnitude to the original TCGA\-BRCA finding \(90\.0%\)\.

### 3\.5Validation is gene\-specific, not universal

We tested fifteen additional genes for generalization \(Table[4](https://arxiv.org/html/2608.05359#S3.T4)\)\. Genes entered the panel three ways, disclosed here against any appearance of post\-hoc cherry\-picking: an initial blind screen against the eligibility criteria in Section[2\.3](https://arxiv.org/html/2608.05359#S2.SS3)\(E2F3, ERBB2, SOX9, FOXA1, GATA3\); systematic screens of proliferation\-machinery/cell\-cycle candidates against those same unchanged criteria, testing every eligible gene regardless of expected outcome \(CCND1, AURKA, CCNE1, MDM2, CCND3, CCND2, FOXM1, TOP2A, RPS6KB1; four further candidates screened – CDK4, MYBL2, E2F1, BIRC5 – failed eligibility and were not tested\); and one gene chosen specifically to falsify rather than confirm the working hypothesis, ESR1, selected after the proliferation\-machinery/lineage\-identity hypothesis \(Section[4](https://arxiv.org/html/2608.05359#S4)\) had already formed, because that hypothesis predicted it would fail \(it did, Table[4](https://arxiv.org/html/2608.05359#S3.T4)\)\.

Of the ten proliferation\-machinery genes tested, nine validate somewhere \(E2F3, CCND1, CCND3, AURKA, CCNE1, MDM2, FOXM1, TOP2A, RPS6KB1; 68–100%, Table[4](https://arxiv.org/html/2608.05359#S3.T4)\) and only CCND2 does not\. None of the five lineage\-identity/RTK genes validate as a category \(GATA3, FOXA1, SOX9, ESR1, ERBB2\) – though ERBB2’s STAD result is a significant exception within that group \(93\.6%,q<0\.0001q<0\.0001; discussed below\)\. Transcription\-factor status does not separate the two groups – Section[4](https://arxiv.org/html/2608.05359#S4)walks through the gene\-by\-gene split, including how it differs from CASCADE’s own network\-derived role classification for some of these genes\.

CCND2 fails outright, unlike its paralogs, in both cancer types tested\.CCND2 is the sharpest exception to any clean gene\-class rule: eligible under criteria identical to CCND1 and CCND3 \(16 amplified samples, network out\-degree 34\), it fails outright \(19/50, 38\.0%,p=0\.9675p=0\.9675, below its own 43\.4% permutation baseline\) in BRCA, the cancer type where both paralogs validate\. Extended to COAD – where its out\-degree \(48\) and amplified\-sample count \(16\) independently clear the same eligibility criteria – CCND2 again fails to validate \(26/48, 54\.2%,p=0\.333p=0\.333, essentially chance\), while CCND3 replicates cleanly in its own second cancer type, STAD \(39/49, 79\.6%,p<0\.0001p<0\.0001\)\. CCND2 is therefore the only proliferation\-machinery gene in the panel that fails to validate in every cancer type tested\.

The cross\-cancer\-type extension surfaces further inconsistency\.Seven non\-MYC/ERBB2 proliferation\-machinery genes were extended to a second cancer type wherever eligibility criteria were met \(Section[2\.3](https://arxiv.org/html/2608.05359#S2.SS3)\): AURKA \(COAD\); CCND1, CCNE1, MDM2, and TOP2A \(STAD\); and CCND2 and CCND3 \(COAD and STAD respectively, above\)\. Four replicate outside BRCA \(AURKA 92\.0%, CCNE1 98\.0%, TOP2A 100\.0%, eachp<0\.0001p<0\.0001; CCND3 79\.6%,p<0\.0001p<0\.0001\), but three do not: CCND1 falls from 96\.0% in BRCA to 57\.1% in STAD \(p=0\.196p=0\.196\), MDM2 falls from 68\.0% in BRCA to 42\.0% in STAD \(p=0\.899p=0\.899, below its own permutation baseline\), and CCND2 fails in both cancer types tested \(above\)\. ERBB2 shows the same pattern even more starkly across all three cancer types tested: strong anti\-concordance in BRCA and COAD \(18\.0%, 14\.3%\) but strong concordance in STAD \(93\.6%\)\. Gene class alone is therefore not sufficient to predict validation; cancer\-type context matters too, for reasons this paper does not resolve\.

Ruling out an embedding\-blend artifact\.Because CASCADE’s default propagation blends network topology with GREmLN embedding similarity \(Section[2\.1](https://arxiv.org/html/2608.05359#S2.SS1)\), we checked whether that embedding component contributes unevenly across the two groups, which could confound the split reported above\. It does not: embedding\-only additions numbered 0 or 1 of 50 in every one of the twenty\-seven gene\-cancer\-type combinations tested, with no difference between validating and failing groups\. GATA3’s result – the most extreme in the panel \(0/50 concordant\) – was checked further under network\-only propagation, with the embedding component removed entirely: the result is an identical 0/50, confirming the anti\-concordance is a property of the network and its annotations, not of CASCADE’s embedding component\.

Table 4:Concordance for additional focal genes tested for generalization, including cross\-cancer\-type extensions for five proliferation\-machinery genes \(Section[3\.5](https://arxiv.org/html/2608.05359#S3.SS5)\)\. Category is a hypothesis\-driven grouping assigned by us \(Section[4](https://arxiv.org/html/2608.05359#S4)\), not a CASCADE output and not a direct function of the Type column: genes of the same molecular Type can fall in different Categories \(e\.g\. E2F3 and FOXM1 versus SOX9, FOXA1, GATA3, and ESR1 are all transcription factors by Type but split across both Categories\)\.“Replicates MYC” refers to concordance rate only, independent of the baseline comparison in Section[3\.2](https://arxiv.org/html/2608.05359#S3.SS2), which was run for E2F3 but not for the other genes in this table\. The Type column reports each gene’s established molecular function \(e\.g\. NCBI Gene, UniProt\) and is independent of, and not always identical to, CASCADE’s own network\-derived role classification \(Table[5](https://arxiv.org/html/2608.05359#S3.T5)\); it is not a CASCADE output\.

Candidate genes with strong, recurrent, cross\-tissue amplification comparable to MYC are themselves rare – lineage\-restricted transcription factors such as KLF5, CTNNB1, and MYB lack sufficient amplified\-sample counts in at least two of the three tested cancer types to run this test at all, which constrained the achievable panel size for this generalization check\. We report the fifteen\-gene, twenty\-seven\-combination comparison in Table[4](https://arxiv.org/html/2608.05359#S3.T4)as evidence bearing on generalizability, not as a comprehensive panel\-level statistic in the style of RegNetAgents’ eleven\- and twelve\-gene combined tests, and discuss a candidate explanation for the observed split in Section[4](https://arxiv.org/html/2608.05359#S4)\.

Table[4](https://arxiv.org/html/2608.05359#S3.T4)’s BH\-FDR column applies Benjamini\-Hochberg correction across all twenty\-seven gene\-cancer\-type combinations tested in this generalization sweep \(the twenty\-four rows shown plus MYC’s three cancer types from Table[2](https://arxiv.org/html/2608.05359#S3.T2)\): seventeen of twenty\-seven combinations remain significant atq<0\.05q<0\.05and ten do not, with CCND1/STAD and CCND2/COAD as the two intermediate cases \(q=0\.29q=0\.29andq=0\.47q=0\.47respectively\)\. The resulting proliferation\-machinery/lineage\-identity split should still be read as a pattern we observed and report, not a pre\-registered or formally tested hypothesis – and, as the cross\-cancer\-type extension above shows, gene class alone is not the whole story; CCND2’s outright failure alongside its validating paralogs, in both cancer types tested, is a further instance of that same point\.

Table[5](https://arxiv.org/html/2608.05359#S3.T5)applies CASCADE’s gene\-role classification \(Table[1](https://arxiv.org/html/2608.05359#S2.T1)\) to every gene\-cancer\-type combination tested in the patient\-data concordance experiments, reported descriptively as a record of what the classifier did, not as a variable systematically tested for association with outcome; with fifteen non\-MYC genes across twenty\-seven combinations \(Section[2\.3](https://arxiv.org/html/2608.05359#S2.SS3)\), any pattern here is an observation, not evidence\. Role classification does not predict validation outcome\. The sharpest single\-gene example is ERBB2: classified as a master regulator in both BRCA and STAD, with no change in role label, yet anti\-concordant in BRCA \(18\.0%\) and concordant in STAD \(93\.6%\) – the same gene, the same role, opposite outcomes, driven by cancer\-type context rather than anything CASCADE’s classifier captures\. The sharpest cross\-gene example is CCND2 and CCND3: both classified as transcription factors in BRCA with comparable out\-degree \(34 and 48\), the same paralog family, the same cancer type, and the same role label, yet CCND2 fails outright \(38\.0%, anti\-concordant\) while CCND3 validates \(76\.0%\)\. Together, these are sufficient to rule out role classification as a strict, standalone explanation, independent of sample size – though neither identifies what the actual explanation is\.

Table 5:CASCADE gene\-role classification \(Table[1](https://arxiv.org/html/2608.05359#S2.T1)\) applied to genes tested in Sections[3\.1](https://arxiv.org/html/2608.05359#S3.SS1)and[3\.5](https://arxiv.org/html/2608.05359#S3.SS5)\.
### 3\.6Agent tool\-call grounding is reliable with a capable model, with three disclosed failure modes – two resolved, one open

Overall accuracy\.Table[6](https://arxiv.org/html/2608.05359#S3.T6)reports exact parameter\-match accuracy across the 35\-query benchmark \(Section[2\.5](https://arxiv.org/html/2608.05359#S2.SS5)\), scored against each model’s raw tool\-call output\.llama3\.1:8b, CASCADE’s own documented local\-deployment default, reaches 71\.4% \(25/35\) exact match;qwen2\.5:72b, tested as a secondary, larger\-model comparison, reaches 85\.7% \(30/35\)\. The gap between them is not uniform across failure types, and each type has a distinct, disclosable explanation rather than being unexplained noise\. Ollama’s sampling is not seeded, sollama3\.1:8b’s exact figure and specific failing queries can vary between runs; we report the one run for which we retained full per\-query results throughout this section, the same run used to test the perturbation\-type fix described below\.

Table 6:Agent tool\-call grounding accuracy by category, exact parameter match out of 5 queries per category, scored against each model’s raw tool\-call output\.†Raw model output; scored instead against what CASCADE’s server would do after its ownresolve\_alias\(\)step, Gene alias rises to 3/5 \(llama3\.1:8b\) and 5/5 \(qwen2\.5:72b\), and Overall rises to 27/35 \(77\.1%\) and 33/35 \(94\.3%\) respectively \(Section[3\.6](https://arxiv.org/html/2608.05359#S3.SS6)\)\.

‡We additionally implemented and tested a fix for the ambiguous\-queryperturbation\_typedefault \(Section[3\.6](https://arxiv.org/html/2608.05359#S3.SS6)\)\. It triggered on 0 of all 35 queries, for both models – and both zeros share one root cause rather than being independent results: neither model ever leftperturbation\_typeunset, so the fix’s trigger condition \(a genuinely omitted field\) never occurred anywhere in the benchmark\. Consequently it flagged*0/3 of the ambiguous queries it was designed to catch*\(a failure to engage, not a null result\) and, for the identical reason,*0/32 of the remaining non\-ambiguous queries*\(no spurious clarifications, but this reflects the same never\-engaged mechanism, not a validated regression check\)\. The fix’s own optionalqueryparameter, needed for it to engage at all, was populated byllama3\.1:8bin 23/35 calls overall \(65\.7%, though never on the 3 ambiguous queries specifically\) and byqwen2\.5:72bin 0/35 calls\.

Network\-parameter errors\.Network\-parameter errors –tcga\_networkornetwork\_sourceleft unset or set to an incorrect value – occurred eight times withllama3\.1:8band zero times withqwen2\.5:72b\(Table[7](https://arxiv.org/html/2608.05359#S3.T7)\): six on queries outside the Gene\-alias category, discussed here, and two more on Gene\-alias\-category queries \(HER3, p53\), discussed separately below where they explainllama3\.1:8b’s two remaining post\-resolution failures\. Of the six discussed here, three are true schema\-adherence failures: a requiredtcga\_networkvalue omitted entirely despitenetwork\_sourcecorrectly set totcga, which would cause CASCADE’s real server to reject the call outright\. The remaining three are wrong\-value substitutions rather than omissions: two cancer\-subtype disambiguation errors in different tissue pairs, and onenetwork\_sourcesubstitution on an implicit\-parameter query \(“What does MYC do?”\) routed to a TCGA network instead of the implied cell\-type default – the same query discussed further below as this benchmark’s shared perturbation\-type ambiguity failure\. All eight errors are specific to the smaller model:qwen2\.5:72bproduced neither an omission nor a wrong\-value network\-parameter error anywhere in the benchmark\.

Table 7:Individual failing queries underlying the three failure modes discussed in this section, by model, field, and value\. Category names match Table[6](https://arxiv.org/html/2608.05359#S3.T6)’s columns directly\. A query with more than one field error \(e\.g\. the HER3 query below, or “What does MYC do?” forllama3\.1:8b, which recurs under both Network\-parameter and Perturbation\-type ambiguity errors\) appears once per error, so row counts within a category can exceed Table[6](https://arxiv.org/html/2608.05359#S3.T6)’s per\-category error count, which counts distinct failing queries\. “Server\-resolved” marks gene\-field mismatches that CASCADE’s realresolve\_alias\(\)step corrects \(Section[3\.6](https://arxiv.org/html/2608.05359#S3.SS6)\); it is not applicable \(N/A\) to non\-gene fields\.Gene\-alias failures: raw output vs\. server\-resolved\.Gene\-alias failures illustrate the gap between a model’s raw output and what CASCADE’s server actually does with it \(Table[7](https://arxiv.org/html/2608.05359#S3.T7)\)\. Scored against raw output,llama3\.1:8bgets 1/5 gene\-alias queries exactly right andqwen2\.5:72bgets 2/5: both emit the literal informal name rather than the network’s official symbol for three of the five queries, while both correctly produce the alias already in CASCADE’s table without needing correction \(p53→\\toTP53\) and the alias resolvable via general reasoning rather than a table lookup \(“the estrogen receptor”→\\toESR1\)\. Scored against what CASCADE’s server would actually do – running each model’s rawgenevalue through the realresolve\_alias\(\)function before comparing it – every one of these three aliases resolves correctly for both models, since all three are covered by CASCADE’s alias table \(Section[2\.5](https://arxiv.org/html/2608.05359#S2.SS5)\), lifting the category to 3/5 and 5/5 respectively \(Table[6](https://arxiv.org/html/2608.05359#S3.T6), note†\\dagger\)\.

llama3\.1:8b’s two remaining post\-resolution failures are not gene\-alias errors at all – both are network\-parameter errors on the same queries \(Table[7](https://arxiv.org/html/2608.05359#S3.T7)\)\. In other words, once CASCADE’s own resolution step is accounted for, no gene name in this benchmark is left genuinely unresolved for either model: what remains is explained entirely by CASCADE’s alias\-table coverage \(complete for every alias tested here\) and, separately, by network\-parameter errors unrelated to gene naming\.

A third failure mode: perturbation\-type ambiguity\.A third failure mode – defaulting to the wrongperturbation\_typeon queries with no explicit directional cue – affected both models, overlapping on one query and diverging on another \(Table[7](https://arxiv.org/html/2608.05359#S3.T7)\): bothllama3\.1:8bandqwen2\.5:72bdefaulted tooverexpressioninstead of CASCADE’s documentedknockdowndefault on the same query \(“What does MYC do?”\), andqwen2\.5:72bshowed the identical pattern separately on a second query thatllama3\.1:8banswered correctly\. That the pattern recurs across both models – on a query they share and one specific to the larger, more capable model – suggests this reflects a genuine mismatch between how such requests read naturally and what CASCADE’s schema silently defaults to – though with only three ambiguous queries tested, this is a suggestive pattern, not a robust estimate – not an artifact of one query’s specific wording or one model’s limitations\. Unlike the first two failure modes, this one is neither resolved by scale nor explained by a fixable gap in CASCADE’s alias coverage, discussed further in Section[4\.1](https://arxiv.org/html/2608.05359#S4.SS1)\.

A targeted fix\.We designed and implemented a targeted, narrowly\-scoped fix for this failure mode \(Section[2\.5](https://arxiv.org/html/2608.05359#S2.SS5)\): a check, added tocascade\_langgraph\_mcp\_server\.py’scomprehensive\_perturbation\_analysishandler, that scans the caller’s original request text for directional cues wheneverperturbation\_typeis omitted from the tool call, returning aclarification\_neededresponse instead of silently defaulting when no single direction is detected; the check never overrides an explicitly\-supplied value, by design, so it cannot introduce a regression on any query where a model already states aperturbation\_type\.

Testing reveals a different problem\.Testing this fix against the same three ambiguous\-directional queries revealed something more fundamental than whether the fix worked: across all six ambiguous\-query calls tested \(three queries×\\timestwo models\), neither model ever leftperturbation\_typeunset – both always supplied a concrete value, correct or not \(Table[6](https://arxiv.org/html/2608.05359#S3.T6), note‡\\ddagger\) – so the check’s trigger condition, a genuinely omitted field, never actually occurred\. This reframes the original failure: it is not an omission\-and\-silent\-default problem but a confident\-wrong\-guess problem\. Instruction\-tuned models, at least the two tested here, appear to reliably populate enum\-typed schema fields even under genuine input ambiguity, rather than signaling uncertainty by leaving a field unset\. A server\-side check gated on a missing field is consequently inert against this failure mode – correctly scoped to the condition we set out to fix, but addressing a condition that, empirically, does not arise with either model tested\. Resolving the underlying failure would require a fix that engages even when a value is present – for example, validating a supplied value against the query’s actual textual support, or requiring an explicit “unspecified” sentinel value rather than allowing silent selection among enum options – which is outside the scope of what we implemented and tested here\.

## 4Discussion

The central claim this paper supports about CASCADE specifically is narrow and specific: for MYC, across three cancer types, under a PAM50 subtype control, in a fully independent patient cohort, and stable across a fourfold range of panel size, CASCADE’s downstream knockdown\-propagation predictions show real, statistically robust, biologically coherent concordance with patient tumor data\. This is a different – and, for the specific claim being validated, more direct – form of evidence than curated\-annotation enrichment: it tests whether a*specific predicted direction of change*matches what happens in actual disease, rather than whether a predicted gene is plausible in general\.

That concordance result establishes CASCADE’s predictions are trustworthy once correctly invoked; Section[3\.6](https://arxiv.org/html/2608.05359#S3.SS6)addresses a distinct, upstream claim – whether an agent calling CASCADE via MCP correctly constructs that invocation in the first place, using a locally\-hosted model with no dependency on a proprietary API\. This claim is more qualified than the biological one: it holds only partially for CASCADE’s own smaller documented default, and one identified failure mode – perturbation\-type ambiguity – persists regardless of model size \(Section[3\.6](https://arxiv.org/html/2608.05359#S3.SS6)\)\.

This result complements RegNetAgents’ validation of the converse analytical direction \(upstream regulator\-candidate identification\): where RegNetAgents tests*membership*via curated\-annotation enrichment, CASCADE’s task makes a*directional*claim that only patient\-data concordance tests directly\. The biologically coherent, patient\-data\-confirmed MYC targets in Table[3](https://arxiv.org/html/2608.05359#S3.T3)\(NOP14, RIOK1, WDR43, KAT2A, and the like\) illustrate why: they are ribosome\-biogenesis components with no OncoKB cancer\-gene annotation, yet their predicted direction of change is independently confirmed correct\. The eligibility\-passing genes that failed to validate in Section[3\.5](https://arxiv.org/html/2608.05359#S3.SS5)\(GATA3, FOXA1, SOX9, ESR1\) serve as a retrospective specificity check: each satisfies the identical structural and statistical\-power prerequisites as the validating genes, yet produces null\-to\-anti\-concordant results\.

No single mechanism yet explains why MYC, E2F3, AURKA, CCNE1, FOXM1, TOP2A, RPS6KB1, and CCND3 validate consistently wherever tested \(68–100%\); why CCND1 and MDM2 validate in BRCA but not STAD, while ERBB2 shows the reverse pattern – anti\-concordant in BRCA and COAD, concordant in STAD; why CCND3 validates in both cancer types tested but its paralog CCND2 does not, in either, despite an identical eligibility profile; or why SOX9, FOXA1, GATA3, and ESR1 fail outright\. A transcription\-factor/non\-transcription\-factor split does not explain it: none of CCND1, CCND3, AURKA, CCNE1, MDM2, TOP2A, or RPS6KB1 is a transcription factor by molecular function, yet all validate in at least one cancer type, while SOX9, FOXA1, GATA3, and ESR1 are genuine transcription factors that all fail – and CCND2, also not a transcription factor, fails too, so transcription\-factor status predicts neither outcome\.

We propose, as a hedged post\-hoc hypothesis rather than an established mechanism, a second explanation that fits the data more closely, though it remains speculative and untested by direct mechanistic evidence\. MYC, E2F3, CCND1, CCND3, AURKA, CCNE1, MDM2, FOXM1, TOP2A, and RPS6KB1 all sit within or adjacent to core proliferation\-machinery, whose output is plausibly expected to scale with regulator dosage across many tissue contexts\. GATA3, FOXA1, SOX9, and ESR1, by contrast, are lineage\-identity transcription factors that establish cell state via chromatin\-level mechanisms that may not respond linearly to added gene copies\. CCND2 does not fit either side: it shares CCND1 and CCND3’s exact molecular function, yet fails outright in every cancer type tested \(38\.0% in BRCA, 54\.2% in COAD\) where its paralogs validate, so some factor beyond this distinction must separate it from its own paralogs\.

GATA3’s result is the sharpest failure in the panel \(0 of 50 predicted targets concordant\), and not an artifact of the embedding blend or unreliable ARACNe annotation – its highest\-confidence target edges \(ESR1, MLPH, XBP1\) are correctly annotated as activating \(Section[3\.5](https://arxiv.org/html/2608.05359#S3.SS5)\)\. One candidate explanation, untested here, is that copy\-number amplification is simply the wrong dosage proxy for GATA3, whose oncogenic role in breast cancer is driven predominantly by point mutation rather than amplification\.

The cross\-cancer\-type extension \(Section[3\.5](https://arxiv.org/html/2608.05359#S3.SS5)\) adds a further wrinkle: proliferation\-machinery membership predicts validation in*at least one*cancer type well \(nine of ten non\-MYC genes validate somewhere\), but guarantees neither validation in every eligible cancer type \(CCND1, MDM2\) nor, as CCND2 shows, validation in any cancer type at all despite sharing its validating paralogs’ eligibility profile\. Candidate factors – tumor purity, stromal contamination, amplicon co\-selection – remain untested here\. Proliferation\-machinery membership is necessary but not sufficient for validation, and ERBB2’s earlier cancer\-type inconsistency \(Section[3\.5](https://arxiv.org/html/2608.05359#S3.SS5)\) turns out to be an early hint of this more general pattern\.

The baseline comparison \(Section[3\.2](https://arxiv.org/html/2608.05359#S3.SS2)\) tempers this picture: neither the MYC\-identity nor the E2F\-identity comparison distinguishes CASCADE’s accuracy from its curated baseline at conventional significance, before or after multiple\-testing correction, though CASCADE reaches this from zero gene\-specific curation, unlike the curated lists’ dedicated perturbation datasets and formal Hallmark curation procedure\(Liberzon et al\.,[2015](https://arxiv.org/html/2608.05359#bib.bib13)\)\. CASCADE’s gene\-specific direction\-calling is not redundant with a naive guess, however, as the ablation reported above shows\.111This pattern is not unique to CASCADE: GEARS’ own benchmarking includes a baseline adapted from CellOracle\(Kamimoto et al\.,[2023](https://arxiv.org/html/2608.05359#bib.bib19)\)that infers a gene regulatory network and linearly propagates perturbation signal along it – structurally similar to CASCADE’s own network\-propagation approach – and a recent broader benchmark reports that a naive average\-perturbation\-effect baseline \(the “perturbed mean”\) often matches or exceeds GEARS, scGPT, and CPA\(Viñas Torné et al\.,[2025](https://arxiv.org/html/2608.05359#bib.bib20)\)\. A simple baseline matching or exceeding a more sophisticated model recurs across the perturbation\-prediction literature generally, and is not an isolated weakness of the comparison in Section[3\.2](https://arxiv.org/html/2608.05359#S3.SS2)\.

### 4\.1Limitations

The TCGA ARACNe networks CASCADE uses are population\-averaged per cancer type and do not capture tumor subclone heterogeneity or single\-patient regulatory state\. Copy\-number amplification is a proxy for gene dosage, not a direct experimental perturbation, and the copy\-number/activity relationship is not perfectly linear for every gene – partially, not fully, mitigated by the PAM50 control and METABRIC replication\.

Our generalization panel \(fifteen genes, twenty\-seven gene\-cancer\-type combinations\) is smaller than RegNetAgents’ twelve\-gene COAD panel but larger than its eleven\-gene BRCA panel; RegNetAgents did not test STAD\. Panel size in this generalization sweep was constrained by the rarity of genes with MYC\-comparable cross\-tissue amplification\. We attempted a direct experimental test of the lineage\-identity hypothesis using LINCS L1000 shRNA data \(GSE106127, MCF7\) for GATA3, FOXA1, SOX9, and ESR1, but target coverage was insufficient: only 2–5 of each gene’s top\-50 predicted targets are on the 978\-gene L1000 panel, and SOX9 was absent from the shRNA reagent library entirely \(scripts/experiment5\_lincs\_coverage\_check\.py\)\.

The baseline comparison \(Section[3\.2](https://arxiv.org/html/2608.05359#S3.SS2)\) is similarly limited – two curated MSigDB gene lists \(MYC\_TARGETS\_V1,E2F\_TARGETS; Appendix[A](https://arxiv.org/html/2608.05359#A1)\), the latter tested only in BRCA – too few to characterize CASCADE’s accuracy against curated annotation more broadly\. A head\-to\-head comparison against CellOracle\(Kamimoto et al\.,[2023](https://arxiv.org/html/2608.05359#bib.bib19)\), the closest existing network\-propagation approach, remains open future work\. Because perturbation prediction is CASCADE’s central function rather than a secondary capability, this gap matters more here than it would for a narrower claim; the patient\-data concordance results above test that central function against real outcomes directly, but they do not establish CASCADE’s accuracy relative to the nearest comparable tool\.

A related concern is that the sequence of gene selections in this paper could reflect post\-hoc pattern\-fitting rather than a genuine effect: some later focal genes were chosen after a working hypothesis predicted their outcome \(Section[3\.5](https://arxiv.org/html/2608.05359#S3.SS5)\)\. No eligible gene was dropped along the way, and each gene’s selection provenance – blind screen, hypothesis\-motivated addition, or hypothesis\-falsifying test case – is disclosed at the point it is introduced\. This does not establish the hypothesis is correct, and the sequence includes a gene \(ESR1\) selected specifically to falsify the pattern, not only genes selected to confirm it\.

The agent tool\-call grounding benchmark \(Section[3\.6](https://arxiv.org/html/2608.05359#S3.SS6)\) is similarly bounded: 35 queries against two local models characterizes distinct failure modes but not a precise, generalizable accuracy estimate, and we did not test proprietary frontier models, which plausibly perform better on schema\-adherence\. CASCADE’s gene\-alias table \(tools/gene\_id\_mapper\.py\) covers a small, fixed set of common informal gene names; the server\-resolved figures \(77\.1%, 94\.3%\) should not be read as evidence of generalization to aliases beyond that coverage\. The ambiguous\-query perturbation\-type fix \(Section[3\.6](https://arxiv.org/html/2608.05359#S3.SS6)\) remains open: it is correctly scoped but inert against the confident\-wrong\-guess failure actually observed, and a fix that validates a supplied value against the query’s textual support remains future work\. This benchmark also evaluates a single agentic decision point in isolation, not sustained multi\-turn behavior or recovery from a failed tool call – a fuller behavioral evaluation is future work\.

## 5Acknowledgements

CASCADE’s embedding\-enhanced propagation path, described in Section[2\.1](https://arxiv.org/html/2608.05359#S2.SS1)and used throughout this paper’s experiments, uses pre\-trained gene embeddings from the GREmLN model developed by the Chan Zuckerberg Initiative AI team\(Zhang et al\.,[2025](https://arxiv.org/html/2608.05359#bib.bib14)\)\. We acknowledge the TCGA Research Network, the ARACNe network compilation\(Lim & Califano,[2018](https://arxiv.org/html/2608.05359#bib.bib4)\), the cBioPortal team, the METABRIC consortium, and the OncoKB team for the public data and APIs this work depends on\.

## 6Funding

No funding was received for this work\.

## 7Competing Interests

The author declares no competing interests\.

## 8AI Usage Disclosure

Development of the validation scripts described in this paper, and drafting of this manuscript, were assisted by Claude Code \(Anthropic\), an AI coding tool\. All experimental design decisions, statistical framework choices, interpretation of results, and the decision of which findings to report \(including negative and inconsistent results\) are the author’s own\. All code was reviewed and results independently inspected by the author before inclusion\.

## 9Data and Code Availability

CASCADE is available at[https://github\.com/jab57/CASCADE](https://github.com/jab57/CASCADE)under the MIT license, archived at Zenodo \(DOI:[https://doi\.org/10\.5281/zenodo\.21774774](https://doi.org/10.5281/zenodo.21774774), v1\.3\.0\)\. The repository README documents CASCADE’s full tool set and usage examples beyond the propagation\-and\-embedding component validated in this paper\. Validation scripts for this paper and their full results are included in the repository underscripts/andoutputs/respectively:

- •scripts/experiment4\_tcga\_myc\_concordance\.py– primary patient\-data concordance test, parametrized by cancer type and focal gene
- •scripts/experiment4\_pam50\_control\.py– PAM50 subtype\-controlled subgroup test
- •scripts/experiment4\_metabric\_replication\.py– independent\-cohort replication
- •scripts/experiment4\_hallmark\_baseline\.py– curated, identity\-matched gene\-set baseline comparison \(Section[3\.2](https://arxiv.org/html/2608.05359#S3.SS2)\)
- •scripts/experiment4\_baseline\_comparison\_stats\.py– combines CASCADE’s and each baseline’s concordance counts into the Fisher’s\-exact and BH\-FDR statistics reported in Table[8](https://arxiv.org/html/2608.05359#A1.T8)\(Section[3\.2](https://arxiv.org/html/2608.05359#S3.SS2)\)
- •scripts/experiment4\_direction\_split\_check\.py– checks whether CASCADE’s own predicted\-target panels are themselves direction\-skewed \(Section[3\.2](https://arxiv.org/html/2608.05359#S3.SS2)\)
- •scripts/experiment5\_lincs\_coverage\_check\.py– checks LINCS L1000 shRNA target\-coverage feasibility for a direct experimental test of the lineage\-identity hypothesis \(Section[4\.1](https://arxiv.org/html/2608.05359#S4.SS1)\)
- •scripts/experiment6\_agent\_tool\_grounding\.py– agent tool\-call grounding benchmark against CASCADE’s real MCP tool schema \(Section[2\.5](https://arxiv.org/html/2608.05359#S2.SS5)\)

Patient expression and copy\-number data were accessed via the public cBioPortal REST API \([https://www\.cbioportal\.org/api](https://www.cbioportal.org/api)\); OncoKB annotations via the public OncoKB API \([https://www\.oncokb\.org](https://www.oncokb.org/)\)\.

## Appendix AFull Baseline Comparison Results

Section[3\.2](https://arxiv.org/html/2608.05359#S3.SS2)describes this comparison’s methodology and interpretation; this appendix reports the full numeric results\. The two MSigDB gene sets used areMYC\_TARGETS\_V1\(tested against all three cancer types\) andE2F\_TARGETS\(tested in BRCA only, against E2F3\)\(Liberzon et al\.,[2015](https://arxiv.org/html/2608.05359#bib.bib13)\), each compared against CASCADE’s ownN=50N\{=\}50panel using every resolvable gene in the set \(188–197 genes depending on set and cancer type\)\. Table[8](https://arxiv.org/html/2608.05359#A1.T8)reports all four resulting comparisons; the comparison is not exhaustive – only two curated gene sets were tested, and E2F\-identity only in BRCA – so it should be read as a partial check, not a complete characterization\.

Table 8:Baseline comparison: CASCADE’s predictions vs\. curated, identity\-matched non\-CASCADE gene\-set baselines\. Fisher’s exactpp\(two\-sided\) tests whether CASCADE’s concordance count and the baseline’s concordance count, each out of their own respective sample size \(CASCADEN=50N\{=\}50; baselineN=188N\{=\}188–197197\), differ significantly\. BH\-FDRqqapplies Benjamini\-Hochberg correction across all four comparisons in this table\.None of the four comparisons is statistically significant, before or after Benjamini\-Hochberg correction\.

CASCADE’s own predicted\-target panels are themselves not close to a uniform “down” guess: full\-panel down/up splits range from a near\-even 47%/53% \(MYC/STAD\) to 78%/22% \(E2F3/BRCA\); the MYC/BRCA panel underlying Table[3](https://arxiv.org/html/2608.05359#S3.T3)itself is 74\.0%/26\.0% \(37/50 down\), markedly less skewed than the 86\.7% “down” share visible in that table’s top\-15\-by\-magnitude slice\. Forcing every predicted “up” gene to “down” – removing CASCADE’s direction information while keeping its gene selection unchanged – drops concordance by 14–29 percentage points in every panel \(MYC/BRCA: 90\.0%→\\to68\.0%; MYC/COAD: 72\.0%→\\to54\.0%; MYC/STAD: 85\.7%→\\to57\.1%; E2F3/BRCA: 96\.0%→\\to82\.0%\), Considering only the subset of genes CASCADE predicts will go up \(rather than down\) upon knockdown, those up\-predictions are themselves correct 76\.5–92\.3% of the time across panels – higher than the down\-predictions in the same panel for MYC/BRCA \(92\.3% vs\. 89\.2%\) and MYC/COAD \(76\.5% vs\. 69\.7%\), lower for MYC/STAD \(76\.9% vs\. 95\.7%\) and E2F3/BRCA \(81\.8% vs\. 100\.0%\), with no consistent direction across panels but never close to the 50% chance level\. This rules out the possibility that CASCADE’s minority “up” calls are just noise being carried by the accuracy of its more common “down” calls\.

## References

- Bird \[2026\]Bird, J\.A\. \(2026\)\. RegNetAgents: A Multi\-Agent Framework for Cross\-Network Regulatory Driver Identification in Cancer Genomics\.*arXiv:2607\.14097*\.
- Califano & Alvarez \[2017\]Califano, A\., Alvarez, M\.J\. \(2017\)\. The recurrent architecture of tumour initiation, progression and drug sensitivity\.*Nature Reviews Cancer*, 17, 116–130\.
- LangChain AI \[2024\]LangChain AI \(2024\)\. LangGraph Documentation\.[https://langchain\-ai\.github\.io/langgraph](https://langchain-ai.github.io/langgraph)
- Lim & Califano \[2018\]Lim, W\.K\., Califano, A\. \(2018\)\. aracne\.networks: ARACNe\-inferred gene networks from TCGA tumor datasets\.*Bioconductor*, v1\.36\.0\.
- Lachmann et al\. \[2016\]Lachmann, A\., et al\. \(2016\)\. ARACNe\-AP: gene network reverse engineering through adaptive partitioning inference of mutual information\.*Bioinformatics*, 32, 2233–2235\.
- Hoadley et al\. \[2018\]Hoadley, K\.A\., et al\. \(2018\)\. Cell\-of\-Origin Patterns Dominate the Molecular Classification of 10,000 Tumors from 33 Types of Cancer\.*Cell*, 173\(2\), 291–304\.
- Mermel et al\. \[2011\]Mermel, C\.H\., Schumacher, S\.E\., Hill, B\., Meyerson, M\., Beroukhim, R\., Getz, G\. \(2011\)\. GISTIC2\.0 Facilitates Sensitive and Confident Localization of the Targets of Focal Somatic Copy\-Number Alteration in Human Cancers\.*Genome Biology*, 12\(4\), R41\.
- Cerami et al\. \[2012\]Cerami, E\., et al\. \(2012\)\. The cBio Cancer Genomics Portal: An Open Platform for Exploring Multidimensional Cancer Genomics Data\.*Cancer Discovery*, 2, 401–404\.
- Gao et al\. \[2013\]Gao, J\., et al\. \(2013\)\. Integrative analysis of complex cancer genomics and clinical profiles using the cBioPortal\.*Science Signaling*, 6\(269\)\.
- Curtis et al\. \[2012\]Curtis, C\., et al\. \(2012\)\. The genomic and transcriptomic architecture of 2,000 breast tumours reveals novel subgroups\.*Nature*, 486, 346–352\.
- Pereira et al\. \[2016\]Pereira, B\., et al\. \(2016\)\. The somatic mutation profiles of 2,433 breast cancers refine their genomic and transcriptomic landscapes\.*Nature Communications*, 7, 11479\.
- Chakravarty et al\. \[2017\]Chakravarty, D\., et al\. \(2017\)\. OncoKB: A precision oncology knowledge base\.*JCO Precision Oncology*, 1, 1–16\.
- Liberzon et al\. \[2015\]Liberzon, A\., Birger, C\., Thorvaldsdóttir, H\., Ghandi, M\., Mesirov, J\.P\., Tamayo, P\. \(2015\)\. The Molecular Signatures Database \(MSigDB\) hallmark gene set collection\.*Cell Systems*, 1\(6\), 417–425\.
- Zhang et al\. \[2025\]Zhang, M\., Swamy, V\., Cassius, R\., Dupire, L\., Kanatsoulis, C\., Paull, E\., AlQuraishi, M\., Karaletsos, T\., Califano, A\. \(2025\)\. GREmLN: A Cellular Graph Structure Aware Transcriptomics Foundation Model\.*bioRxiv*\.[https://doi\.org/10\.1101/2025\.07\.03\.663009](https://doi.org/10.1101/2025.07.03.663009)
- Alvarez et al\. \[2016\]Alvarez, M\.J\., Shen, Y\., Giorgi, F\.M\., Lachmann, A\., Ding, B\.B\., Ye, B\.H\., Califano, A\. \(2016\)\. Functional characterization of somatic mutations in cancer using network\-based inference of protein activity\.*Nature Genetics*, 48\(8\), 838–847\.
- van Riggelen et al\. \[2010\]van Riggelen, J\., Yetil, A\., Felsher, D\.W\. \(2010\)\. MYC as a regulator of ribosome biogenesis and protein synthesis\.*Nature Reviews Cancer*, 10\(4\), 301–309\.
- Roohani et al\. \[2023\]Roohani, Y\., Huang, K\., Leskovec, J\. \(2023\)\. Predicting transcriptional outcomes of novel multigene perturbations with GEARS\.*Nature Biotechnology*\.[https://doi\.org/10\.1038/s41587\-023\-01905\-6](https://doi.org/10.1038/s41587-023-01905-6)
- Lotfollahi et al\. \[2023\]Lotfollahi, M\., Klimovskaia Susmelj, A\., De Donno, C\., Hetzel, L\., Ji, Y\., Ibarra, I\.L\., Srivatsan, S\.R\., Naghipourfar, M\., Daza, R\.M\., Martin, B\., Shendure, J\., McFaline\-Figueroa, J\.L\., Boyeau, P\., Wolf, F\.A\., Yakubova, N\., Günnemann, S\., Trapnell, C\., Lopez\-Paz, D\., Theis, F\.J\. \(2023\)\. Predicting cellular responses to complex perturbations in high\-throughput screens\.*Molecular Systems Biology*, 19\(6\), e11517\.
- Kamimoto et al\. \[2023\]Kamimoto, K\., Stringa, B\., Hoffmann, C\.M\., Jindal, K\., Solnica\-Krezel, L\., Morris, S\.A\. \(2023\)\. Dissecting cell identity via network inference and in silico gene perturbation with CellOracle\.*Nature*, 614\(7949\), 742–751\.
- Viñas Torné et al\. \[2025\]Viñas Torné, R\., et al\. \(2025\)\. Systema: a framework for evaluating genetic perturbation response prediction beyond systematic variation\.*Nature Biotechnology*\.[https://doi\.org/10\.1038/s41587\-025\-02777\-8](https://doi.org/10.1038/s41587-025-02777-8)
- Patil et al\. \[2023\]Patil, S\.G\., Zhang, T\., Wang, X\., Gonzalez, J\.E\. \(2023\)\. Gorilla: Large Language Model Connected with Massive APIs\.*arXiv:2305\.15334*\.
- Patil et al\. \[2025\]Patil, S\.G\., Mao, H\., Yan, F\., Ji, C\.C\., Suresh, V\., Stoica, I\., Gonzalez, J\.E\. \(2025\)\. The Berkeley Function Calling Leaderboard \(BFCL\): From Tool Use to Agentic Evaluation of Large Language Models\.*Proceedings of the 42nd International Conference on Machine Learning \(ICML\)*, PMLR 267, 48371–48392\.

Similar Articles

TRAPS: Therapeutic Response Analysis via Pathway-informed Stratification

arXiv cs.LG

This paper presents the first unified benchmark for pathway-guided therapy response modeling, evaluating three biologically informed architectures (BINN, GraphPath, PATH) across five cancer cohorts from The Cancer Genome Atlas for multi-label prediction of targeted therapy, radiation therapy, and survival outcomes.