# 反向项目反应理论:面向碎片化癌症药物响应矩阵的稀疏鲁棒排序方法

arXiv cs.LG 论文

摘要

本文将反向项目反应理论(reverse Item Response Theory)引入药物基因组学研究,将癌种视为潜在的“受试者”,将药物视为“项目”,从而在稀疏、碎片化的药物响应矩阵中实现稳健的排序恢复,并在多种缺失机制下,于 GDSC2 数据上表现优于简单的平均聚合方法。

arXiv:2610.00002v1 Announce Type: new Abstract: We introduce reverse Item Response Theory (IRT) to pharmacogenomic drug-response analysis by treating cancer types as latent "subjects" with resistance ability and drugs as "items" with evasion difficulty. Applied to 242,036 drug sensitivity measurements from the Genomics of Drug Sensitivity in Cancer (GDSC2) database, the model estimates cancer-type-level in-vitro resistance and drug-level broad activity on a shared latent scale. Validation across four missingness regimes demonstrates that reverse IRT better recovers the full-data latent ranking than simple averaging, with advantages of Delta-rho = +0.089 to +0.095 at 60% missingness under MCAR, cancer-biased, and drug-biased sparsity. Held-out prediction confirms IRT achieves the best Brier score among five evaluated methods. Bootstrap confidence intervals show 19 of 28 cancer types have stable resistant/sensitive classifications. Cross-platform PRISM replication shows 82% directional agreement but weak rank-order correlation (rho = 0.25), indicating the contribution is methodological robustness under fragmented evaluation, not a universal clinical resistance leaderboard.
查看原文
查看缓存全文

缓存时间: 2026/10/03 09:50

# Reverse Item Response Theory for Sparsity-Robust Ranking in Fragmented Cancer Drug-Response Matrices
Source: [https://arxiv.org/html/2610.00002](https://arxiv.org/html/2610.00002)
\(May 2026\)

###### Abstract

We introduce reverse Item Response Theory \(IRT\) to pharmacogenomic drug\-response analysis by treating cancer types as latent “subjects” with resistance ability and drugs as “items” with evasion difficulty\. Applied to 242,036 drug sensitivity measurements from the Genomics of Drug Sensitivity in Cancer \(GDSC2\) database, the model estimates cancer\-type\-level in\-vitro resistance and drug\-level broad activity on a shared latent scale\. Validation across four missingness regimes demonstrates that reverse IRT better recovers the full\-data latent ranking than simple averaging, with advantages ofΔ​ρ=\+0\.089\\Delta\\rho=\+0\.089to\+0\.095\+0\.095at 60% missingness under MCAR, cancer\-biased, and drug\-biased sparsity\. Held\-out prediction confirms IRT achieves the best Brier score among five evaluated methods\. Bootstrap confidence intervals show 19 of 28 cancer types have stable resistant/sensitive classifications\. Cross\-platform PRISM replication shows 82% directional agreement but weak rank\-order correlation \(ρ=0\.25\\rho=0\.25\), indicating the contribution is methodological robustness under fragmented evaluation, not a universal clinical resistance leaderboard\.

## 1Introduction

Large\-scale pharmacogenomic screening efforts, including the Genomics of Drug Sensitivity in Cancer\(GDSC; Yang et al\.,[2012](https://arxiv.org/html/2610.00002#bib.bib14); Iorio et al\.,[2016](https://arxiv.org/html/2610.00002#bib.bib7)\)and the PRISM Repurposing dataset\(Corsello et al\.,[2020](https://arxiv.org/html/2610.00002#bib.bib3)\), have generated comprehensive drug\-response matrices spanning hundreds of drugs and thousands of cancer cell lines\. Existing approaches include ANOVA\-based biomarker discovery\(Garnett et al\.,[2012](https://arxiv.org/html/2610.00002#bib.bib6)\), machine learning prediction\(Costello et al\.,[2014](https://arxiv.org/html/2610.00002#bib.bib4)\), deep learning\(Liu et al\.,[2020](https://arxiv.org/html/2610.00002#bib.bib10)\), and recent work inferring general principles of drug sensitivity with experimental validation\(Carli et al\.,[2025](https://arxiv.org/html/2610.00002#bib.bib2)\)\.

These approaches model each drug–cell\-line pair independently or predict sensitivity from genomic features, without jointly estimating a cancer type’s global resistance and a drug’s global activity on a common measurement scale\.

Item Response Theory\(IRT; Lord and Novick,[1968](https://arxiv.org/html/2610.00002#bib.bib11); Baker and Kim,[2004](https://arxiv.org/html/2610.00002#bib.bib1)\)provides this capability\. IRT jointly estimates subject ability and item difficulty on a shared latent scale\.Kang \([2026a](https://arxiv.org/html/2610.00002#bib.bib8)\)demonstrated that IRT\-based ranking outperforms simple averaging under sparse evaluation in AI benchmarking\.Rodriguez et al\. \([2021](https://arxiv.org/html/2610.00002#bib.bib13)\)applied IRT to NLP evaluation, andPolo et al\. \([2024](https://arxiv.org/html/2610.00002#bib.bib12)\)used IRT for efficient LLM benchmarking\.

We propose a conceptual inversion: cancer types become “subjects” with latent resistanceθj\\theta\_\{j\}, drugs become “items” with evasion difficultybib\_\{i\}:

P​\(sensitive\)=σ​\(bi−θj\)P\(\\text\{sensitive\}\)=\\sigma\(b\_\{i\}\-\\theta\_\{j\}\)\(1\)The primary contribution is not the specific rankings—which are platform\-dependent—but demonstrating that reverse IRT provides sparsity\-robust ranking recovery in fragmented drug\-response matrices\. This is relevant because real\-world therapeutic evidence matrices are sparse: drugs, indications, and trial populations are unevenly evaluated\.

## 2Data

GDSC2\.Release 8\.5 \(October 2023\), Wellcome Sanger Institute\. Raw: 242,036 drug–cell\-line measurements \(969 cell lines, 286 drugs, 32 TCGA cancer types\)\. After removing unclassified types \(196,345 remaining\) and excluding CLL \(9 drugs, insufficient coverage\), we aggregate to a28×28628\\times 286matrix with 7,821 cells \(97\.7% coverage\)\. Binarization: sensitive if fittedLN\_IC50<3\.297\\text\{LN\\\_IC50\}<3\.297\(global median,≈27​μ\{\\approx\}27~\\muM\)\. A global threshold avoids drug\-specific normalization; drug parameters reflect apparent broad in\-vitro activity rather than absolute pharmacological potency\.

PRISM\.Secondary dose\-response dataset\(Corsello et al\.,[2020](https://arxiv.org/html/2610.00002#bib.bib3)\): 701,004 IC50 entries, 1,448 compounds, 499 cell lines\. Cell lines mapped via DepMap lineage metadata, yielding 17 overlapping cancer types\.

## 3Methods

Reverse IRT\.1PL model \(Eq\.[1](https://arxiv.org/html/2610.00002#S1.E1)\) with analytical gradients \(finite\-difference error<0\.02<0\.02\), L\-BFGS\-B optimization, Gaussian priors \(σ2=4\\sigma^\{2\}=4\)\. Parameters:28\+286=31428\+286=314for 7,821 observations\.

Sparsity test\.Remove 20–60% of cells under four regimes: MCAR, cancer\-biased \(harder cancers lose more\), drug\-biased \(weaker drugs lose more\), pathway\-block \(entire pathways removed\)\. Evaluate Spearmanρ\\rhovs\. full\-data IRT ranking \(15 seeds\)\.

Held\-out prediction\.20% cells held out, 10\-fold CV\. Brier score against cancer\-only averaging, drug\-only averaging, two\-way additive \(global\+cancer effect\+drug effect\\text\{global\}\+\\text\{cancer effect\}\+\\text\{drug effect\}\), logistic fixed effects, and reverse IRT\.

Bootstrap CIs\.200 drug\-panel \(column\) resamples for 95% CIs onθ\\theta\.

PRISM replication\.Independent reverse IRT on PRISM; Spearmanρ\\rhowith GDSC2 ranking\.

## 4Results

### 4\.1Sparsity Robustness

Reverse IRT outperforms averaging across all regimes \(Table[1](https://arxiv.org/html/2610.00002#S4.T1), Figure[1](https://arxiv.org/html/2610.00002#S4.F1)\)\. The advantage generally increases with missingness, except under pathway\-block missingness where it remains positive but modest\.

Table 1:Δ​ρ\\Delta\\rho= IRT−\-averaging for recovery of the full\-data IRT ranking under induced sparsity \(mean±\\pmSD, 15 seeds\)\.![Refer to caption](https://arxiv.org/html/2610.00002v1/x1.png)Figure 1:Ranking recovery advantage \(Δ​ρ\\Delta\\rho\) of reverse IRT over averaging under four missingness regimes\. Error bars show±1\\pm 1SD across 15 seeds\. IRT outperforms averaging in all conditions\.
### 4\.2Held\-Out Prediction

IRT achieves the best Brier score among all five evaluated methods \(Table[2](https://arxiv.org/html/2610.00002#S4.T2), Figure[2](https://arxiv.org/html/2610.00002#S4.F2)\)\.

Table 2:Mean Brier score across 10 held\-out folds \(lower = better\)\.![Refer to caption](https://arxiv.org/html/2610.00002v1/x2.png)Figure 2:Held\-out Brier score among five evaluated methods \(10\-fold CV\)\. Reverse IRT outperforms all methods including the fairer two\-way additive and logistic fixed\-effect baselines\.
### 4\.3Cancer Resistance Ranking

The ranking shows face\-valid concordance with known clinical difficulty patterns \(Figure[3](https://arxiv.org/html/2610.00002#S4.F3)\)\. Pancreatic adenocarcinoma \(PAAD\) ranks most resistant, directionally consistent with clinical difficulty \(13\.7% five\-year relative survival; SEER\)\. Hematological malignancies rank most sensitive, consistent with therapeutic advances in ALL \(∼90%\{\\sim\}90\\%childhood cure rate; NCI PDQ\)\. Nineteen of 28 cancer types have stable resistant/sensitive classifications \(CIs not crossing zero\); nine middle\-tier cancers remain uncertain\.

![Refer to caption](https://arxiv.org/html/2610.00002v1/x3.png)Figure 3:Cancer in\-vitro resistance with bootstrap 95% CIs \(200 drug\-panel resamples\)\. Red: highly resistant \(θ\>0\.5\\theta\>0\.5\); orange: moderately resistant; light blue: moderately sensitive; dark blue: highly sensitive\. CIs crossing zero indicate uncertain classification\.
### 4\.4External Replication

PRISM replication yieldsρ=0\.252\\rho=0\.252\(p=0\.33p=0\.33\) with 82% directional agreement \(14/17 cancers; Figure[4](https://arxiv.org/html/2610.00002#S4.F4)\)\. Directional agreement is strongest at the extremes\. Three middle\-tier cancers \(STAD, HNSC, NB\) show disagreement, consistent with the bootstrap uncertainty zone\.

![Refer to caption](https://arxiv.org/html/2610.00002v1/x4.png)Figure 4:GDSC2 vs PRISM resistance estimates\. Blue: directional agreement; red crosses: disagreement\. Only extremes and disagreements are labeled; full table in repository outputs\. Weak rank\-order correlation \(ρ=0\.25\\rho=0\.25\) despite 82% directional agreement\.

## 5Discussion

The primary finding is methodological: reverse IRT provides sparsity\-robust ranking recovery in drug\-response matrices\. The Evaluation Failure Scaling Law mechanism\(Kang,[2026a](https://arxiv.org/html/2610.00002#bib.bib8)\), originally demonstrated in AI benchmark evaluation, transfers to pharmacogenomic data\.

Limitations\.\(1\) Cell\-line in\-vitro resistance does not equal clinical resistance\. \(2\) Binarization at a globalLN\_IC50median discards continuous information\. \(3\) The 1PL model assumes unidimensional resistance; the LLTM\(Fischer,[1973](https://arxiv.org/html/2610.00002#bib.bib5); Kang,[2026b](https://arxiv.org/html/2610.00002#bib.bib9)\)with mutation features could decompose resistance into interpretable components\. \(4\) PRISM replication is directionally consistent but rank\-order weak, reflecting platform differences\.

## 6Conclusion

Reverse IRT provides a sparsity\-robust framework for ranking cancer types and drugs on a shared latent scale\. The method outperforms averaging under all tested missingness regimes, achieves the best held\-out calibration among five evaluated methods, and produces rankings with face\-valid clinical concordance\. The contribution is methodological: when drug\-response matrices become fragmented, reverse IRT preserves ranking structure better than averaging\. Cross\-platform replication confirms this is a ranking methodology contribution, not a universal biological discovery\.

## Data and Code Availability

## References

- Baker and Kim \[2004\]Baker, F\. B\. and Kim, S\.\-H\. \(2004\)\.*Item Response Theory*\. Marcel Dekker\.
- Carli et al\. \[2025\]Carli, F\. et al\. \(2025\)\. Learning and actioning general principles of cancer cell drug sensitivity\.*Nat\. Commun\.*, 16, 1654\.
- Corsello et al\. \[2020\]Corsello, S\. M\. et al\. \(2020\)\. Discovering the anticancer potential of non\-oncology drugs\.*Nat\. Cancer*, 1, 235–248\.
- Costello et al\. \[2014\]Costello, J\. C\. et al\. \(2014\)\. A community effort to assess and improve drug sensitivity prediction\.*Nat\. Biotechnol\.*, 32, 1202–1212\.
- Fischer \[1973\]Fischer, G\. H\. \(1973\)\. The linear logistic test model\.*Acta Psychol\.*, 37, 359–374\.
- Garnett et al\. \[2012\]Garnett, M\. J\. et al\. \(2012\)\. Systematic identification of genomic markers of drug sensitivity\.*Nature*, 483, 570–575\.
- Iorio et al\. \[2016\]Iorio, F\. et al\. \(2016\)\. A landscape of pharmacogenomic interactions in cancer\.*Cell*, 166, 740–754\.
- Kang \[2026a\]Kang, J\. M\. \(2026a\)\. The scaling law of evaluation failure\.*arXiv:2605\.11205*\.
- Kang \[2026b\]Kang, J\. M\. \(2026b\)\. Explaining benchmark difficulty: LLTM for feature\-based AI evaluation\. Preprint\.
- Liu et al\. \[2020\]Liu, Q\. et al\. \(2020\)\. DeepCDR: hybrid graph convolutional network for cancer drug response\.*Bioinformatics*, 36, i911–i918\.
- Lord and Novick \[1968\]Lord, F\. M\. and Novick, M\. R\. \(1968\)\.*Statistical Theories of Mental Test Scores*\. Addison\-Wesley\.
- Polo et al\. \[2024\]Polo, F\. M\. et al\. \(2024\)\. Efficient multi\-prompt evaluation of LLMs\.*NeurIPS 2024*\.
- Rodriguez et al\. \[2021\]Rodriguez, P\. et al\. \(2021\)\. Evaluation examples are not equally informative\.*ACL\-IJCNLP*, 4486–4503\.
- Yang et al\. \[2012\]Yang, W\. et al\. \(2012\)\. Genomics of Drug Sensitivity in Cancer\.*Nucleic Acids Res\.*, 41, D955–D961\.

相似文章

TRAPS: 基于通路信息分层的治疗反应分析

arXiv cs.LG

本文提出了首个用于通路引导的治疗反应建模的统一基准,评估了三种生物学信息驱动的架构(BINN、GraphPath、PATH),在来自癌症基因组图谱的五个癌症队列上,对靶向治疗、放射治疗和生存结局进行多标签预测。

面向AI安全的项目反应理论

arXiv cs.AI

本文将项目反应理论应用于192个语言模型在八个安全基准上的评估,识别出三个潜在因素,通过自适应测试实现97-99%的成本削减,并支持故意表现不佳检测和模型审计。

面向跨基准医学问答的临床结构化秩门控LoRA

arXiv cs.CL

本文提出BiRG-LoRA,一种用于医学问答的秩门控LoRA方法,利用临床结构化先验选择稀疏秩子集,在四个基准上达到69.31%的宏平均准确率,同时使用的参数少于混合专家方法。

基于项目反应理论的评分标准奖励

Hugging Face Daily Papers

论文提出 Rubric Response Theory(RRT),用两参数项目反应模型和通过在线 EM 更新的 Response Parameter Network(RPN)取代加法式评分标准打分。在 Medical 与 Science 基准上,RRT 取得比 GRPO 更高的宏平均评分标准得分,同时将 judge 请求量减半。