Open-source Python library + no-code web dashboard for evaluating oncology AI models at clinical decision thresholds. [P]
Summary
This article introduces oncothresh, a Python library and web dashboard for evaluating oncology AI models at specific clinical decision thresholds, providing metrics like sensitivity, specificity, and decision-curve analysis to fill gaps in existing benchmarks.
Similar Articles
Expert Evaluation of Clinical AI Tools on Real Point-of-Care Clinical Queries
This paper presents a blinded evaluation of clinical AI tools using real point-of-care queries from physicians, comparing specialized and general-purpose models across five dimensions. The specialized tool (OpenEvidence) outperformed general-purpose models on all axes, and the authors release the Real-POCQi benchmark.
OncoTriad-QA: A Patient-Level Radiology-Pathology-Genomics Benchmark for Pan-Cancer Reasoning
Introduces OncoTriad-QA, a patient-level benchmark integrating radiology, pathology, genomics, and clinical data for pan-cancer reasoning, along with OncoVLM, a reference multimodal model that outperforms existing medical LLMs after fine-tuning.
Treatment Response Optimized Clinical Decision Support AI System via Digital Twin Simulation
This paper presents an online adaptive clinical decision support AI system that integrates treatment effect estimation, digital twin simulation, and reinforcement learning to recommend treatments in a safe, clinician-supervised manner, validated on a synthetic simulator and the TCGA ovarian cancer dataset.
"OncoAgent: A Dual-Tier Multi-Agent Framework for Privacy-Preserving Oncology Clinical Decision Support"
The article introduces OncoAgent, a dual-tier multi-agent framework designed for privacy-preserving clinical decision support in oncology. It details a system architecture that combines corrective RAG, a reflexion safety loop, and dual-tier QLoRA fine-tuning optimized for AMD hardware.
Introducing HealthBench
OpenAI introduces HealthBench, a new benchmark for evaluating AI systems in healthcare contexts, created with 262 physicians across 60 countries. The benchmark includes 5,000 realistic health conversations with physician-written rubrics to assess model performance on meaningful, trustworthy, and improvable metrics.