From Analytics to Tumor Boards: An Evidence-Linked Multi-Agent Workflow for Oncology Feature Extraction

arXiv cs.AI Papers

Summary

The paper evaluates a configurable multi-agent system (nMAS) for extracting structured oncology data from fragmented clinical documents, achieving high performance compared to a baseline model.

arXiv:2608.28974v1 Announce Type: new Abstract: Clinically relevant oncology information is distributed across heterogeneous, longitudinal documentation, creating substantial abstraction burden and requiring accurate attribution across specimens, tumors, biomarkers, and time points, while manual cancer-registry abstraction can require 27.2 minutes per case, highlighting the need for scalable methods that preserve clinical context while converting documentation into structured data. We evaluate the Nimblemind Multi-Agent System (nMAS), a configurable oncology information-extraction workflow which extracts clinically relevant structured fields from fragmented oncology documentation. The extraction task uses a clinician-informed schema of 328 attributes spanning report metadata, diagnosis, staging, and cancer-type-specific information. nMAS separates clinician-defined field specifications from model execution and combines complexity-aware extraction, report-level consolidation, and source-grounded validation. The retrospective evaluation included 230 de-identified oncology documents from 40 patients and 418 clinician-reviewed document-field pairs containing 1,126 non-empty reference values. Evaluation focused on fields identified by clinicians as present in the source documents rather than exhaustively annotating all 328 schema fields. nMAS achieved a rank-weighted value-level precision of 82.6%, recall of 87.5%, and F1 of 85.0%, compared with an F1 of 66.4% for an independently implemented UMA-style MiniMax M2.5 comparator. These findings support the feasibility of using a configurable, source-grounded extraction workflow to convert fragmented oncology documentation into reusable structured data.
Original Article
View Cached Full Text

Cached at: 09/01/26, 12:45 PM

# From Analytics to Tumor Boards: An Evidence-Linked Multi-Agent Workflow for Oncology Feature Extraction
Source: [https://arxiv.org/html/2608.28974](https://arxiv.org/html/2608.28974)
Daniel KangAffiliation:Nimblemind and Nimblemind and Nimblemind and Nimblemind and Nimblemind and Independent Researcher and School of Interactive Computing, Georgia Institute of Technology and Department of Technology Management and Innovation, NYU Tandon School of Engineering, New York University and Florida International University and Duke\-NUS Medical School and Department of Anatomical Pathology, Singapore General Hospital and University of Illinois Urbana\-Champaign and OncoLens and OncoLens and Nimblemind and NimblemindSoorya Ram ShimgekarAffiliation:Shayan VassefAffiliation:Yufan WangAffiliation:Anit Kumar SahuAffiliation:Munmun De ChoudhuryAffiliation:Vedant Das SwainAffiliation:Christian PoellabauerAffiliation:Li Yan KhorAffiliation:Koustuv SahaAffiliation:Robert WojciechowskiAffiliation:Elliot KiddAffiliation:Piyum ZonoozAffiliation:Navin KumarAffiliation:

###### Abstract

Clinically relevant oncology information is distributed across heterogeneous, longitudinal documentation, creating substantial abstraction burden and requiring accurate attribution across specimens, tumors, biomarkers, and time points, while manual cancer\-registry abstraction can require 27\.2 minutes per case, highlighting the need for scalable methods that preserve clinical context while converting documentation into structured data\. We evaluate the Nimblemind Multi\-Agent System \(nMAS\), a configurable oncology information\-extraction workflow which extracts clinically relevant structured fields from fragmented oncology documentation\. The extraction task uses a clinician\-informed schema of 328 attributes spanning report metadata, diagnosis, staging, and cancer\-type\-specific information\. nMAS separates clinician\-defined field specifications from model execution and combines complexity\-aware extraction, report\-level consolidation, and source\-grounded validation\. The retrospective evaluation included 230 de\-identified oncology documents from 40 patients and 418 clinician\-reviewed document\-field pairs containing 1,126 non\-empty reference values\. Evaluation focused on fields identified by clinicians as present in the source documents rather than exhaustively annotating all 328 schema fields\. nMAS achieved a rank\-weighted value\-level precision of 82\.6%, recall of 87\.5%, and F1 of 85\.0%, compared with an F1 of 66\.4% for an independently implemented UMA\-style MiniMax M2\.5 comparator\. These findings support the feasibility of using a configurable, source\-grounded extraction workflow to convert fragmented oncology documentation into reusable structured data\.

††proceedings:: Submitted to ML4H 2026:\\mlhtrackname††workshop:Machine Learning for Health \(ML4H\) 2026###### keywords

Multi\-Agent Systems, Medical Data Processing, Feature Identification and Enrichment, Data Extraction/Retrieval, Machine Learning Pipeline Optimization

Similar Articles

Skill-Augmented AI Agents for Medical Research Analysis: An Exploratory Multi-Model Human Evaluation in an NSCLC Transcriptomic Biomarker Task

arXiv cs.AI

This exploratory study evaluates whether augmenting AI agents with a medical research skill package improves the quality of transcriptomic research analysis outputs compared to native AI, using a multi-model human evaluation in an NSCLC biomarker task. Results show a directional but statistically non-significant improvement, highlighting the need for larger, more robust evaluations.