Human-in-the-Loop Large Language Model Framework for Identification of Cutaneous Immune-Related Adverse Events
Summary
This paper presents a retrieval-augmented, multi-agent LLM framework with human-in-the-loop for detecting cutaneous immune-related adverse events from clinical notes, achieving higher accuracy, improved inter-rater agreement, and halved review time compared to manual review.
View Cached Full Text
Cached at: 07/24/26, 05:15 AM
# Human-in-the-Loop Large Language Model Framework for Identification of Cutaneous Immune-Related Adverse Events Source: [https://arxiv.org/abs/2607.20428](https://arxiv.org/abs/2607.20428) Authors:[Charles Lu](https://arxiv.org/search/cs?searchtype=author&query=Lu,+C),[Olivia Burke](https://arxiv.org/search/cs?searchtype=author&query=Burke,+O),[Debby Cheng](https://arxiv.org/search/cs?searchtype=author&query=Cheng,+D),[Adam Kashlan](https://arxiv.org/search/cs?searchtype=author&query=Kashlan,+A),[Caitlyn Duffy](https://arxiv.org/search/cs?searchtype=author&query=Duffy,+C),[Zeyun Lu](https://arxiv.org/search/cs?searchtype=author&query=Lu,+Z),[Lirit Fuksman](https://arxiv.org/search/cs?searchtype=author&query=Fuksman,+L),[Jin Ning Tian](https://arxiv.org/search/cs?searchtype=author&query=Tian,+J+N),[Andrew Sedlack](https://arxiv.org/search/cs?searchtype=author&query=Sedlack,+A),[Priya Katyal](https://arxiv.org/search/cs?searchtype=author&query=Katyal,+P),[Eudora Lee](https://arxiv.org/search/cs?searchtype=author&query=Lee,+E),[Ralina Karagenova](https://arxiv.org/search/cs?searchtype=author&query=Karagenova,+R),[Chuck Lin](https://arxiv.org/search/cs?searchtype=author&query=Lin,+C),[Kun\-Hsing Yu](https://arxiv.org/search/cs?searchtype=author&query=Yu,+K),[Nicole LeBoeuf](https://arxiv.org/search/cs?searchtype=author&query=LeBoeuf,+N),[Alexander Gusev](https://arxiv.org/search/cs?searchtype=author&query=Gusev,+A),[Yevgeniy R\. Semenov](https://arxiv.org/search/cs?searchtype=author&query=Semenov,+Y+R) [View PDF](https://arxiv.org/pdf/2607.20428) > Abstract:This study evaluated a retrieval\-augmented, multi\-agent large language model \(LLM\)\-driven, human\-in\-the\-loop framework for detecting cutaneous immune\-related adverse events \(cirAEs\) from clinical notes\. Compared with unassisted manual review, the LLM\-assisted workflow improved accuracy \(F1 = 0\.88 vs 0\.77\), inter\-rater agreement measured by Cohen's kappa \(kappa = 0\.82 vs 0\.50\), and reduced average review time by approximately half\. This framework pilots how LLMs can be applied to identify immune\-related toxicities across organ systems and, more broadly, enable accurate, scalable, and transparent adverse event data extraction\. ## Submission history From: Lirit Fuksman \[[view email](https://arxiv.org/show-email/e33cc725/2607.20428)\] **\[v1\]**Sat, 9 May 2026 16:37:49 UTC \(829 KB\)
Similar Articles
A Multi-Domain Red Teaming Framework for Safety, Robustness, and Fairness Evaluation of Medical Large Language Models
This paper presents a multi-domain red teaming framework for evaluating safety, robustness, and fairness of medical LLMs across 690 clinically grounded scenarios. Results show that high aggregate accuracy can mask critical failures, and hybrid evaluation with clinician oversight is necessary for credible safety assessment.
Large Language Models as Unified Multimodal Learners for Clinical Prediction
The paper proposes converting multimodal patient data (text, labs, vitals) into a single natural language sequence and fine-tuning LLMs for clinical prediction, achieving comparable or better performance than specialized fusion architectures across three tasks.
AIPatient Arena: EHR-grounded evaluation of large language models in end-to-end clinical consultation workflows
Introduces AIPatient Arena, an EHR-grounded evaluation framework for assessing LLMs across multiple dimensions of clinical competence. The study reveals strengths in interviewing and ethics but weaknesses in handling ambiguity and diagnostic accuracy.
Language Models as Interfaces, Not Oracles: A Hybrid LLM-ML System for Pediatric Appendicitis
This paper presents ClaMPAPP, a hybrid architecture that uses an LLM as an interface to extract features from clinical narratives, which are then passed to an XGBoost classifier for pediatric appendicitis diagnosis, demonstrating improved robustness and safety over end-to-end LLM baselines.
Specialty-Specific Medical Language Model for Immune-Mediated Diseases
This paper presents a specialty-specific medical language model for extracting information from clinical narratives about immune-mediated and infectious diseases, using a BiLSTM-CNN-Char architecture trained on a curated corpus of 371 case reports, achieving an F1 score of 0.89.