Human-in-the-Loop Large Language Model Framework for Identification of Cutaneous Immune-Related Adverse Events

arXiv cs.CL Papers

Summary

This paper presents a retrieval-augmented, multi-agent LLM framework with human-in-the-loop for detecting cutaneous immune-related adverse events from clinical notes, achieving higher accuracy, improved inter-rater agreement, and halved review time compared to manual review.

arXiv:2607.20428v1 Announce Type: new Abstract: This study evaluated a retrieval-augmented, multi-agent large language model (LLM)-driven, human-in-the-loop framework for detecting cutaneous immune-related adverse events (cirAEs) from clinical notes. Compared with unassisted manual review, the LLM-assisted workflow improved accuracy (F1 = 0.88 vs 0.77), inter-rater agreement measured by Cohen's kappa (kappa = 0.82 vs 0.50), and reduced average review time by approximately half. This framework pilots how LLMs can be applied to identify immune-related toxicities across organ systems and, more broadly, enable accurate, scalable, and transparent adverse event data extraction.
Original Article
View Cached Full Text

Cached at: 07/24/26, 05:15 AM

# Human-in-the-Loop Large Language Model Framework for Identification of Cutaneous Immune-Related Adverse Events
Source: [https://arxiv.org/abs/2607.20428](https://arxiv.org/abs/2607.20428)
Authors:[Charles Lu](https://arxiv.org/search/cs?searchtype=author&query=Lu,+C),[Olivia Burke](https://arxiv.org/search/cs?searchtype=author&query=Burke,+O),[Debby Cheng](https://arxiv.org/search/cs?searchtype=author&query=Cheng,+D),[Adam Kashlan](https://arxiv.org/search/cs?searchtype=author&query=Kashlan,+A),[Caitlyn Duffy](https://arxiv.org/search/cs?searchtype=author&query=Duffy,+C),[Zeyun Lu](https://arxiv.org/search/cs?searchtype=author&query=Lu,+Z),[Lirit Fuksman](https://arxiv.org/search/cs?searchtype=author&query=Fuksman,+L),[Jin Ning Tian](https://arxiv.org/search/cs?searchtype=author&query=Tian,+J+N),[Andrew Sedlack](https://arxiv.org/search/cs?searchtype=author&query=Sedlack,+A),[Priya Katyal](https://arxiv.org/search/cs?searchtype=author&query=Katyal,+P),[Eudora Lee](https://arxiv.org/search/cs?searchtype=author&query=Lee,+E),[Ralina Karagenova](https://arxiv.org/search/cs?searchtype=author&query=Karagenova,+R),[Chuck Lin](https://arxiv.org/search/cs?searchtype=author&query=Lin,+C),[Kun\-Hsing Yu](https://arxiv.org/search/cs?searchtype=author&query=Yu,+K),[Nicole LeBoeuf](https://arxiv.org/search/cs?searchtype=author&query=LeBoeuf,+N),[Alexander Gusev](https://arxiv.org/search/cs?searchtype=author&query=Gusev,+A),[Yevgeniy R\. Semenov](https://arxiv.org/search/cs?searchtype=author&query=Semenov,+Y+R)

[View PDF](https://arxiv.org/pdf/2607.20428)

> Abstract:This study evaluated a retrieval\-augmented, multi\-agent large language model \(LLM\)\-driven, human\-in\-the\-loop framework for detecting cutaneous immune\-related adverse events \(cirAEs\) from clinical notes\. Compared with unassisted manual review, the LLM\-assisted workflow improved accuracy \(F1 = 0\.88 vs 0\.77\), inter\-rater agreement measured by Cohen's kappa \(kappa = 0\.82 vs 0\.50\), and reduced average review time by approximately half\. This framework pilots how LLMs can be applied to identify immune\-related toxicities across organ systems and, more broadly, enable accurate, scalable, and transparent adverse event data extraction\.

## Submission history

From: Lirit Fuksman \[[view email](https://arxiv.org/show-email/e33cc725/2607.20428)\] **\[v1\]**Sat, 9 May 2026 16:37:49 UTC \(829 KB\)

Similar Articles

Specialty-Specific Medical Language Model for Immune-Mediated Diseases

arXiv cs.CL

This paper presents a specialty-specific medical language model for extracting information from clinical narratives about immune-mediated and infectious diseases, using a BiLSTM-CNN-Char architecture trained on a curated corpus of 371 case reports, achieving an F1 score of 0.89.