An Agentic AI Framework Overcomes Fundamental Limitations of Large Language Models for Glaucoma Detection from Fundus Photography
Summary
This arXiv paper presents an agentic AI framework that integrates LLMs with specialized deep learning tools for glaucoma detection from fundus images, improving accuracy by 16-47 percentage points and reducing run-to-run variability compared to LLM-alone approaches.
View Cached Full Text
Cached at: 08/11/26, 08:03 AM
# An Agentic AI Framework Overcomes Fundamental Limitations of Large Language Models for Glaucoma Detection from Fundus Photography Source: [https://arxiv.org/abs/2608.07651](https://arxiv.org/abs/2608.07651) Authors:[Jalil Jalili](https://arxiv.org/search/cs?searchtype=author&query=Jalili,+J),[Hossein Taghizad](https://arxiv.org/search/cs?searchtype=author&query=Taghizad,+H),[Anuwat Jiravarnsirikul](https://arxiv.org/search/cs?searchtype=author&query=Jiravarnsirikul,+A),[Christopher Bowd](https://arxiv.org/search/cs?searchtype=author&query=Bowd,+C),[Akram Belghith](https://arxiv.org/search/cs?searchtype=author&query=Belghith,+A),[Raheleh Kafieh](https://arxiv.org/search/cs?searchtype=author&query=Kafieh,+R),[Christopher A\. Girkin](https://arxiv.org/search/cs?searchtype=author&query=Girkin,+C+A),[Sally L\. Baxter](https://arxiv.org/search/cs?searchtype=author&query=Baxter,+S+L),[Robert N\. Weinreb](https://arxiv.org/search/cs?searchtype=author&query=Weinreb,+R+N),[Linda M\. Zangwill](https://arxiv.org/search/cs?searchtype=author&query=Zangwill,+L+M),[Mark Christopher](https://arxiv.org/search/cs?searchtype=author&query=Christopher,+M) [View PDF](https://arxiv.org/pdf/2608.07651) > Abstract:Large language models \(LLMs\) show promise in medical image interpretation but suffer from hallucination, limited accuracy, and run\-to\-run inconsistency\. We developed and validated an agentic AI framework integrating LLMs with specialized deep learning tools for glaucoma detection from fundus photography\. The workflow had three steps: \(1\) LLM initial assessment; \(2\) function calling to invoke specialized tools for image quality \(QAModel, FundaQ\-8\), glaucoma classification \(SwinV2\-Tiny\), and optic disc/cup segmentation \(SegFormer\-B0\); and \(3\) LLM reflection integrating the initial impression with tool outputs\. Two LLMs \(Gemini 2\.5 Flash, GPT\-5\.4 mini\) were evaluated on two public datasets \(ORIGA, n=100; RIM\-ONE\-v3, n=100\) under uncropped and cropped fields of view; all images were independently graded by a masked fellowship\-trained glaucoma specialist\. The agentic workflow improved classification accuracy by 16 to 47 percentage points across all conditions, reaching within 6 points of the specialist; on RIM\-ONE\-v3 the best configurations matched the specialist accuracy of 88%\. LLM\-alone approaches failed in two ways: GPT\-5\.4 mini showed positive bias \(sensitivity 95\-100%, specificity 0\-5%\), while Gemini 2\.5 Flash varied stochastically between runs; the agentic workflow corrected both\. Cup\-to\-disc ratio error fell 15\-50% \(MAE 0\.156\-0\.228 to 0\.104\-0\.132\), and correlation with specialist grading rose from weak \(r=0\.12\-0\.39\) to moderate\-strong \(r=0\.59\-0\.84\)\. Run\-to\-run consistency rose from near\-random \(kappa as low as \-0\.01\) to near\-perfect \(kappa up to 0\.96\)\. Integrating LLMs with specialized tools addressed key limitations of LLM\-alone approaches, including over\-diagnosis and run\-to\-run variability\. Gains held for both LLMs, suggesting generalizability across backbones, and may signal a shift from monolithic models toward orchestrated multi\-agent systems in medical AI\. ## Submission history From: Jalil Jalili \[[view email](https://arxiv.org/show-email/345ac49b/2608.07651)\] **\[v1\]**Fri, 7 Aug 2026 17:33:12 UTC \(17,491 KB\)
Similar Articles
DeepLens Diagnosis Agent: Agentic Workflow Design Lets a Small Reasoning Model Compete with Frontier LLMs
The DeepLens Diagnosis Agent uses a five-stage agentic workflow with a small medical reasoning model (7B) to achieve 60.14% diagnostic accuracy on a 915-case benchmark, outperforming frontier LLMs like Claude Sonnet 4.5 and Gemini 3.1 Pro at lower cost. The workflow design alone yields a 36-point gain over the base model, demonstrating that structured process constraints are key for diagnostic reasoning.
Human-in-the-Loop Large Language Model Framework for Identification of Cutaneous Immune-Related Adverse Events
This paper presents a retrieval-augmented, multi-agent LLM framework with human-in-the-loop for detecting cutaneous immune-related adverse events from clinical notes, achieving higher accuracy, improved inter-rater agreement, and halved review time compared to manual review.
Agentic Large Language Models for Automated Structural Analysis of 3D Frame Systems
This paper proposes an agentic LLM framework for automated structural analysis of 3D frame systems from natural language inputs, achieving 90% accuracy on ten representative 3D frames through a multi-agent pipeline.
Agentic AI for operating scientific instruments for nanoscale characterization
The paper presents an agentic AI framework using large language models and the Model Context Protocol to automate atomic force microscope operation for nanoscale characterization, achieving expert-level performance with safe execution.
Designing a Robust LLM-Based Evaluation System for Agentic AI in Drug Discovery Through Human Alignment
This paper presents an LLM-as-a-Judge evaluation framework for agentic AI in drug discovery, validated through human alignment studies with expert annotators. It optimizes the judge to improve alignment with human judgment and provides insights for reusable evaluation in scientific domains.