An Agentic AI Framework Overcomes Fundamental Limitations of Large Language Models for Glaucoma Detection from Fundus Photography
Summary
This arXiv paper presents an agentic AI framework that integrates LLMs with specialized deep learning tools for glaucoma detection from fundus images, improving accuracy by 16-47 percentage points and reducing run-to-run variability compared to LLM-alone approaches.
View Cached Full Text
Cached at: 08/11/26, 08:03 AM
# An Agentic AI Framework Overcomes Fundamental Limitations of Large Language Models for Glaucoma Detection from Fundus Photography Source: [https://arxiv.org/abs/2608.07651](https://arxiv.org/abs/2608.07651) Authors:[Jalil Jalili](https://arxiv.org/search/cs?searchtype=author&query=Jalili,+J),[Hossein Taghizad](https://arxiv.org/search/cs?searchtype=author&query=Taghizad,+H),[Anuwat Jiravarnsirikul](https://arxiv.org/search/cs?searchtype=author&query=Jiravarnsirikul,+A),[Christopher Bowd](https://arxiv.org/search/cs?searchtype=author&query=Bowd,+C),[Akram Belghith](https://arxiv.org/search/cs?searchtype=author&query=Belghith,+A),[Raheleh Kafieh](https://arxiv.org/search/cs?searchtype=author&query=Kafieh,+R),[Christopher A\. Girkin](https://arxiv.org/search/cs?searchtype=author&query=Girkin,+C+A),[Sally L\. Baxter](https://arxiv.org/search/cs?searchtype=author&query=Baxter,+S+L),[Robert N\. Weinreb](https://arxiv.org/search/cs?searchtype=author&query=Weinreb,+R+N),[Linda M\. Zangwill](https://arxiv.org/search/cs?searchtype=author&query=Zangwill,+L+M),[Mark Christopher](https://arxiv.org/search/cs?searchtype=author&query=Christopher,+M) [View PDF](https://arxiv.org/pdf/2608.07651) > Abstract:Large language models \(LLMs\) show promise in medical image interpretation but suffer from hallucination, limited accuracy, and run\-to\-run inconsistency\. We developed and validated an agentic AI framework integrating LLMs with specialized deep learning tools for glaucoma detection from fundus photography\. The workflow had three steps: \(1\) LLM initial assessment; \(2\) function calling to invoke specialized tools for image quality \(QAModel, FundaQ\-8\), glaucoma classification \(SwinV2\-Tiny\), and optic disc/cup segmentation \(SegFormer\-B0\); and \(3\) LLM reflection integrating the initial impression with tool outputs\. Two LLMs \(Gemini 2\.5 Flash, GPT\-5\.4 mini\) were evaluated on two public datasets \(ORIGA, n=100; RIM\-ONE\-v3, n=100\) under uncropped and cropped fields of view; all images were independently graded by a masked fellowship\-trained glaucoma specialist\. The agentic workflow improved classification accuracy by 16 to 47 percentage points across all conditions, reaching within 6 points of the specialist; on RIM\-ONE\-v3 the best configurations matched the specialist accuracy of 88%\. LLM\-alone approaches failed in two ways: GPT\-5\.4 mini showed positive bias \(sensitivity 95\-100%, specificity 0\-5%\), while Gemini 2\.5 Flash varied stochastically between runs; the agentic workflow corrected both\. Cup\-to\-disc ratio error fell 15\-50% \(MAE 0\.156\-0\.228 to 0\.104\-0\.132\), and correlation with specialist grading rose from weak \(r=0\.12\-0\.39\) to moderate\-strong \(r=0\.59\-0\.84\)\. Run\-to\-run consistency rose from near\-random \(kappa as low as \-0\.01\) to near\-perfect \(kappa up to 0\.96\)\. Integrating LLMs with specialized tools addressed key limitations of LLM\-alone approaches, including over\-diagnosis and run\-to\-run variability\. Gains held for both LLMs, suggesting generalizability across backbones, and may signal a shift from monolithic models toward orchestrated multi\-agent systems in medical AI\. ## Submission history From: Jalil Jalili \[[view email](https://arxiv.org/show-email/345ac49b/2608.07651)\] **\[v1\]**Fri, 7 Aug 2026 17:33:12 UTC \(17,491 KB\)
Similar Articles
DeepLens Diagnosis Agent: Agentic Workflow Design Lets a Small Reasoning Model Compete with Frontier LLMs
The DeepLens Diagnosis Agent uses a five-stage agentic workflow with a small medical reasoning model (7B) to achieve 60.14% diagnostic accuracy on a 915-case benchmark, outperforming frontier LLMs like Claude Sonnet 4.5 and Gemini 3.1 Pro at lower cost. The workflow design alone yields a 36-point gain over the base model, demonstrating that structured process constraints are key for diagnostic reasoning.
Human-in-the-Loop Large Language Model Framework for Identification of Cutaneous Immune-Related Adverse Events
This paper presents a retrieval-augmented, multi-agent LLM framework with human-in-the-loop for detecting cutaneous immune-related adverse events from clinical notes, achieving higher accuracy, improved inter-rater agreement, and halved review time compared to manual review.
Agentic Large Language Models for Automated Structural Analysis of 3D Frame Systems
This paper proposes an agentic LLM framework for automated structural analysis of 3D frame systems from natural language inputs, achieving 90% accuracy on ten representative 3D frames through a multi-agent pipeline.
DoctorAgents: an agentic framework to iteratively refine AutoML pipeline for small clinical temporal data
The paper proposes DoctorAgents, an agentic AI framework that uses specialized LLM agents to iteratively generate, validate, and refine end-to-end machine learning pipelines for small, heterogeneous clinical temporal datasets, outperforming established AutoML baselines.
Large Language Models as Unified Multimodal Learners for Clinical Prediction
The paper proposes converting multimodal patient data (text, labs, vitals) into a single natural language sequence and fine-tuning LLMs for clinical prediction, achieving comparable or better performance than specialized fusion architectures across three tasks.