
Common AI Models Fall Short of Acceptable Sensitivity and Specificity Rates in DR
Published on August 15, 2025
With everyone currently in pursuit of AI applications that might save time, improve accuracy or both, you may be inclined to upload fundus photos to a chatbot and ask for a diagnosis. While that’s likely to give you a reasonable reply, it’s not good enough to rely upon. So says a recent study that tested the potential of four multimodal large language models (LLMs) to automate screenings for diabetic retinopathy (DR).In a retrospective study published in Ophthalmology Science, researchers from UC San Diego took ultrawide field (UWF) fundus images from patients diagnosed with pre-diabetes and diabetes and had four different LLMs (ChatGPT-4o, Claude 3.5 Sonnet, Google Gemini 1.5 Pro, Perplexity Llama 3.1 Sonar/Default) test them, using four distinct prompts. The photos were also graded for DR severity by three retina specialists using the ETDRS classification system to establish a comparison. The prompts included assessment of multiple choice disease diagnosis, binary disease classification and disease severity, with the LLMs assessed for accuracy, sensitivity and specificity in DR presence or absence as well as disease severity.
Even with this research team using fairly detailed prompts and including the ETDRS grid as an overlay, the four AI chatbots struggled to match human grading. A real-world scenario where a doctor uploads a photo without the overlay and simply asks “What is this?” is likely to yield even lower accuracy. Photo: Most JA, et al. Ophthalmol Sci. August 11, 2025. Click image to enlarge.
The AI tools evaluated a total of 309 eyes from 188 patients, with an average patient age of 58.9 years. Specialist grading found 70.2% of eyes to have DR with varying severities; the remaining 29.8% had no DR. Compared against expert human grading, diagnostic accuracy from fundus imaging alone exceeded 60% in some cases with the AI models. Claude and ChatGPT scored higher in terms of multiple choice disease identification than the other two LLMs for accuracy and sensitivity. ChatGPT and Perplexity had the highest accuracy in binary DR vs. no DR classification. Sensitivity varied, but specificity for all models was relatively high. The DR severity prompt with best overall results did not display significant difference between the models in accuracy, and all models had low sensitivity.The authors of the paper then outlined the potential clinical implementations. “In general, the observed sensitivity rates were inadequate and could lead to missed diagnoses, and suboptimal specificities could lead to overdiagnosis and resource inefficiencies.” They propose that combining AI screening with human grading “offers one path to more balanced sensitivity and specificity, but improvement to sensitivity at the levels we observed would be needed regardless.”They also elaborate on their work, explaining that poor performance and LLM errors were most often due to DR feature omissions and hallucinations rather than logical errors or poor interpretations of the grading system. Poor performance wasn’t limited to certain severities of disease but was most pronounced in the identification of severe nonproliferative DR and both proliferative DR forms. Most commonly hallucinated features in false positives were microaneurysms, hemorrhages and exudates. With disease severity, under-classification errors (likely feature omissions) or over-classification errors (like hallucinations) were mostly responsible for misses, depending on the prompt used.Succinctly summarized, the authors relay that large language models “are powerful tools which may eventually assist retinal image analysis. Currently, however, there is variability in the accuracy of image analysis, and diagnostic performance falls short of clinical standards for safe implementation in diabetic retinopathy diagnosis and grading.”Click here for the journal source.
Most JA, Walker EH, Mehta NN, et al. Can multimodal large language models diagnose diabetic retinopathy from fundus photos? A quantitative evaluation. Ophthalmol Sci. August 11, 2025. [Epub ahead of print]. This article was developed by the editorial staff in conjunction with experts in the field. In the process, AI may have been among the editorial tools used to meet the goals of human editors, who approved all content.
