Foundation Models Narrow the Knowledge Gap in Ophthalmology but Struggle with Images
Highlight
• In a cross-sectional evaluation of offline FRCOphth Part 2 preparation questions, seven foundation models (FMs) showed strong performance on textual multiple-choice items; the best-performing FM (Claude 3.5 Sonnet) achieved 77.7% accuracy, comparable with expert ophthalmologists.
• Multimodal performance (questions that included images or other non-text inputs) remained substantially lower: the top multimodal FM (GPT-4o) scored 57.5%, underperforming expert clinicians and trainees.
Sign in free to continue reading
Create or use your MedXY account to unlock the complete article.
This article was created using several editorial tools, including AI, as part of the process. Human editors reviewed this content before publication.