2608004987
  • Open Access
  • Review

Large Language Models for Differential Diagnosis: A Survey of Performance, Collaboration, and Technical Strategies

  • Yunjia Wu †,   
  • Qi Yan †,   
  • Dingcheng Tian *

Received: 25 Jan 2026 | Revised: 17 Aug 2026 | Accepted: 21 Aug 2026 | Published: 25 Aug 2026

Abstract

Errors in differential diagnosis often arise while clinicians are generating and comparing candidate explanations. This review examines the use of large language models (LLMs) for this part of diagnostic reasoning. Internal medicine and pediatrics are the main focus; evidence from radiology, surgical subspecialties, infectious disease, and mental health is used to examine how findings change across specialties. Reported performance depends on the clinical setting, the quality of the input, the prompt, model adaptation, and the evaluation design. Some studies place LLMs near trainees and find that they produce wider, better-organized differentials. Experienced clinicians, however, remain more reliable overall. Domain adaptation, external knowledge, and interactive workflows have improved performance in specific evaluations, but hallucinations and automation bias remain, alongside unresolved questions of governance. Current evidence therefore supports clinician-supervised use of artificial intelligence (AI) systems rather than autonomous diagnosis, pending prospective and specialty-specific evaluation.

References 

  • 1.

    Jung, K.H. Large Language Models in Medicine: Clinical Applications, Technical Challenges, and Ethical Considerations. Healthc. Inform. Res. 2025, 31, 114–124. https://doi.org/10.4258/hir.2025.31.2.114.

  • 2.

    Yang, X.; Li, T.; Su, Q.; et al. Application of Large Language Models in Disease Diagnosis and Treatment. Chin. Med. J. 2025, 138, 130–142. https://doi.org/10.1097/cm9.0000000000003456.

  • 3.

    Ríos-Hoyo, A.; Shan, N.L.; Li, A.; et al. Evaluation of Large Language Models as a Diagnostic Aid for Complex Medical Cases. Front. Med. 2024, 11, 1380148. https://doi.org/10.3389/fmed.2024.1380148.

  • 4.

    McDuff, D.; Schaekermann, M.; Tu, T.; et al. Towards Accurate Differential Diagnosis with Large Language Models. Nature 2025, 642, 451–457. https://doi.org/10.1038/s41586-025-08869-4.

  • 5.

    Wang, X.; Ye, H.; Zhang, S.; et al. Evaluation of the Performance of Three Large Language Models in Clinical Decision Support: A Comparative Study Based on Actual Cases. J. Med. Syst. 2025, 49, 23. https://doi.org/10.1007/s10916-025-02152-9.

  • 6.

    Mansoor, M.A.; Ibrahim, A.F.; Grindem, D.J.; et al. Evaluating the Accuracy and Reliability of Large Language Models in Assisting with Pediatric Differential Diagnoses: A Multicenter Diagnostic Study. medRxiv 2024. https://doi.org/10.1101/2024.08.09.24311777.

  • 7.

    Akhondi-Asl, A.; Yang, Y.; Luchette, M.; et al. Comparing the Quality of Domain-Specific versus General Language Models for Artificial Intelligence-Generated Differential Diagnoses in PICU Patients. Pediatr. Crit. Care Med. 2024, 25, e273–e282. https://doi.org/10.1097/pcc.0000000000003468.

  • 8.

    Sun, S.H.; Huynh, K.; Cortes, G.; et al. Testing the Ability and Limitations of ChatGPT to Generate Differential Diagnoses from Transcribed Radiologic Findings. Radiology 2024, 313, e232346. https://doi.org/10.1148/radiol.232346.

  • 9.

    Spitzer, P.; Hendriks, D.; Rudolph, J.; et al. The Effect of Medical Explanations from Large Language Models on Diagnostic Accuracy in Radiology. npj Digit. Med. 2026, 9, 333. https://doi.org/10.1038/s41746-026-02619-0.

  • 10.

    Kim, S.H.; Wihl, J.; Schramm, S.; et al. Human-AI Collaboration in Large Language Model-Assisted Brain MRI Differential Diagnosis: A Usability Study. Eur. Radiol. 2025, 35, 5252–5263. https://doi.org/10.1007/s00330-025-11484-6.

  • 11.

    AlShenaiber, A.; Datta, S.; Mosa, A.J.; et al. Large Language Models in the Diagnosis of Hand and Peripheral Nerve Injuries: An Evaluation of ChatGPT and the Isabel Differential Diagnosis Generator. J. Hand Surg. Glob. Online 2024, 6, 847–854. https://doi.org/10.1016/j.jhsg.2024.07.011.

  • 12.

    Shanmugam, S.K.; Browning, D. Comparison of Large Language Models in Diagnosis and Management of Challenging Clinical Cases. Clin. Ophthalmol. 2024, 18, 3239–3247. https://doi.org/10.2147/opth.s488232.

  • 13.

    Mondal, A.; Karad, R.; Bhattacharjee, B.; et al. Evaluating the Performance of Artificial Intelligence in Generating Differential Diagnoses for Infectious Diseases Cases: A Comparative Study of Large Language Models. medRxiv 2024. https://doi.org/10.1101/2024.06.28.24309694.

  • 14.

    Wu, Y.; Wan, G.; Li, J.; et al. WiseMind: A Knowledge-Guided Multi-Agent Framework for Accurate and Empathetic Psychiatric Diagnosis. npj Digit. Med. 2026, 9, 575. https://doi.org/10.1038/s41746-026-02559-9.

  • 15.

    Gargari, O.K.; Fatehi, F.; Mohammadi, I.; et al. Diagnostic Accuracy of Large Language Models in Psychiatry. Asian J. Psychiatry 2024, 100, 104168. https://doi.org/10.1016/j.ajp.2024.104168.

  • 16.

    Levkovich, I. Evaluating Diagnostic Accuracy and Treatment Efficacy in Mental Health: A Comparative Analysis of Large Language Model Tools and Mental Health Professionals. Eur. J. Investig. Health Psychol. Educ. 2025, 15, 9. https://doi.org/10.3390/ejihpe15010009.

  • 17.

    Thomas, J.; Elyoseph, Z.; Kuchinke, L.; et al. Large Language Model Performance versus Human Expert Ratings in Automated Suicide Risk Assessment. Sci. Rep. 2025, 15, 39231. https://doi.org/10.1038/s41598-025-22402-7.

  • 18.

    Takita, H.; Kabata, D.; Walston, S.L.; et al. A Systematic Review and Meta-Analysis of Diagnostic Performance Comparison between Generative AI and Physicians. npj Digit. Med. 2025, 8, 175. https://doi.org/10.1038/s41746-025-01543-z.

  • 19.

    Shan, G.; Chen, X.; Wang, C.; et al. Comparing Diagnostic Accuracy of Clinical Professionals and Large Language Models: Systematic Review and Meta-Analysis. JMIR Med. Inform. 2025, 13, e64963. https://doi.org/10.2196/64963.

  • 20.

    Wu, C.K.; Chen, W.L.; Chen, H.H. Large Language Models Perform Diagnostic Reasoning. arXiv 2023, arXiv:2307.08922.

  • 21.

    Zhou, S.; Lin, M.; Ding, S.; et al. Explainable Differential Diagnosis with Dual-Inference Large Language Models. npj Health Syst. 2025, 2, 12. https://doi.org/10.1038/s44401-025-00015-6.

  • 22.

    Sayin, B.; Schlicht, I.B.; Hong, N.V.; et al. MedSyn: Enhancing Diagnostics with Human-AI Collaboration. In Proceedings of the Hybrid Human-Artificial Intelligence (HHAI) 2025 Workshops, Pisa, Italy, 9–10 June 2025; pp. 494–505.

  • 23.

    Zöller, N.; Berger, J.; Lin, I.; et al. Human–AI Collectives Most Accurately Diagnose Clinical Vignettes. Proc. Natl. Acad. Sci. USA 2025, 122, e2426153122. https://doi.org/10.1073/pnas.2426153122.

  • 24.

    Savage, T.; Nayak, A.; Gallo, R.; et al. Diagnostic Reasoning Prompts Reveal the Potential for Large Language Model Interpretability in Medicine. npj Digit. Med. 2024, 7, 20. https://doi.org/10.1038/s41746-024-01010-1.

  • 25.

    Cabral, S.; Restrepo, D.; Kanjee, Z.; et al. Clinical Reasoning of a Generative Artificial Intelligence Model Compared with Physicians. JAMA Intern. Med. 2024, 184, 581–583. https://doi.org/10.1001/jamainternmed.2024.0295.

  • 26.

    Goh, E.; Gallo, R.; Hom, J.; et al. Large Language Model Influence on Diagnostic Reasoning: A Randomized Clinical Trial. JAMA Netw. Open 2024, 7, e2440969. https://doi.org/10.1001/jamanetworkopen.2024.40969.

  • 27.

    Qazi, I.A.; Ali, A.; Khawaja, A.U.; et al. Large Language Model Diagnostic Assistance for Physicians in a Lower-Middle-Income Country: A Randomized Controlled Trial. Nat. Health 2026, 1, 198–205. https://doi.org/10.1038/s44360-025-00007-8.

  • 28.

    Reese, J.T.; Danis, D.; Caufield, J.H.; et al. On the Limitations of Large Language Models in Clinical Diagnosis. medRxiv 2024. https://doi.org/10.1101/2023.07.13.23292613.

  • 29.

    World Health Organization. Ethics and Governance of Artificial Intelligence for Health: Guidance on Large Multi-Modal Models; World Health Organization: Geneva, Switzerland, 2024.

  • 30.

    Choudhury, A.; Chaudhry, Z. Large Language Models and User Trust: Consequence of Self-Referential Learning Loop and the Deskilling of Health Care Professionals. J. Med. Internet Res. 2024, 26, e56764. https://doi.org/10.2196/56764.

  • 31.

    Zondag, A.G.M.; Rozestraten, R.; Grimmelikhuijsen, S.G.; et al. The Effect of Artificial Intelligence on Patient-Physician Trust: Cross-Sectional Vignette Study. J. Med. Internet Res. 2024, 26, e50853. https://doi.org/10.2196/50853.

  • 32.

    Mello, M.M.; Guha, N. Understanding Liability Risk from Using Health Care Artificial Intelligence Tools. N. Engl. J. Med. 2024, 390, 271–278. https://doi.org/10.1056/nejmhle2308901.

  • 33.

    Hirosawa, T.; Harada, Y.; Mizuta, K.; et al. Evaluating ChatGPT-4’s Accuracy in Identifying Final Diagnoses within Differential Diagnoses Compared with Those of Physicians: Experimental Study for Diagnostic Cases. JMIR Form. Res. 2024, 8, e59267. https://doi.org/10.2196/59267.

  • 34.

    Kämmer, J.E.; Hautz, W.E.; Krummrey, G.; et al. Effects of Interacting with a Large Language Model Compared with a Human Coach on the Clinical Diagnostic Process and Outcomes among Fourth-Year Medical Students: Study Protocol for a Prospective, Randomised Experiment Using Patient Vignettes. BMJ Open 2024, 14, e087469. https://doi.org/10.1136/bmjopen-2024-087469.

Share this article:
How to Cite
Wu, Y.; Yan, Q.; Tian, D. Large Language Models for Differential Diagnosis: A Survey of Performance, Collaboration, and Technical Strategies. AI Medicine 2026, 3 (2), 7. https://doi.org/10.53941/aim.2026.100007.
RIS
BibTex
Copyright & License
article copyright Image
Copyright (c) 2026 by the authors.