Accuracy and Reliability of Chatbot Responses to Physician Questions

JAMA Network Open, 2023

Plain-language summary

This study tested how well a chatbot (ChatGPT) answered real medical questions posed by clinicians. Thirty-three physicians across 17 specialties wrote 284 questions of varying difficulty, then graded the chatbot's answers for accuracy on a 6-point scale and completeness on a 3-point scale, also comparing GPT-3.5 with GPT-4 and consistency over time. The answers were largely accurate, with a median accuracy score of 5.5 (between almost completely and completely correct), suggesting such tools could make medical information more accessible while still carrying limitations that matter in clinical settings.

Read the full paper (DOI: 10.1001/jamanetworkopen.2023.36483).

This is a plain-language summary written for discoverability; the authoritative version is the published paper. Part of Travis Osterman's peer-reviewed publications.