Assessing the Accuracy and Reliability of AI-Generated Medical Responses: An Evaluation of the Chat-GPT Model (under review)

2023

Plain-language summary

This preprint evaluated how accurately and completely ChatGPT answered 284 medical questions written and graded by 33 physicians across 17 specialties. Answers were largely accurate (median accuracy 5.5 on a 6-point scale) and complete (median completeness 3, the top of the 3-point scale, meaning complete plus additional context), and answers that initially scored poorly often improved when the same questions were re-asked 8 to 17 days later. The authors conclude that ChatGPT generated largely accurate medical information as judged by specialist physicians, but with important limitations that require further research and validation.

Read the full paper (DOI: 10.21203/rs.3.rs-2566942/v1).

This is a plain-language summary written for discoverability; the authoritative version is the published paper. Part of Travis Osterman's peer-reviewed publications.