Can AI see what the surgeons see? Analyzing the accuracy of ChatGPT: publicly available multimodal AI for radiographic interpretation and treatment plan in oral surgery
DOI:
https://doi.org/10.18203/2394-6040.ijcmph20263196Keywords:
Artificial intelligence, ChatGPT, Maxillofacial surgeryAbstract
Background: With the rise of AI, patients have been actively doing their own research. With emergence of easily accessible multimodal AI, such as ChatGPT-5 that can interpret images, there is a growing need to evaluate their accuracy, reliability, and potential clinical application. Therefore, this study aimed to evaluate the accuracy of radiographic diagnostics and surgical treatment planning of ChatGPT-5 in the context of oral and maxillofacial surgery.
Methods: We conducted a cross-sectional study using ChatGPT-5’s free version and departmental OPG. Oral surgeons opinions regarding diagnosis and treatment plan were recorded, and a Cohen’s Kappa was calculated for inter-rater agreeability. Furthermore, Cohen’s Kappa was calculated amongst the Surgeons and AI’s response for inter-rater reliability analysis.
Results: The mean number of positive findings per radiograph was significantly higher for ChatGPT (3.69±1.28) compared to oral surgeons (2.8±1.18), with this difference reaching statistical significance (p=0.001). Indicating that AI consistently reported more diagnostic and treatment related observations per OPG.
Conclusions: The findings demonstrated that ChatGPT-5 identified a significantly higher number of diagnostic and treatment related findings per radiograph compared to clinicians. While this suggests high sensitivity, qualitative and agreement analysis (Cohen’s Kappa) revealed that many of these additional findings represented false positive interpretations.
References
Xu Y, Liu X, Cao X, Huang C, Liu E, Qian S, et al. Artificial intelligence: A powerful paradigm for scientific research. Innovation. 2021;2(4):100179.
Iqbal U, Tanweer A, Rahmanti AR, Greenfield D, Lee LT, Li YJ. Impact of large language model (ChatGPT) in healthcare: an umbrella review and evidence synthesis. J Biomed Sci. 2025;32(1):45.
Wang Q, Amugo I, Rajakaruna H, Irudayam MJ, Xie H, Shanker A, et al. Evaluating GPT-5 for melanoma detection using dermoscopic images. Diagnostics. 2025;15(23):3052.
Choi WC. Chang CI. ChatGPT-5 in Education: New Capabilities and Opportunities for Teaching and Learning. Preprints.org10. 2025.
Dursun D, Bilici Geçer R. Dental age estimation from panoramic radiographs: a comparison of orthodontist and ChatGPT-4 evaluations using the London Atlas, Nolla, and Haavikko Methods. Diagnostics. 2025;15(18):2389.
World Health Organization, 2023; Nature Medicine, 2023; JAMA Network Open. 2023.
Izzetti R, Nisi M, Aringhieri G, Crocetti L, Graziani F, Nardi C. Basic knowledge and new advances in panoramic radiography imaging techniques: a narrative review on what dentists and radiologists should know. Applied Sciences. 2021;11(17):7858.
Chen J, Mullins CD, Novak P, Thomas SB. Personalized strategies to activate and empower patients in health care and reduce health disparities. Health Educ Behav. 2016;43:25-34.
Shahsavar Y, Choudhury A. User intentions to use ChatGPT for self-diagnosis and health-related purposes: cross-sectional survey study. JMIR Hum Fact. 2023;10:e47564.
Mentzou A, Rogers A, Carvalho E, Daly A, Malone M, Kerasidou X. Artificial intelligence in digital self-diagnosis tools: a narrative overview of reviews. Mayo Clin Proc Digit Health. 2025;3(3):100242.
Kamerman J. AI in healthcare has moved past the hype- now the real work begins. Available from: https://www.usa.philips.com/healthcare/article/ai-in-healthcare-has-moved-past-the-hype-now-the-real-work-begins. Accessed on 13 August 2026.
Foresman G, Biro J, Tran A, MacRae K, Kazi S, Schubel L, et al. Patient perspectives on artificial intelligence in health care: focus group study for diagnostic communication and tool implementation. J Particip Med. 2025;17:e69564.
Suárez A, Arena S, Herranz Calzada A, Castillo Varón AI, Diaz-Flores García V, Freire Y. Decoding wisdom: evaluating ChatGPT's accuracy and reproducibility in analyzing orthopantomographic images for third molar assessment. Comput Struct Biotechnol J. 2025;28:141-7.
Jeong H, Han SS, Yu Y, Kim S, Jeon KJ. How well do large language model-based chatbots perform in oral and maxillofacial radiology? Dentomaxillofac Radiol. 2024;53(6):390-5.