[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-217083-en":3,"doc-seo-217083-105":30,"detail-sidebar-cat-0-en-105":94},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},217083,8796093062539,"8796093062539","",7,"Healthcare","Performance of the Artificial Intelligence large language models ChatGPT 3.5, Gemini (Google Bard), ChatGPT 4.0, and Gemini 2.5 flash in surgical subspecialty questions of Brazilian medical residency exams","Objective: With rapid advances in artificial intelligence and its growing influence on medical education, large language models such as ChatGPT and Gemini increasingly support clinical reasoning. Yet evidence remains limited regarding their performance on Brazilian medical residency entrance examinations and their value as educational tools for trainees. Therefore, this study evaluated AI performance on surgical subspecialty residency entrance exams across six programs using structured, single-best-answer multiple-choice items. Methods: Performance was assessed using 464 practice questions drawn from major institutions in São Paulo.","Official Publication of the Instituto Israelita de Ensino e Pesquisa Albert Einstein  \nPerformance of the Artificial Intelligence large language models ChatGPT 3.5, Gemini (Google Bard), ChatGPT 4.0, and Gemini 2.5 flash in surgical subspecialty questions of Brazilian medical residency exams  \n❚ Authors  \nMaria Clara Pimenta de Figueiredo, Victor Hugo Alves Diniz, Ana Clara de Campos Granado, Gabriel Chagas Lutfala Paulino, Gabryella Rodrigues de Oliveira  \n❚ Correspondence  \nE-mail: [pimenta8mariaclara@gmail.com](pimenta8mariaclara@gmail.com)  \n❚ DOI  \nDOI: 10.31744/einstein_journal/2026AO1436  \n❚ In Brief  \nThe application of Artificial Intelligence has been expanded to medicine and presents a promising future for medical education. Despite technological advances, it is still important to consider the role of professionals in developing the essential clinical judgments required for medical practice.  \n❚ Highlights  \n■ ChatGPT and Gemini are showing increased ability to accurately answer multiple-choice questions on medical exams.  \n■ There was no statistical significance in the rate of correct answers by ChatGPT 3.5 and Gemini 1.5. However, we observed that ChatGPT 4.0 performed significantly better, and so did Gemini 2.5 Flash, when comparing to the literature.  \n■ The question taxonomy did not appear to be a relevant factor regarding the success rate of the models.  \n❚ How to cite this article:  \nFigueiredo MC, Diniz VH, Granado AC, Paulino GC, Oliveira GR. Performance of the Artificial Intelligence large language models ChatGPT 3 .5, Gemini (Google Bard), ChatGPT 4.0, and Gemini 2.5 flash in surgical subspecialty questions of Brazilian medical residency exams. einstein (São Paulo) . 2026;24:eAO1436 .  \neinstein (São Paulo)  \nOR IG INAL ARTICLE  \nOfficial Publication of the Instituto Israelita de Ensino e Pesquisa Albert Einstein  \ne-ISSN: 2317-6385  \nHow to cite this article:  \nFigueiredo MC, Diniz VH, Granado AC, Paulino GC, Oliveira GR. Performance of the Artificial Intelligence large language models ChatGPT 3 . 5, Gemini (Google Bard), ChatGPT 4 .0, and Gemini 2.5 flash in surgical subspecialty questions of Brazilian medical residency exams. einstein (São Paulo) . 2026;24:eAO1436 .  \nAssociate Editor:  \nHelder I Nakaya  \nHospital Israelita Albert Einstein, São Paulo, SP, Brazil  \nORCID: [https://orcid.org/0000-0001-5297-9108](https://orcid.org/0000-0001-5297-9108)  \nCorresponding Author:  \nMaria Clara Pimenta de Figueiredo Rua Tessalia Vieira de Camargo, 126, Cidade Universitária Zeferino Vaz  \nZip code: 13084-971, Campinas, SP, Brazil  \nPhone: (55 21) 99494-3537  \nE-mail: [pimenta8mariaclara@gmail.com](pimenta8mariaclara@gmail.com)  \nReceived on:  \nOct 4, 2024  \nAccepted on:  \nAug 6, 2025  \nConflict of interest:  \nnone.  \nCopyright the authors  \nThis content is licensed  \nunder a Creative Commons Attribution 4.0 International License.  \nORIGINAL ARTICLE  \nPerformance of the Artificial Intelligence large language models ChatGPT 3.5, Gemini (Google Bard), ChatGPT 4.0, and Gemini 2.5 flash in surgical subspecialty questions of Brazilian medical residency exams  \nMaria Clara Pimenta de Figueiredo1, Victor Hugo Alves Diniz1, Ana Clara de Campos Granado1, Gabriel Chagas Lutfala Paulino1, Gabryella Rodrigues de Oliveira1  \n1 Universidade Estadual de Campinas, Campinas, SP, Brazil.  \nDOI: 10.31744/einstein_journal/2026AO1436  \n❚ ABSTRACT  \nObjective: Given the rapid advancement of Artificial Intelligence and its significant impact on medical education, particularly with the development of large language models, such as ChatGPT and Gemini, ChatBots have shown an increasing ability to support clinical reasoning. Regarding this, there has been a growing interest in assessing the performance of ChatBots in medical examinations. However, there are insufficient data on these tools for addressing Brazilian medical exam-related inquiries, as well as their potential as educational tools for medical students. Therefore, we aimed to eva","cbCaimuF6YQSbDGg","https://ap.wps.com/l/cbCaimuF6YQSbDGg","pdf",955462,1,6,"English","en",105,"# Abstract\n## Objective\n## Methods\n## Results\n## Conclusion\n# Introduction\n# How to cite this article","[{\"question\":\"What question does the study aim to answer about large language models?\",\"answer\":\"The study evaluates how ChatGPT and Gemini perform on surgical subspecialty residency entrance exam questions in Brazil and whether they could function as educational adjuncts.\"},{\"question\":\"Which models were tested and how were the questions structured?\",\"answer\":\"ChatGPT 3.5, Gemini (Google Bard), ChatGPT 4.0, and Gemini 2.5 Flash were tested using 464 single-correct-answer multiple-choice practice questions from six surgical residency programs in São Paulo.\"},{\"question\":\"What were the main findings regarding accuracy?\",\"answer\":\"All models showed substantial performance differences, with higher correctness for ChatGPT 4.0 (77.6%) and Gemini 2.5 Flash (81%) compared with ChatGPT 3.5 (55.4%) and Gemini (51.1%).\"},{\"question\":\"How should the tools be used in medical education, according to the conclusion?\",\"answer\":\"The findings suggest advanced large language models may support medical education when used appropriately, while not replacing the clinical decision-making abilities developed through formal surgeon training.\"}]","Performance of the Artificial Intelligence large language models ChatGPT 3.5, Gemini (Google Bard), ChatGPT 4.0, and Gemini 2.5 flash in surgical subspecialty questions of Brazilian medical residency exams | PDF",1788830426,15,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":10,"description":14,"schema_data":34,"social_meta":89,"head_meta":91,"extra_data":93,"updated_unix":28},"performance-of-the-artificial-intelligence-large-language-models-chatgpt-35-gemini-google-bard-chatgpt-40-and-gemini-25-flash-in-surgical-subspecialty-questions-of-brazilian-medical-residency-exams",{"@graph":35,"@context":88},[36,53,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/healthcare/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/performance-of-the-artificial-intelligence-large-language-models-chatgpt-35-gemini-google-bard-chatgpt-40-and-gemini-25-flash-in-surgical-subspecialty-questions-of-brazilian-medical-residency-exams/217083/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-09-08",true,{"@type":64,"interactionType":65,"userInteractionCount":4},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80,84],{"name":71,"@type":72,"acceptedAnswer":73},"What question does the study aim to answer about large language models?","Question",{"text":74,"@type":75},"The study evaluates how ChatGPT and Gemini perform on surgical subspecialty residency entrance exam questions in Brazil and whether they could function as educational adjuncts.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"Which models were tested and how were the questions structured?",{"text":79,"@type":75},"ChatGPT 3.5, Gemini (Google Bard), ChatGPT 4.0, and Gemini 2.5 Flash were tested using 464 single-correct-answer multiple-choice practice questions from six surgical residency programs in São Paulo.",{"name":81,"@type":72,"acceptedAnswer":82},"What were the main findings regarding accuracy?",{"text":83,"@type":75},"All models showed substantial performance differences, with higher correctness for ChatGPT 4.0 (77.6%) and Gemini 2.5 Flash (81%) compared with ChatGPT 3.5 (55.4%) and Gemini (51.1%).",{"name":85,"@type":72,"acceptedAnswer":86},"How should the tools be used in medical education, according to the conclusion?",{"text":87,"@type":75},"The findings suggest advanced large language models may support medical education when used appropriately, while not replacing the clinical decision-making abilities developed through formal surgeon training.","https://schema.org",{"og:url":51,"og:type":90,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":92,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":95},[96,100,104,108,113,117,120,125,130,133,137],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":105,"show_sort_weight":106,"slug":107},"Exam",70,"exam",{"id":109,"doc_module":4,"doc_module_name":45,"category_name":110,"show_sort_weight":111,"slug":112},5,"Comic",60,"comic",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":114,"show_sort_weight":115,"slug":116},"Technology",50,"technology",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":118,"slug":119},40,"healthcare",{"id":121,"doc_module":4,"doc_module_name":45,"category_name":122,"show_sort_weight":123,"slug":124},8,"Research & Report",30,"research-report",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":128,"slug":129},9,"Religion & Spirituality",20,"religion-spirituality",{"id":128,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":128,"slug":132},"World Cup","world-cup",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":134,"slug":136},10,"Lifestyle","lifestyle",{"id":138,"doc_module":4,"doc_module_name":45,"category_name":139,"show_sort_weight":109,"slug":140},19,"General","general"]