[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-160422-en":3,"doc-seo-160422-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},160422,962084928432,"Maya Linwood","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Performance of Language Models on the Family Medicine In-Training Exam","Artificial intelligence, including ChatGPT and Bard, has become widely used in medical education, yet its role in family medicine has not been systematically assessed. This study evaluates three large language models—ChatGPT 3.5, ChatGPT 4.0, and Google Bard—using the 2022 family medicine in-training exam (ITE) composed of 193 multiple-choice questions. Model responses were scored and scaled, then benchmarked against postgraduate year 3 resident performance.","ORIGINAL ARTICLE  \nPerformance of Language Models on the Family Medicine In-Training Exam  \nRana E. Hanna, BSa ; Logan R. Smith, BAa ; RahulMhaskar, PhD b ; Karim Hanna, MDa,c  \nAUTHOR AFFILIATIONS:  \na Morsani College of Medicine, University of South Florida, Tampa, FL b Department of Medical Education, Morsani College of Medicine, University of South Florida, Tampa, FL c Department of Family Medicine, Morsani College of Medicine, University of South Florida, Tampa, FL  \nCORRESPONDING AUTHOR:  \nKarim Hanna, Morsani College of Medicine, University of South Florida, Tampa, FL, [khanna@usf.edu](khanna@usf.edu)  \nHOW TO CITE: Hanna RE, Smith LR, Mhaskar R, Hanna K. Performance of Language Models on the Family Medicine In-Training Exam. Fam Med.  \n2024;56(X):1-6 .  \ndoi: 10.22454/FamMed.2024.233738  \nPUBLISHED: 12 August 2024  \nKEYWORDS: artiﬁcial intelligence, board exams, family medicine, intraining exams  \n© Society of Teachers of Family Medicine  \nABSTRACT  \nBackground and Objectives: Artiﬁcial intelligence (AI), such as ChatGPT and Bard, has gained popularity as a tool in medical education. The use of AI in family medicine has not yet been assessed. The objective of this study is to compare the performance of three large language models (LLMs; ChatGPT 3.5, ChatGPT 4.0, and Google Bard) on the family medicine in-training exam (ITE) .  \nMethods: The 193 multiple-choice questions of the 2022 ITE, written by the American Board of Family Medicine, were inputted inChatGPT 3.5, ChatGPT 4.0, and Bard. The LLMs’ performance was then scored and scaled.  \nResults: ChatGPT 4 .0 scored 167/193 (86 . 5%) with a scaled score of 730 out of 800 . According to the Bayesian score predictor, ChatGPT 4.0 has a 100% chance of passing the family medicine board exam. ChatGPT 3.5 scored 66.3%, translating to a scaled score of 400 and an 88% chance of passing the family medicine board exam. Bard scored 64.2%, with a scaled score of 380 and an 85% chance of passing the boards. Compared to the national average of postgraduate year 3 residents, only ChatGPT 4.0 surpassed the residents’ mean of 68.4% .  \nConclusions: ChatGPT 4.0 was the only LLM that outperformed the family medicine postgraduate year 3 residents’ national averages on the 2022 ITE, providing robust explanations and demonstrating its potential use in delivering background information on common medical concepts that appear on board exams.  \nINTRODUCTION  \nArtiﬁcial intelligence (AI) has grown in popularity recently with increased application in many ﬁelds, including medicine. Large language models (LLMs) are deep learning models that aim to generate humanlike responses; LLMs are pretrained on a vast amount of information, and unlike search engines, they produce de novo responses to the inputs they receive. ChatGPT and Bard are publicly available chat-based generative AI developed by OpenAI and Google, respectively. The newest model, ChatGPT 4.0, has been shown to outperform ChatGPT 3.5 and other LLMs on most exams taken, including the bar exam, LSAT, SAT, Medical Knowledge Self-Assessment Program, and many others. 1 Interestingly, other researchers have investigated ChatGPT’s performance on ophthalmology 2 and neurosurgery 3 board review questions; however, LLM performance on family medicine board exams has not been evaluated. Given that ChatGPT 4.0 is trained using larger parameters than previous models, this LLM scored in the 90th percentile on a sample bar exam, while ChatGPT 3.5 scored in the bottom 10% . 1 ChatGPT 4.0 has limitations similar to previous models, yet fewer hallucinations. A hallucination is a term used in the AI ﬁeld to refer to a coherent yet untrue AI-generated response. 1  \nUnlike its predecessor, ChatGPT 4.0 utilizes computer vision to analyze images uploaded by users. For instance, AI also can help diagnose pathologies like diabetic retinopathy and skin lesions; 4 computer vision is useful in medicine because physical exam ﬁndings drive many diagnoses. The addition o","cbCaigRwjLMvMk8H","https://ap.wps.com/l/cbCaigRwjLMvMk8H","pdf",968767,3,1,6,"English","en",105,"# Abstract\n## Background and Objectives\n## Methods\n## Results\n## Conclusions\n# Introduction","[{\"question\":\"Which large language models were assessed for the family medicine in-training exam?\",\"answer\":\"The study compared ChatGPT 3.5, ChatGPT 4.0, and Google Bard on the 2022 family medicine in-training exam questions.\"},{\"question\":\"How were the models evaluated in the study?\",\"answer\":\"The 193 multiple-choice questions from the 2022 ITE were input into each model, and performance was scored and scaled to generate comparable results.\"},{\"question\":\"What were the main performance outcomes for ChatGPT 4.0 versus others?\",\"answer\":\"ChatGPT 4.0 achieved the highest score (167/193; 86.5%) and was the only model that surpassed the national average performance of postgraduate year 3 residents.\"}]","Performance of Language Models on the Family Medicine In-Training Exam | PDF",1788063219,15,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"performance-of-language-models-on-the-family-medicine-in-training-exam","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,51],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":20},"https://docshare.wps.com/document/research-report/",{"item":52,"name":13,"@type":44,"position":53},"https://docshare.wps.com/document/performance-of-language-models-on-the-family-medicine-in-training-exam/160422/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-09-05","2026-08-30",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Which large language models were assessed for the family medicine in-training exam?","Question",{"text":76,"@type":77},"The study compared ChatGPT 3.5, ChatGPT 4.0, and Google Bard on the 2022 family medicine in-training exam questions.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How were the models evaluated in the study?",{"text":81,"@type":77},"The 193 multiple-choice questions from the 2022 ITE were input into each model, and performance was scored and scaled to generate comparable results.",{"name":83,"@type":74,"acceptedAnswer":84},"What were the main performance outcomes for ChatGPT 4.0 versus others?",{"text":85,"@type":77},"ChatGPT 4.0 achieved the highest score (167/193; 86.5%) and was the only model that surpassed the national average performance of postgraduate year 3 residents.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":22,"doc_module":4,"doc_module_name":47,"category_name":112,"show_sort_weight":113,"slug":114},"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":47,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":47,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":47,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":47,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":47,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]