[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-212232-en":3,"doc-seo-212232-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},212232,5909877438554,"Maeve","https://ap-avatar.wpscdn.com/avatar/5600025385ad2bf12a7?_k=1778553567797529272",8,"Research & Report","Class of LLMs - Benchmarking Large Language Models on the Brazilian National Medical Examination","The evaluation of Large Language Models (LLMs) in medicine has largely depended on English benchmarks tied to North American guidelines, reducing transferability to other healthcare systems. This paper benchmarks twenty-two proprietary and open-weight LLMs on ENAMED 2025, Brazil’s high-stakes National Examination for the Evaluation of Medical Training. The dataset contains 90 multiple-choice questions grounded in Brazilian public health policy, clinical practice, and Portuguese terminology, assessed via accuracy and ENAMED’s official IRT framework.","Class of LLMs: Benchmarking Large Language Models on the Brazilian  \nNational Medical Examination  \nJoão Vitor Mariano Correia 1 , Pedro Henrique Alves de Castro2 , Gabriel Lino Garcia 1 ,  \nPedro Henrique Paiola 1 , João Paulo Papa 1  \n1Department of Computing, Faculty of Sciences, São Paulo State University  \n2Department of Medical Sciences, Nove de Julho University  \nAbstract  \nThe evaluation of Large Language Models (LLMs) in medicine has predominantly relied on English-language benchmarks aligned with North American clinical guidelines, limiting their applicability to other healthcare systems.  \nIn this paper, we evaluate twenty-two proprietary and open-weight LLMs on the 2025 National Examination for the Evaluation of Medical Training (ENAMED), a high-stakes, government-standardized assessment used to evaluate medical graduates in Brazil. The benchmark comprises 90 multiple-choice questions grounded in Brazilian public health policy, clinical practice, and Portuguese medical terminology, and is released as an open dataset.  \nModel performance is measured using both standard accuracy and the official Item Response Theory (IRT) framework employed by ENAMED, enabling direct comparison with human proficiency thresholds. Results reveal a clear stratification of model capabilities: proprietary frontier models achieve the highest performance, whereas many open-weight and smallerdomain-adapted models fail to meet the minimum proficiency criterion. Across comparable scales, large generalist models consistently outperform specialized medical fine-tunes, suggesting that general reasoning capacity is a stronger predictor of success than narrow domain adaptation in this setting. These findings establish ENAMED as a rigorous benchmark for evaluating medical LLMs in Portuguese and highlight both the potential and current limitations of such models for educational assessment.  \n1 Introduction  \nThe integration of Artificial Intelligence into clinical practice has advanced Large Language Models (LLMs) from experimental systems to evaluated tools for tasks like summarization and question answering. While frontier models now pass global, English-language benchmarks such as the USMLE  \n(Nori et al., 2023 ; Singhal et al., 2023), these assessments ignore the organizational structure of the Brazilian Unified Health System (SUS), regional epidemiology, and the Federal Council of Medicine (CFM) professional norms.  \nIn 2025, Brazil introduced the National Examination for the Evaluation of Medical Training (ENAMED), a centralized assessment consolidating undergraduate and residency evaluations. Aligned with National Curricular Guidelines (DCNs), ENAMED uniquely prioritizes public health policies and primary care. Its inaugural administration yielded unsatisfactory institutional scores, providing a rigorous, governmentstandardized context for evaluating AI preparedness and clinical reasoning in Brazilian medical education.  \nThis work evaluates the performance of generalpurpose and domain-adapted Large Language Models on ENAMED. Our results show that recent generalist models consistently outperform several specialized medical models. The main contributions of this study are twofold: (i) a systematic, domainlevel analysis of LLM behavior in a national medical assessment setting, highlighting both their potential and the challenges of deploying such models in context-specific healthcare environments, and (ii) the release of a structured dataset derived from ENAMED 2025, enabling reproducible evaluation and future research on LLM performance in Brazilian medical education.  \n2 Related Work  \nEarly evaluations of medical LLMs relied primarily on English-language benchmarks, including general-purpose and biomedical questionanswering datasets such as MMLU (Hendryckset al., 2021), PubMedQA (Jin et al., 2019), and MedQA (Jin et al., 2021) . Although these datasets effectively assess biomedical knowledge and multi-  \n101  \nProceedings of the 17th Internat","cbCailp6rkpLCQjj","https://ap.wps.com/l/cbCailp6rkpLCQjj","pdf",270778,1,11,"English","en",105,"# Abstract\n# 1 Introduction\n# 2 Related Work\n# 3 Methodology","[{\"question\":\"What benchmark does the paper use to evaluate medical LLMs in Brazil?\",\"answer\":\"The study evaluates 22 proprietary and open-weight LLMs on ENAMED 2025, Brazil’s National Examination for the Evaluation of Medical Training.\"},{\"question\":\"How is model performance measured in the study?\",\"answer\":\"Performance is measured using standard accuracy and the official Item Response Theory (IRT) framework used by ENAMED, allowing comparison to human proficiency thresholds.\"},{\"question\":\"What does the paper find about different types of LLMs on ENAMED?\",\"answer\":\"Proprietary frontier models perform best, while many open-weight and smaller domain-adapted models fail to reach the minimum proficiency criterion; generalist models outperform specialized medical fine-tunes at comparable scales.\"}]","Class of LLMs - Benchmarking Large Language Models on the Brazilian National Medical Examination | PDF",1788721006,28,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"class-of-llms-benchmarking-large-language-models-on-the-brazilian-national-medical-examination","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/class-of-llms-benchmarking-large-language-models-on-the-brazilian-national-medical-examination/212232/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-09-06",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What benchmark does the paper use to evaluate medical LLMs in Brazil?","Question",{"text":75,"@type":76},"The study evaluates 22 proprietary and open-weight LLMs on ENAMED 2025, Brazil’s National Examination for the Evaluation of Medical Training.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is model performance measured in the study?",{"text":80,"@type":76},"Performance is measured using standard accuracy and the official Item Response Theory (IRT) framework used by ENAMED, allowing comparison to human proficiency thresholds.",{"name":82,"@type":73,"acceptedAnswer":83},"What does the paper find about different types of LLMs on ENAMED?",{"text":84,"@type":76},"Proprietary frontier models perform best, while many open-weight and smaller domain-adapted models fail to reach the minimum proficiency criterion; generalist models outperform specialized medical fine-tunes at comparable scales.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]