[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-216503-en":3,"doc-seo-216503-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},216503,687207024643,"Rhys","https://ap-avatar.wpscdn.com/davatar_3d24733baf745e90a7e4bdd5f77d97b2",8,"Research & Report","Zero-shot Performance of Generative AI in Brazilian Portuguese Medical Exam","Artificial intelligence can improve healthcare through better diagnostic accuracy, optimized clinical workflows, and more personalized treatment planning. Large Language Models and Multimodal Large Language Models have progressed in medical NLP, yet most evaluations emphasize English, creating uncertainty about cross-language reliability. This study measures zero-shot performance of multiple LLMs and MLLMs on Brazilian spoken Portuguese questions from the HCFMUSP medical residency entrance exam, benchmarking accuracy, processing time, and explanation coherence against human candidates.","arXiv :2507 . 19885v 1 [ cs .CL] 26 Jul 2025  \nZero-shot Performance of Generative AI in Brazilian Portuguese Medical Exam  \nC´esar Augusto Madid Truyts 1,2,* , Amanda Gomes Rabelo 1,2 , Gabriel Mesquita de Souza 1 , Daniel Scaldaferri Lages 1 , Adriano Jos´e Pereira 1,2 , Uri Adrian Prync Flato2 , Eduardo Pontes dos Reis 1,3 , Joaquim Edson Vieira4,6 , Paulo Sergio Panse Silveira5 , and  \nEdson Amaro Junior 1  \n1 Einstein Global Advanced Technologies for Equity, Hospital Israelita Albert Einstein  \n2 Departamento de Pacientes Graves, Hospital Israelita Albert Einstein  \n3 Stanford Center for Artificial Intelligence in Medicine and  \nImaging, Stanford University  \n4 Anestesiologia, Departmento de Cirurgia, Faculdade de Medicina, Universidade de S˜ao Paulo, SP, Brasil  \n5 Inform´atica M´edica, Departamento de Patologia, Faculdade de Medicina da Universidade de S˜ao Paulo  \n6 Faculdade Israelita de Ciˆencias da Sa´ude Albert Einstein  \nHospital Israelita Albert Einstein, SP, Brasil  \n* Corresponding author:  \nEinstein Global Advanced Technologies for Equity Hospital Israelita Albert Einstein  \nC´esar Augusto Madid Truyts  \nAv. Albert Einstein, 627/701, Sao Paulo, SP, Brazil postcode: 05652-900  \nJuly 29, 2025  \n2  \nAbstract  \nArtificial intelligence (AI) has shown the potential to revolutionize healthcare by improving diagnostic accuracy, optimizing workflows, and personalizing treatment plans. Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) have achieved notable advancements in natural language processing and medical applications. However, the evaluation of these models has focused predominantly on the English language, leading to potential biases in their performance across different languages.  \nThis study investigates the capability of six LLMs (GPT-4.0 Turbo, LLaMA-3-8B, LLaMA-3-70B, Mixtral 8x7B Instruct, Titan Text G1-Express, and Command R+) and four MLLMs (Claude-3.5-Sonnet, Claude-3-Opus, Claude-3-Sonnet, and Claude-3-Haiku) to answer questions written in Brazilian spoken portuguese from the medical residency entrance exam of the Hospital das Cl´ınicas da Faculdade de Medicina da Universidade de S˜ao Paulo (HCFMUSP) - the largest health complex in South America. The performance of the models was benchmarked against human candidates, analyzing accuracy, processing time, and coherence of the generated explanations.  \nThe results show that while some models, particularly Claude-3.5-Sonnet and Claude-3-Opus, achieved accuracy levels comparable to human candidates, performance gaps persist, particularly in multimodal questions requiring image interpretation. Furthermore, the study highlights language disparities, emphasizing the need for further fine-tuning and data set augmentation for non-English medical AI applications.  \nOur findings reinforce the importance of evaluating generative AI in various linguistic and clinical settings to ensure a fair and reliable deployment in healthcare. Future research should explore improved training methodologies, improved multimodal reasoning, and real-world clinical integration of AI-driven medical assistance.  \nKeywords: Generative Artificial Intelligence, Large Language Models, Medical Education  \n1 Introduction  \nLarge language models (LLM), have revolutionized the interpretation of data [1, 2] . Artificial intelligence (AI) has the promise to transform healthcare by improving diagnostic accuracy, personalizing treatment plans, and optimizing workflows in medical practice by extracting value and information from unstructured data, which predominate in electronic health records [1, 3–7] .  \nAlthough LLMs have been tested on benchmarks such as Massive Multitask Language Understanding and BIG Bench [8, 9], these evaluations are conducted predominantly in English, reflecting the overwhelming dominance of this language in training data. Specialized datasets, such as the MedQA-US Medical Licensing Examination (USMLE), have been used to assess the capabilities of ","cbCaivevwZn5nAqy","https://ap.wps.com/l/cbCaivevwZn5nAqy","pdf",1154895,1,26,"English","en",105,"# Abstract\n## Introduction","[{\"question\":\"What problem does the study address in evaluating generative AI for healthcare?\",\"answer\":\"It examines how performance of generative AI can vary by language, since prior evaluations mostly focus on English and may introduce biases in high-stakes medical contexts.\"},{\"question\":\"Which tasks and benchmarks are used to test the models?\",\"answer\":\"Models are tested in a zero-shot setting on Brazilian spoken Portuguese questions from the HCFMUSP medical residency entrance exam, with evaluation based on accuracy, processing time, and coherence of generated explanations.\"},{\"question\":\"How do the results compare to human candidate performance?\",\"answer\":\"Some models, particularly Claude-3.5-Sonnet and Claude-3-Opus, reach accuracy comparable to human candidates, while performance gaps remain, especially for multimodal questions requiring image interpretation.\"}]","Zero-shot Performance of Generative AI in Brazilian Portuguese Medical Exam | PDF",1788816650,66,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"zero-shot-performance-of-generative-ai-in-brazilian-portuguese-medical-exam","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/zero-shot-performance-of-generative-ai-in-brazilian-portuguese-medical-exam/216503/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-09-11","2026-09-07",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does the study address in evaluating generative AI for healthcare?","Question",{"text":76,"@type":77},"It examines how performance of generative AI can vary by language, since prior evaluations mostly focus on English and may introduce biases in high-stakes medical contexts.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"Which tasks and benchmarks are used to test the models?",{"text":81,"@type":77},"Models are tested in a zero-shot setting on Brazilian spoken Portuguese questions from the HCFMUSP medical residency entrance exam, with evaluation based on accuracy, processing time, and coherence of generated explanations.",{"name":83,"@type":74,"acceptedAnswer":84},"How do the results compare to human candidate performance?",{"text":85,"@type":77},"Some models, particularly Claude-3.5-Sonnet and Claude-3-Opus, reach accuracy comparable to human candidates, while performance gaps remain, especially for multimodal questions requiring image interpretation.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]