[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-213701-en":3,"doc-seo-213701-105":31,"detail-sidebar-cat-0-en-105":97},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},213701,962084925502,"Emma Mercer","https://ap-avatar.wpscdn.com/davatar_6f874abed73319feea01a86fa6f0fab8",8,"Research & Report","HealthQA-BR - A System-Wide Benchmark Reveals Critical Knowledge Gaps in Large Language Models","Healthcare evaluation of Large Language Models (LLMs) has been dominated by physician-centric, English-language benchmarks, creating an illusion of competence that overlooks interprofessional patient care. HealthQA-BR introduces a first large-scale, system-wide Portuguese benchmark for Brazilian healthcare, using 5,632 questions from national licensing and residency exams. Zero-shot testing across 20+ leading models shows “spiky” gaps: high overall accuracy can hide severe specialty-specific weaknesses. Releasing HealthQA-BR and the evaluation suite enables granular, safety-relevant audits for the entire healthcare team.","arXiv :2506 .21578v1 [ cs .CL] 16 Jun 2025  \nHealthQA-BR: A System-Wide Benchmark Reveals Critical Knowledge Gaps in Large  \nLanguage Models  \nAndrew Maranhão Ventura D’addario 1  \n1 Independent Researcher  \nAbstract  \nThe evaluation of Large Language Models (LLMs) in healthcare has been dominated by physician-centric, English-language benchmarks, creating a dangerous illusion of competence that ignores the interprofessional nature of patient care. To provide a more holistic and realistic assessment, we introduce HealthQA-BR, the first large-scale, system-wide benchmark for Portuguese-speaking healthcare. Comprising 5,632 questions from Brazil’s national licensing and residency exams, it uniquely assesses knowledge not only in medicine and its specialties but also in nursing, dentistry, psychology, social work, and other allied health professions.  \nWe conducted a rigorous zero-shot evaluation of over 20 leading LLMs. Our results reveal that while state-of-the-art models like GPT 4.1 achieve high overall accuracy (86.6%), this top-line score masks alarming, previously unmeasured deficiencies. A granular analysis shows performance plummets from near-perfect in specialties like Ophthalmology (98.7%) to barely passing in Neurosurgery (60.0%) and, most notably, Social Work (68.4%) . This\"spiky\" knowledge profile is a systemic issue observed across all models, demonstrating that high-level scores are insufficient for safety validation. By publicly releasing HealthQABR and our evaluation suite, we provide a crucial tool to move beyond single-score evaluations and toward a more honest, granular audit of AI readiness for the entire healthcare team.  \n1 Introduction  \nThe rapid ascent of Large Language Models (LLMs) has been marked by remarkable achievements on physician-centric benchmarks, with leading models now meeting or exceeding the passing threshold on high-stakes assessments like the United States Medical Licensing Examination (USMLE) [1] . While impressive, these headline figures can create a dangerous illusion of competence. A single accuracy score, derived from a narrow set of tasks, risks masking critical, specialty-specific knowledge gaps and fundamentally misrepresents an AI’s readiness for the complex reality of patient care.  \nThis illusion of competence stems from two foundational biases in current evaluation paradigms. The first is a well-documented geographical and linguistic myopia. As recent work on the AfriMed-QA benchmark has shown, model performance can degrade significantly when tested outside of the high-resource, English-language contexts that dominate training data [2] . Our work complements this geographical perspective by introducing a new, orthogonal dimension of evaluation: the professional diversity within a single, integrated healthcare system. Modern healthcare is not the domain of a single physician but a collaborative, interprofessional  \neffort where nurses, pharmacists, therapists, and social workers contribute essential expertise. This system-wide reality of care—a team sport—has been largely ignored by physician-centric benchmarks, leaving a critical blind spot in our understanding of AI capabilities [3] .  \nTo address these gaps, we introduce HealthQA-BR, the first large-scale, system-wide benchmark for Portuguese-speaking healthcare. Comprising 5,632 questions from Brazil’s national licensing and residency exams, HealthQA-BR moves beyond a monolithic focus on medicine to include dentistry, psychology, nursing, social work, and other allied health professions. Its design enables the granular, specialty-level analysis required to move past a single score and uncover the “spiky” knowledge profile of an LLM, pinpointing specific areas of both strength and alarming weakness.  \nUltimately, this work is motivated by the “last mile” problem of AI adoption: the challenge of moving from passing an exam to earning the trust of clinicians, patients, and health systems [4 , 5] . True and safe adoption requ","cbCaicN5LTXcSfLB","https://ap.wps.com/l/cbCaicN5LTXcSfLB","pdf",1263253,2,1,14,"English","en",105,"# Abstract\n# Introduction\n# The HealthQA-BR Dataset\n## Data Sources and Rationale","[{\"question\":\"What problem does HealthQA-BR address in evaluating healthcare LLMs?\",\"answer\":\"It addresses the overreliance on physician-centric, English-language benchmarks that can mask critical knowledge gaps in the real interprofessional nature of patient care.\"},{\"question\":\"What does the HealthQA-BR dataset include and how many questions are there?\",\"answer\":\"HealthQA-BR contains 5,632 multiple-choice questions sourced from Brazil’s national licensing and residency exams, covering medicine and multiple allied health disciplines.\"},{\"question\":\"What do the zero-shot evaluation results show about model performance?\",\"answer\":\"Overall accuracy can look high (e.g., 86.6%), yet a granular analysis reveals large “spiky” weaknesses across specialties, including notably lower performance in areas like Neurosurgery and Social Work.\"}]","HealthQA-BR - A System-Wide Benchmark Reveals Critical Knowledge Gaps in Large Language Models | PDF",1788766179,35,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":92,"head_meta":94,"extra_data":96,"updated_unix":29},"healthqa-br-a-system-wide-benchmark-reveals-critical-knowledge-gaps-in-large-language-models","",{"@graph":37,"@context":91},[38,54,74],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,48,51],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":20},"https://docshare.wps.com/document/","Document",{"item":49,"name":12,"@type":44,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":44,"position":53},"https://docshare.wps.com/document/healthqa-br-a-system-wide-benchmark-reveals-critical-knowledge-gaps-in-large-language-models/213701/",4,{"url":52,"name":13,"@type":55,"image":56,"author":61,"headline":13,"publisher":63,"fileFormat":66,"inLanguage":24,"description":14,"dateModified":67,"datePublished":68,"encodingFormat":66,"isAccessibleForFree":69,"interactionStatistic":70},"DigitalDocument",{"url":57,"@type":58,"width":59,"height":60},"https://docshare.wps.com/thumbnails/healthqa-br-a-system-wide-benchmark-reveals-critical-knowledge-gaps-in-large-language-models/213701.png","ImageObject",300,407,{"name":9,"@type":62},"Person",{"url":42,"name":64,"@type":65},"DocShare","Organization","application/pdf","2026-09-11","2026-09-07",true,{"@type":71,"interactionType":72,"userInteractionCount":20},"InteractionCounter",{"@type":73},"ViewAction",{"@type":75,"mainEntity":76},"FAQPage",[77,83,87],{"name":78,"@type":79,"acceptedAnswer":80},"What problem does HealthQA-BR address in evaluating healthcare LLMs?","Question",{"text":81,"@type":82},"It addresses the overreliance on physician-centric, English-language benchmarks that can mask critical knowledge gaps in the real interprofessional nature of patient care.","Answer",{"name":84,"@type":79,"acceptedAnswer":85},"What does the HealthQA-BR dataset include and how many questions are there?",{"text":86,"@type":82},"HealthQA-BR contains 5,632 multiple-choice questions sourced from Brazil’s national licensing and residency exams, covering medicine and multiple allied health disciplines.",{"name":88,"@type":79,"acceptedAnswer":89},"What do the zero-shot evaluation results show about model performance?",{"text":90,"@type":82},"Overall accuracy can look high (e.g., 86.6%), yet a granular analysis reveals large “spiky” weaknesses across specialties, including notably lower performance in areas like Neurosurgery and Social Work.","https://schema.org",{"og:url":52,"og:type":93,"og:title":13,"og:site_name":64,"og:description":14},"article",{"robots":95,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":98},[99,103,107,111,116,121,126,129,134,137,141],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":100,"show_sort_weight":101,"slug":102},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":104,"show_sort_weight":105,"slug":106},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},"Exam",70,"exam",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},5,"Comic",60,"comic",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},6,"Technology",50,"technology",{"id":122,"doc_module":4,"doc_module_name":47,"category_name":123,"show_sort_weight":124,"slug":125},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":127,"slug":128},30,"research-report",{"id":130,"doc_module":4,"doc_module_name":47,"category_name":131,"show_sort_weight":132,"slug":133},9,"Religion & Spirituality",20,"religion-spirituality",{"id":132,"doc_module":4,"doc_module_name":47,"category_name":135,"show_sort_weight":132,"slug":136},"World Cup","world-cup",{"id":138,"doc_module":4,"doc_module_name":47,"category_name":139,"show_sort_weight":138,"slug":140},10,"Lifestyle","lifestyle",{"id":142,"doc_module":4,"doc_module_name":47,"category_name":143,"show_sort_weight":112,"slug":144},19,"General","general"]