[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-213541-en":3,"doc-seo-213541-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},213541,1236954412713,"McGucket","https://us-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","WorldMedQA-V - a multilingual, multimodal medical examination dataset for multimodal language models evaluation","Multimodal/vision language models are increasingly used in healthcare, requiring reliable benchmarks to verify safety, efficacy, and fairness. Existing medical QA datasets are often text-only and limited in languages and countries. WorldMedQA-V presents a multilingual, multimodal benchmark built from national medical examinations, pairing 568 multiple-choice questions with 568 clinically validated medical images across Brazil, Israel, Japan, and Spain, with local language items and verified English translations. Baselines evaluate models with and without images in both local language and English, supporting more equitable and representative deployments.","WorldMedQA-V: a multilingual, multimodal medical examination dataset for multimodal language models evaluation  \nJoão Matos 1 * , Shan Chen2 ,3 ,4 * , Siena Placino5 , Yingya Li2 ,4 , Juan Carlos Climent Pardo2 ,3  \nDaphna Idan6 , Takeshi Tohyama7 ,9 , David Restrepo7 , Luis F. Nakayama7 Jose M. M. Pascual-Leone8 , Guergana Savova2 ,4 , Hugo Aerts2 ,3 , 10 , Leo A. Celi2 ,7 , 11  \nA. Ian Wong 12 , Danielle S. Bitterman2 ,3 ,4 , Jack Gallifant2 ,3†  \n1 Oxford, 2Harvard, 3Mass General Brigham, 4Boston Children’s Hospital,  \n5 St. Luke’s Medical Center, 6Ben-Gurion University of the Negev, 7MIT, 8Alcalá University, 9International University of Health and Welfare, 10Maastricht University, 11BIDMC, 12Duke  \nAbstract  \nMultimodal/vision language models (VLMs) are increasingly being deployed in healthcare settings worldwide, necessitating robust benchmarks to ensure their safety, efficacy, and fairness. Multiple-choice question and answer (QA) datasets derived from national medical examinations have long served as valuable evaluation tools, but existing datasets are largely text-only and available in a limited subset of languages and countries. To address these challenges, we present WorldMedQA-V, an updated multilingual, multimodal benchmarking dataset designed to evaluate VLMs in healthcare. WorldMedQA-V includes 568 labeled multiple-choice QAs paired with 568 medical images from four countries (Brazil, Israel, Japan, and Spain), covering original languages and validated English translations by native clinicians, respectively. Baseline performance for common open-and closed-source models are provided in the local language and English translations, and with and without images provided to the model. The WorldMedQA-V benchmark aims to better match AI systems to the diverse healthcare environments in which they are deployed, fostering more equitable, effective, and representative applications.1  \n1 Introduction  \nGenerative artificial intelligence (AI) models are increasingly being adopted in healthcare, highlighting the need for robust benchmarks to assess their safety, efficacy, and fairness (Thirunavukarasuet al., 2023 ; Clusmann et al., 2023 ; Abbasian et al., 2024 ; Wiggers, 2024) .  \nOne of the key evaluation tasks in Natural Language Processing (NLP) is Question Answering  \n* Co-first authors: João Matos and Shan Chen  \n†Corresponding author: [jgallifant@bwh.harvard.edu](jgallifant@bwh.harvard.edu)  \n1All code is accessible on [https://github.com/](https://github.com/)[ ](https://github.com/)WorldMedQA/V and the dataset on [https://huggingface](https://huggingface). com/datasets/WorldMedQA/V.  \nFigure 1: WorldMedQA-V dataset generation and evaluation workflows.  \n(QA)(Yu et al., 2024 ; Fan et al., 2023), which involves building systems that can automatically respond to human queries in natural language by combining language understanding with information retrieval (Jin et al., 2020) . Multi-choice QA benchmarks have become essential not only for evaluating large language models (LLMs) but also for assessing vis language models (VLMs) in medicine (Liu et al., 2024) .  \nRecent research has explored the performance of LLMs in medical exams, with ChatGPT being the first AI system to pass the USMLE (Kung et al., 2023), prompting further studies (Gobira et al., 2023 ; Liu et al., 2024 ; Chen et al., 2024b) . A recent review identified 45 studies on ChatGPT’s performance in medical exams (Liu et al., 2024), but VLMs remain underexplored in medical tasks (Yanet al., 2023 ; Wu et al., 2023) . Despite progress, current models face limitations such as context fragility, biases, and inconsistent multilingual per-  \n7218  \nFindings of the Association for Computational Linguistics: NAACL 2025, pages 7218–7231  \nApril 29-May 4, 2025 ©2025 Association for Computational Linguistics  \nformance (Gallifant et al., 2024 ; Zack et al., 2024 ; Chen et al., 2024a) . There is also a need for more diverse datasets to ensure equitable AI evaluation in hea","cbCaikZz3dyJmmDf","https://ap.wps.com/l/cbCaikZz3dyJmmDf","pdf",2365980,2,1,14,"English","en",105,"# Introduction\n## Evaluation needs for healthcare VLMs\n## WorldMedQA-V contributions\n# Related Work\n## Multilingual and multimodal VLM benchmarks","[{\"question\":\"What is WorldMedQA-V and what does it evaluate?\",\"answer\":\"WorldMedQA-V is an updated multilingual, multimodal medical examination dataset designed to evaluate vision-language models in healthcare using multiple-choice question answering paired with medical images.\"},{\"question\":\"How many questions and images are included, and which countries are covered?\",\"answer\":\"WorldMedQA-V includes 568 labeled multiple-choice QAs paired with 568 medical images from four countries: Brazil, Israel, Japan, and Spain.\"},{\"question\":\"How does the benchmark measure the impact of images and language translations?\",\"answer\":\"It provides baseline performance for models with and without images, and reports results in both local languages and validated English translations by native clinicians, enabling comparisons of performance differentials.\"}]","WorldMedQA-V - a multilingual, multimodal medical examination dataset for multimodal language models evaluation | PDF",1788762594,35,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"worldmedqa-v-a-multilingual-multimodal-medical-examination-dataset-for-multimodal-language-models-evaluation","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,48,51],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":20},"https://docshare.wps.com/document/","Document",{"item":49,"name":12,"@type":44,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":44,"position":53},"https://docshare.wps.com/document/worldmedqa-v-a-multilingual-multimodal-medical-examination-dataset-for-multimodal-language-models-evaluation/213541/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-09-11","2026-09-07",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What is WorldMedQA-V and what does it evaluate?","Question",{"text":76,"@type":77},"WorldMedQA-V is an updated multilingual, multimodal medical examination dataset designed to evaluate vision-language models in healthcare using multiple-choice question answering paired with medical images.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How many questions and images are included, and which countries are covered?",{"text":81,"@type":77},"WorldMedQA-V includes 568 labeled multiple-choice QAs paired with 568 medical images from four countries: Brazil, Israel, Japan, and Spain.",{"name":83,"@type":74,"acceptedAnswer":84},"How does the benchmark measure the impact of images and language translations?",{"text":85,"@type":77},"It provides baseline performance for models with and without images, and reports results in both local languages and validated English translations by native clinicians, enabling comparisons of performance differentials.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]