[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-133535-en":3,"doc-seo-133535-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},133535,687197207057,"Sage","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",7,"Healthcare","Foundation Models Evaluation and Regulation Merlin Abdominal CT FM","This document discusses the evaluation and regulation of Foundation Models (FMs), with a specific focus on Merlin, a Vision Language FM for 3D Computed Tomography (CT). Merlin was trained on 15.5k CT scans and radiology reports, pre-trained using ICD diagnosis codes, and evaluated on a significant number of internal and external studies. The evaluation criteria for Merlin include zero-shot classification, phenotype prediction, retrieval, disease prediction, report generation, and segmentation, utilizing metrics such as F1, AUROC, AUPRC, Recall@k, ROUGE, GREEN, LLM-as-a-judge, Dice, and ASSD. The document also touches upon contemporary LLM evaluation methods, referencing the Chatbot Arena platform and its challenges, particularly the 'Leaderboard Illusion'. A key question posed is whether to evaluate and regulate the model itself or the tasks it performs. The healthcare LLM evaluation landscape is questioned for potential improvements. The document also includes a visual representation of an abdominal CT scan, illustrating the application context of such FMs in medical imaging.","Foundation Models Evaluation and Regulation  \nMerlin Abdominal CT FM  \n• Trained on 15.5k CT scans and corresponding radiology reports (6M tokens)  \n• Pre-trained using ICD diagnosis codes  \n• Evaluated on 5k internal and 5k external studies  \nBlankemeier et al. Merlin: A Vision Language Foundation Model for 3D Computed Tomography. arXiv 2024.  \nMerlin: 3D Abdominal CT FM  \nBlankemeier et al. Merlin: A Vision Language Foundation Model for 3D Computed Tomography. arXiv 2024.  \nMerlin Evaluation Criteria  \n• Zero-Shot Classification: F1, AUROC, etc  \n• Phenotype Prediction: AUROC, AUPRC, etc  \n• Retrieval: Recall @k  \n• Disease Prediction: AUROC, AUPRC, etc  \n• Report Generation: ROUGE, GREEN, LLM-as-a-judge. etc  \n• Segmentation: Dice, ASSD, etc  \nFoundation Models  \n• Do we evaluate/regulate the model or the tasks?  \nContemporary LLM Evaluation  \nChiang et al. \"Chatbot arena: an open platform for evaluating LLMs by human preference”. ICML 2024.  \nContemporary LLM Evaluation  \nChiang et al. \"Chatbot arena: an open platform for evaluating  \nLLMs by human preference”. ICML 2024.  \nChallenges with Chatbot Arena  \nSingh et al. The Leaderboard Illusion. 2025.  \nHealthcare LLM Eval Any Better?","cbCaiaWZWIif5mfl","https://ap.wps.com/l/cbCaiaWZWIif5mfl","pdf",4152128,3,1,30,"English","en",105,"# Foundation Models Evaluation and Regulation\nMerlin Abdominal CT FM\nMerlin Evaluation Criteria\nContemporary LLM Evaluation\nChallenges with Chatbot Arena\nHealthcare LLM Eval Any Better?","[{\"question\":\"What is Merlin and what is its primary function?\",\"answer\":\"Merlin is a Vision Language Foundation Model specifically designed for 3D Computed Tomography (CT) scans, focusing on abdominal imaging. It is trained to process CT scans and their corresponding radiology reports.\"},{\"question\":\"What are the key evaluation criteria used for Merlin?\",\"answer\":\"Merlin is evaluated on several criteria including zero-shot classification, phenotype prediction, retrieval, disease prediction, report generation, and segmentation, using a variety of metrics like F1, AUROC, ROUGE, and Dice.\"},{\"question\":\"What are some of the challenges in contemporary LLM evaluation mentioned in the document?\",\"answer\":\"The document references challenges such as the 'Leaderboard Illusion' in LLM evaluation platforms like the Chatbot Arena, suggesting that current evaluation methods may not always provide a complete or accurate picture of performance.\"}]","Foundation Models Evaluation and Regulation Merlin Abdominal CT FM | PDF",1787221477,76,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"foundation-models-evaluation-and-regulation-merlin-abdominal-ct-fm","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,51],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":20},"https://docshare.wps.com/document/healthcare/",{"item":52,"name":13,"@type":44,"position":53},"https://docshare.wps.com/document/foundation-models-evaluation-and-regulation-merlin-abdominal-ct-fm/133535/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-09-01","2026-08-20",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What is Merlin and what is its primary function?","Question",{"text":76,"@type":77},"Merlin is a Vision Language Foundation Model specifically designed for 3D Computed Tomography (CT) scans, focusing on abdominal imaging. It is trained to process CT scans and their corresponding radiology reports.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What are the key evaluation criteria used for Merlin?",{"text":81,"@type":77},"Merlin is evaluated on several criteria including zero-shot classification, phenotype prediction, retrieval, disease prediction, report generation, and segmentation, using a variety of metrics like F1, AUROC, ROUGE, and Dice.",{"name":83,"@type":74,"acceptedAnswer":84},"What are some of the challenges in contemporary LLM evaluation mentioned in the document?",{"text":85,"@type":77},"The document references challenges such as the 'Leaderboard Illusion' in LLM evaluation platforms like the Chatbot Arena, suggesting that current evaluation methods may not always provide a complete or accurate picture of performance.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,119,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":117,"slug":118},40,"healthcare",{"id":120,"doc_module":4,"doc_module_name":47,"category_name":121,"show_sort_weight":22,"slug":122},8,"Research & Report","research-report",{"id":124,"doc_module":4,"doc_module_name":47,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":47,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":47,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":47,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]