[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83885-en":3,"doc-seo-83885-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83885,8796095461564,"Liam","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Multi-Large Language Model Orchestrated Severity Assessment of Clinical Records (MOSAIC)","Agentic large language model systems offer a way to synthesize clinical evidence and reason over electronic health records, yet their effectiveness for disease severity phenotyping remains unevaluated. MOSAIC (Multi-agent Orchestrated Severity Assessment In Clinical Records) is a two-phase agentic LLM framework evaluated for type 2 diabetes on a synthetic EHR cohort against established algorithmic ground truths, with alignment assessed against mortality and incident complications. Results show meaningful clinical stratification, biomarker- and social-determinant coverage beyond baselines, and multi-agent reasoning benefits over deterministic rule execution.","Title Page  \nOriginal Article  \nMulti-Large Language Model Orchestrated Severity Assessment of Clinical Records (MOSAIC)  \nManuela Del Castillo Suero 1, Arnault-Quentin Vermillet 1, Nicole Sonne Heckmann 1, Darmendra Ramcharran2,3, Maurizio Sessa 1  \n1Department of Drug Design and Pharmacology, University of Copenhagen, Copenhagen, Denmark  \n2GSK, Providence, RI, USA  \n3 School of Pharmacy, University of Rhode Island, Kingston, RI, USA  \nCorresponding author:  \nMaurizio Sessa,  \nDepartment of Drug Design and Pharmacology, University of Copenhagen,  \nJagtvej 160, 2100 Copenhagen, Denmark [Email: ](Email: maurizio.sessa@sund.ku.dk)[maurizio.sessa@sund.ku.dk](Email: maurizio.sessa@sund.ku.dk)  \nAbstract  \nBackground: Disease severity is a multidimensional construct difficult to capture with rulebased approaches in Electronic Healthcare Records (EHR) . Agentic large language model (LLM) systems offer a potential approach for synthesising clinical evidence and reasoning over EHRs, but their performance for this task has been yet not evaluated.  \nMethods: MOSAIC (Multi-agent Orchestrated Severity Assessment In Clinical Records) is a two-phase agentic LLM framework for severity phenotyping, using type 2 diabetes (T2D) as a proof-of-concept. MOSAIC was evaluated on a synthetic EHR cohort (SyntheticMass; openweight cohort N = 4,886; closed-weight benchmark N = 200) against three established algorithmic ground truths (DCSI GT, DiSSCo, Cooper GT) and by assessing the alignment of its severity classifications with hard clinical outcomes, namely all-cause mortality and incident complications. Open-weight (locally deployable) and proprietary LLM pipelines were also compared.  \nResults: The generated framework extended beyond the domain coverage of the comparators, incorporating biomarker-based glycaemic staging, beta-cell function markers, and social determinants of health not represented in the expert-defined algorithms. Open-weight MOSAIC performed comparably to the proprietary pipeline (closed- vs open-weight κw = 0.773) and achieved moderate agreement with Cooper GT (κw = 0.597) and DCSI GT (κw = 0.534) and fair agreement with DiSSCo GT (κw = 0.320). Agent-based (Type 1) severity tiers showed significant separation of all-cause mortality (log-rank p \u003C 0.001; crude hazard ratios 1.6–2.4 for non-Baseline tiers), with survival declining across tiers although not strictly monotonically at the upper tiers, and a significant inverse gradient for incident complications (log-rank p \u003C 0.001) consistent with depletion of susceptibles in higher-severity patients. Agentic classification also differed substantially from deterministic execution of the same rubric (MOSAIC Frozen; κw = 0.428), indicating a contribution of multi-agent reasoning beyond fixed rule execution.  \nConclusion: MOSAIC demonstrates that agentic LLM systems can generate and apply clinically meaningful severity phenotypes from structured EHR data when applied to T2D. Broadening the application of MOSAIC to assess severity of other disease phenotypes beyond T2D is an area that warrants further research and evaluation.  \nKeywords: pharmacoepidemiology; artificial intelligence; large language model; severity phenotyping; type 2 diabetes; electronic health records; agentic systems.  \n1. Introduction  \nThe growing availability of longitudinal electronic health records (EHRs) as secondary data sources has expanded opportunities for large-scale pharmacoepidemiology (PE) research, but realising this potential depends on the ability to transform raw clinical information into structured, computable patient representations, a process known as phenotyping (1, 2) . Substantial progress has been made in EHR phenotyping using rule-based algorithms, machine learning, and natural language processing (3, 4); however, scalability remains constrained by static representations or extensive manual input, a limitation that becomes particularly problematic when phenotype definitions evolve over time (4, ","cbCaimF48igOcgJx","https://ap.wps.com/l/cbCaimF48igOcgJx","pdf",2799545,5,1,61,"English","en",105,"# Introduction\n## Phenotyping and severity challenges\n## Agentic AI for structured reasoning\n# Methods\n## Data source\n## MOSAIC framework overview\n# Results\n## Agreement with algorithmic ground truths\n## Clinical outcome stratification\n## Comparison to deterministic execution\n# Conclusion","[{\"question\":\"What is MOSAIC and what problem does it address?\",\"answer\":\"MOSAIC is a two-phase agentic LLM framework designed to perform severity phenotyping from structured EHR data. It targets the challenge that disease severity is multidimensional and difficult to capture with static rule-based approaches.\"},{\"question\":\"How was MOSAIC evaluated in the study?\",\"answer\":\"MOSAIC was tested on a synthetic EHR cohort (SyntheticMass) using type 2 diabetes as a proof-of-concept. Performance was compared against three algorithmic ground truths and assessed by how well severity classifications aligned with all-cause mortality and incident complications.\"},{\"question\":\"What were the main findings about MOSAIC’s severity classification?\",\"answer\":\"MOSAIC extended domain coverage by incorporating biomarker-based glycaemic staging, beta-cell function markers, and social determinants of health not present in expert-defined algorithms. It showed survival separation across severity tiers and meaningful agreement with the ground truths, while differing substantially from deterministic execution of the same rubric.\"}]",1784191217,154,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"multi-large-language-model-orchestrated-severity-assessment-of-clinical-records-mosaic","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/multi-large-language-model-orchestrated-severity-assessment-of-clinical-records-mosaic/83885/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What is MOSAIC and what problem does it address?","Question",{"text":76,"@type":77},"MOSAIC is a two-phase agentic LLM framework designed to perform severity phenotyping from structured EHR data. It targets the challenge that disease severity is multidimensional and difficult to capture with static rule-based approaches.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How was MOSAIC evaluated in the study?",{"text":81,"@type":77},"MOSAIC was tested on a synthetic EHR cohort (SyntheticMass) using type 2 diabetes as a proof-of-concept. Performance was compared against three algorithmic ground truths and assessed by how well severity classifications aligned with all-cause mortality and incident complications.",{"name":83,"@type":74,"acceptedAnswer":84},"What were the main findings about MOSAIC’s severity classification?",{"text":85,"@type":77},"MOSAIC extended domain coverage by incorporating biomarker-based glycaemic staging, beta-cell function markers, and social determinants of health not present in expert-defined algorithms. It showed survival separation across severity tiers and meaningful agreement with the ground truths, while differing substantially from deterministic execution of the same rubric.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]