[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118712-en":3,"doc-seo-118712-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},118712,1099513958607,"Jiven","https://ap-avatar.wpscdn.com/avatar/100002390cf8733938c?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778829742770036399",8,"Research & Report","Improving generalization of machine learning-identified biomarkers with causal modeling - an investigation into immune receptor diagnostics","Machine learning increasingly supports discovery of diagnostic and prognostic biomarkers from high-dimensional molecular data, yet experimental choices can limit learnable, clinically generalizable signals. This investigation argues that adopting a causal perspective clarifies how design-related factors relate to robustness and generalization in ML-based diagnostics. Using adaptive immune receptor repertoires (AIRRs) as a concrete high-dimensional biomarker case, it analyzes biological and experimental influences and provides adjustable simulations to study their effects. Results show causal modeling improves robustness by exposing stable variable relations and guiding adjustments for population-specific variations.","Improving generalization of machine learning-identified biomarkers with causal modeling: an investigation into immune receptor diagnostics  \nMilena Pavlović1 ,2 , Ghadi S. Al Hajj1 , Johan Pensar3 , Mollie Wood4 ,5 , Ludvig M. Sollid2 ,6 , Victor Greiff6 , Geir K. Sandve1 ,2  \n1 Centre for Bioinformatics, Department of Informatics, University of Oslo, Norway  \n2 K.G. Jebsen Centre for Coeliac Disease Research, Institute of Clinical Medicine, University of Oslo, Norway  \n3 Department of Mathematics, University of Oslo, Norway  \n4 Department of Pharmacy, University of Oslo, Norway  \n5 Department of Epidemiology, Gillings School of Global Public Health, University of North Carolina at Chapel Hill, USA  \n6 Department of Immunology, University of Oslo and Oslo University Hospital, Norway  \nAbstract  \nMachine learning is increasingly used to discover diagnostic and prognostic biomarkers from high-dimensional molecular data. However, a variety of factors related to experimental design may affect the ability to learn generalizable and clinically applicable diagnostics. Here, we argue that a causal perspective improves the identification of these challenges, and formalizes their relation to the robustness and generalization of machine learning-based diagnostics. To make for a concrete discussion, we focus on a specific, recently established high-dimensional biomarker – adaptive immune receptor repertoires (AIRRs) . We discuss how the main biological and experimental factors of the AIRR domain may influence the learned biomarkers and provide easily adjustable simulations of such effects. In conclusion, we find that causal modeling improves machine learning-based biomarker robustness by identifying stable relations between variables and by guiding the adjustment of the relations and variables that vary between populations.  \nIntroduction  \nHigh-throughput sequencing technologies now allow for the examination of a variety of patient characteristics, such as genetic variation1 , DNA methylation2 , gene expression3 , gut microbiota4 , and adaptive immune receptor repertoires (AIRRs)5 ,6. Proof-of-concept studies showed that such molecular markers hold great promise for disease diagnostics, especially in combination with machine learning ( ML)2 ,3 ,5 ,7 . However, there exist several challenges to using ML for diagnostics. First, the data used in diagnostic studies may be selected based on availability, e.g. , collected from patients visiting the clinic or having a similar genetic background (sometimes referred to as \"convenience sampling\") . Furthermore, rather than originating from a single source, the data might instead be collected at multiple locations or at distinct time points. These factors may introduce systematic differences between datasets, such as measurement errors and batch effects, which need to be taken into account when designing a new study, or adjusted for when the data are already collected. A failure to do so can introduce selection and confounding biases that lead to models failing in real-world application despite showing promising performance during diagnostic development8–11. Finally, biomarker data are typically high-dimensional, which makes it more challenging to disentangle noise and biases from the true markers associated with the disease12 .  \nIn ML-based diagnostics, these challenges are examined from two perspectives. One perspective is purely statistical13 , 14: it attempts to solve the challenges by anticipating how the distributions of features or labels will change (a phenomenon called dataset shift) but does not consider causal relations between them, as discussed by Whalen and colleagues in a genomics setting10. An alternative perspective investigates these challenges using the causal inference framework15–17 , describing dataset shift using formal definitions with respect to a proposed causal model of the underlying process18.  \nThe causal inference framework described by Pearl, known as do-calculus15 ","cbCaitWQT9rQTwh8","https://ap.wps.com/l/cbCaitWQT9rQTwh8","pdf",1288634,1,22,"English","en",105,"# Abstract\n# Introduction\n## Challenges in ML-based diagnostics\n## Two perspectives: statistical vs causal inference\n## Do-calculus, causal graphs, and identifiability\n## Causal models for robustness under dataset shift","[{\"question\":\"Why is generalization a key challenge for machine learning-based biomarker diagnostics?\",\"answer\":\"Diagnostic studies may suffer selection and confounding biases, and high-dimensional biomarker data make it difficult to separate noise and biases from true disease markers. These issues can cause models to fail in real-world use despite good development performance.\"},{\"question\":\"How does causal modeling improve robustness compared with purely statistical approaches?\",\"answer\":\"Causal modeling formalizes dataset shift in terms of a proposed causal model and identifies stable relations between variables. It also guides how to adjust relations and variables that change across populations.\"},{\"question\":\"Why are adaptive immune receptor repertoires (AIRRs) used as the focus biomarker?\",\"answer\":\"AIRRs are a recently established high-dimensional biomarker, providing a concrete setting to discuss how biological and experimental factors influence learned biomarkers and to run easily adjustable simulations of those effects.\"}]","Improving generalization of machine learning-identified biomarkers with causal modeling - an investigation into immune receptor diagnostics | PDF",1785719859,55,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"improving-generalization-of-machine-learning-identified-biomarkers-with-causal-modeling-an-investigation-into-immune-receptor-diagnostics","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/improving-generalization-of-machine-learning-identified-biomarkers-with-causal-modeling-an-investigation-into-immune-receptor-diagnostics/118712/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04","2026-08-03",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why is generalization a key challenge for machine learning-based biomarker diagnostics?","Question",{"text":76,"@type":77},"Diagnostic studies may suffer selection and confounding biases, and high-dimensional biomarker data make it difficult to separate noise and biases from true disease markers. These issues can cause models to fail in real-world use despite good development performance.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does causal modeling improve robustness compared with purely statistical approaches?",{"text":81,"@type":77},"Causal modeling formalizes dataset shift in terms of a proposed causal model and identifies stable relations between variables. It also guides how to adjust relations and variables that change across populations.",{"name":83,"@type":74,"acceptedAnswer":84},"Why are adaptive immune receptor repertoires (AIRRs) used as the focus biomarker?",{"text":85,"@type":77},"AIRRs are a recently established high-dimensional biomarker, providing a concrete setting to discuss how biological and experimental factors influence learned biomarkers and to run easily adjustable simulations of those effects.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]