[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-117565-en":3,"doc-seo-117565-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},117565,2336464648746,"Skyler","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Machine learning and statistical inference in microbial population genomics","Large-scale genome datasets have transformed microbiology research, but extracting insight requires computationally intensive analysis and new methodological frameworks. Machine learning and statistical inference share discovery goals yet differ in emphasis: machine learning optimizes prediction, whereas statistical inference aims to interpret the processes connecting variables. This review compares aspirations, principles, and resulting methods with examples from microbial genomics, arguing that combining both perspectives can strengthen pathogen research in the big-data era.","Sheppard et al. Genome Biology (2025) 26:313 [https://doi.org/10.1186/s13059-025-03775-4](https://doi.org/10.1186/s13059-025-03775-4)  \nGenome Biology  \nREVIEW Open Access  \nMachine learning and statistical inference in microbial population genomics  \nSamuel K. Sheppard 1, Nicolas Arning2, David W. Eyre2,3,4 and Daniel J. Wilson2,5*  \n*Correspondence: [daniel.wilson@bdi.ox.ac.uk](daniel.wilson@bdi.ox.ac.uk)  \n1 Ineos Oxford Institute for Antimicrobial Research, Department of Biology, University of Oxford, Oxford, United Kingdom  \n2 Big Data Institute, Oxford Population Health, University of Oxford, Oxford, United Kingdom  \n3 NIHR Oxford Biomedical Research Centre, Oxford, United Kingdom  \n4 NIHR Health Protection Research Unit in Healthcare Associated Infections and Antimicrobial Resistance, University of Oxford, Oxford, United Kingdom  \n5 Oxford University Department for Continuing Education, Oxford, United Kingdom  \nAbstract  \nThe availability of large genome datasets has changed the microbiology research landscape. Analyzing such data requires computationally demanding analyses, and new approaches have come from different data analysis philosophies. Machine learning and statistical inference have overlapping knowledge discovery aims and approaches. However, machine learning focuses on optimizing prediction, whereas statistical inference focuses on understanding the processes relating variables. In this review, we outline the different aspirations, precepts, and resulting methodologies, with examples from microbial genomics. Emphasizing complementarity, we argue that the combination and synthesis of machine learning and statistics has potential for pathogen research in the big data era.  \nBackground  \nAdvances in technology and data generation have driven a big data revolution in microbiology, with studies routinely analyzing thousands of whole genome sequences. Datasets generated with ever-increasing volume, variety, and velocity bring tremendous opportunities as well as unique analysis challenges. Inspired by the promise of deeper understanding and driven by high-throughput low-cost DNA sequencing, there are now vast genome libraries of bacterial species approaching one million genomes [1]. Achieving the potential of these resources has required the scaling of conventional statistical methods which face challenges with high-dimensional data, necessitating simplificationsand approximations [2]. This is paradoxical, because the vast information content of modern resources should make it easier to glean biological insights about evolutionary origins, transmission dynamics and the genetic basis of phenotypic diversity. Machine learning (ML) approaches offer a potential solution as they can handle very large and heterogeneous datasets [3]. ML is a multidisciplinary pursuit that draws heavily on statistics and computer science. Quantitative approaches to exploiting data underpin both endeavors, but for the purposes of this review we work with the following distinction: statistical inference is a tool for furthering our scientific understanding of the world,  \n© The Author(s) 2025. Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit [http://](http://)[ ](http://)[creativecommons.org/license","cbCaijtcwbA9J3gM","https://ap.wps.com/l/cbCaijtcwbA9J3gM","pdf",1394464,1,22,"English","en",105,"# Background\n## Principles of machine learning and statistical inference\n## Machine learning and statistical inference in microbial population genomics","[{\"question\":\"机器学习与统计推断在微生物群体基因组研究中的主要侧重点是什么？\",\"answer\":\"机器学习侧重优化预测，而统计推断侧重理解变量之间所反映的过程，从而提升对科学机制的解释能力。\"},{\"question\":\"为什么大规模基因组数据会带来新的分析挑战？\",\"answer\":\"数据在规模、类型和生成速度上持续增长，导致传统统计方法难以直接处理高维数据，因此需要扩展、简化或近似，并引入更强的计算方法。\"},{\"question\":\"将机器学习和统计推断结合有什么潜在价值？\",\"answer\":\"结合两者的目标与方法互补性，能够在大数据时代提升对病原体相关问题的分析与知识发现能力，例如预测事件、评估变量影响和识别数据模式。\"}]","Machine learning and statistical inference in microbial population genomics | PDF",1785677023,55,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"machine-learning-and-statistical-inference-in-microbial-population-genomics","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/machine-learning-and-statistical-inference-in-microbial-population-genomics/117565/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"机器学习与统计推断在微生物群体基因组研究中的主要侧重点是什么？","Question",{"text":75,"@type":76},"机器学习侧重优化预测，而统计推断侧重理解变量之间所反映的过程，从而提升对科学机制的解释能力。","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"为什么大规模基因组数据会带来新的分析挑战？",{"text":80,"@type":76},"数据在规模、类型和生成速度上持续增长，导致传统统计方法难以直接处理高维数据，因此需要扩展、简化或近似，并引入更强的计算方法。",{"name":82,"@type":73,"acceptedAnswer":83},"将机器学习和统计推断结合有什么潜在价值？",{"text":84,"@type":76},"结合两者的目标与方法互补性，能够在大数据时代提升对病原体相关问题的分析与知识发现能力，例如预测事件、评估变量影响和识别数据模式。","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]