[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118667-en":3,"doc-seo-118667-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},118667,13056703020460,"Valentina","https://ap-avatar.wpscdn.com/avatar/be000253dac470eee5d?_k=1778207105932848923",8,"Research & Report","Predicting Biofilm Formation in Pseudomonas aeruginosa Using Machine Learning","This study investigated whether whole genome sequence data could be used to predict biofilm formation in Pseudomonas aeruginosa using machine learning. K-mer frequency profiles were extracted from 870 P. aeruginosa genomes and used as features in supervised classification, regression, and unsupervised clustering. Across extensive modeling experiments, predictive performance remained poor, indicating insufficient data to support robust learning. Exploratory analyses still identified a subset of consistently selected k-mers linked to misclassification behavior, suggesting a weak but potentially meaningful biological signal. The work outlines limitations of applying machine learning to small, high-dimensional biological datasets and motivates improved data quality for future studies.","Predicting Biofilm Formation in Pseudomonas aeruginosa Using  \nMachine Learning  \nby  \nCorinne Jackson  \nA Thesis  \npresented to  \nThe University of Guelph  \nIn partial fulfilment of requirements  \nfor the degree of  \nMaster of Science  \nin  \nBioinformatics  \nCollaborative Specialization in Artificial Intelligence  \nGuelph, Ontario, Canada  \n© Corinne Jackson, October, 2025  \nAbstract  \nPredicting Biofilm Formation in Pseudomonas aeruginosa Using Machine Learning  \nCorinne Jackson Advisor(s):  \nUniversity of Guelph, 2025 Dr. Cezar Khursigara  \nThis study investigated whether whole genome sequence data could be used to predict biofilm formation in Pseudomonas aeruginosa using machine learning. K-mer frequency profiles were extracted from 870 P. aeruginosa genomes and used as features in multiple modeling approaches, including supervised classification, regression, and unsupervised clustering. Despite extensive experimentation, predictive performance remained poor across all methods, suggesting that the available dataset lacked the volume needed to support robust learning. However, exploratory analyses revealed a subset of k-mers that were consistently selected across model runs and associated with misclassification behaviour, indicating the presence of a weak but potentially meaningful biological signal. The study highlights the limitations of applying machine learning to small, high-dimensional biological datasets , underscoring the importance of data quality. While accurate prediction was not achieved, the project provides a framework for future studies and demonstrates the value of interpretable machine learning pipelines in exploratory genomic analysis.  \nAcknowledgements  \nFirst and foremost, I would like to thank my incredible partner, Griffin, for his unwavering love, patience, and encouragement throughout this journey.  \nTo my parents and family , thank you for instilling in me the curiosity and resilience that carried me through this journey. Your constant support and understanding created the foundation on which this thesis stands. I am equally grateful to Griffin’s parents and family for welcoming me as one of their own and cheering me on from the very beginning.  \nTo my wonderful friends, Emily Hyatt, Daniel Peña, Alice Wang, and Jessica Winslade, thank you for lifting me up through the highs and lows, and for reminding me to laugh even on the most challenging days.  \nI owe sincere thanks to the members of the Khusigara lab for fostering such a collaborative and supportive environment. In particular, Alyssa Banaag and Nicole Garnier provided crucial assistance that was instrumental to the success of my research.  \nI am profoundly grateful to my advisor, Dr. Cezar Khursigara, for his guidance , insight, and mentorship throughout this journey. I also thank  \nDr. Andrew Hamilton-Wright for serving on my committee and for offering invaluable feedback and perspective that strengthened this thesis.  \nTo everyone named above , and to the many others who lent a kind word or a listening ear, thank you.  \nDeclaration of Work  \nThis study includes contributions from collaborators whose efforts were essential to the completion of the research:  \nNicole Garnier performed all biofilm assays used in this study. This included obtaining raw absorbance measurements (A590 nm crystal violet signal normalized by OD600 nm) after 24 hours of growth in Lysogeny Broth (LB) medium. These quantitative data were subsequently binned into four equal-frequency categories and served as the basis for the analysis presented in this work.  \nAlyssa Banaag conducted the Kirby-Bauer disk diffusion assays to evaluate the susceptibility of Pseudomonas aeruginosa to tobramycin. The results from these assays were used directly in a comparative analysis and informed several key conclusions within this thesis.  \nTheir contributions are gratefully acknowledged, and their work was instrumental to the successful execution of this research.  \nTable of Contents  \n","cbCaibNL1CTJQAJ0","https://ap.wps.com/l/cbCaibNL1CTJQAJ0","pdf",1033733,1,56,"English","en",105,"# Abstract\n# Acknowledgements\n# Declaration of Work\n# Table of Contents\n# List of Tables\n# List of Figures\n# List of Symbols, Abbreviations or Nomenclature\n# 1 Introduction\n## 1.1 Challenges Posed by Biofilms\n## 1.2 Statement of Research\n# 2 Methodology and Results\n## 2.1 Dataset Acquisition\n## 2.2 Data Preprocessing\n### 2.2.1 Biofilm Formation Categories\n### 2.2.2 K-mer Frequency Extraction\n# 3 Model Development and Results","[{\"question\":\"What biological problem does the thesis address?\",\"answer\":\"The thesis studies predicting biofilm formation in Pseudomonas aeruginosa using machine learning based on whole genome sequence information.\"},{\"question\":\"How were genome data converted into model inputs?\",\"answer\":\"K-mer frequency profiles were extracted from 870 genomes and used as features for classification, regression, and clustering models.\"},{\"question\":\"Why did the models fail to achieve strong predictive accuracy?\",\"answer\":\"Predictive performance stayed poor across methods, suggesting the dataset volume was insufficient for robust learning on small, high-dimensional biological data.\"}]","Predicting Biofilm Formation in Pseudomonas aeruginosa Using Machine Learning | PDF",1785684807,141,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"predicting-biofilm-formation-in-pseudomonas-aeruginosa-using-machine-learning","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/predicting-biofilm-formation-in-pseudomonas-aeruginosa-using-machine-learning/118667/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What biological problem does the thesis address?","Question",{"text":75,"@type":76},"The thesis studies predicting biofilm formation in Pseudomonas aeruginosa using machine learning based on whole genome sequence information.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How were genome data converted into model inputs?",{"text":80,"@type":76},"K-mer frequency profiles were extracted from 870 genomes and used as features for classification, regression, and clustering models.",{"name":82,"@type":73,"acceptedAnswer":83},"Why did the models fail to achieve strong predictive accuracy?",{"text":84,"@type":76},"Predictive performance stayed poor across methods, suggesting the dataset volume was insufficient for robust learning on small, high-dimensional biological data.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]