[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-120662-en":3,"doc-seo-120662-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},120662,1099513958607,"Jiven","https://ap-avatar.wpscdn.com/avatar/100002390cf8733938c?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778829742770036399",8,"Research & Report","FlaKat - A Machine Learning-Based Categorization Framework for Flaky Tests","Flaky tests can exhibit nondeterministic behavior, passing or failing without changes to the software system, reducing the credibility of test suites. Existing work covers defining, locating, categorizing flaky tests, and repairing specific flakiness types, including machine-learning-based detection. A gap remains between detection and category-specific repair. This thesis introduces FlaKat, which parses flaky tests into vector embeddings, applies dimensionality reduction, addresses class imbalance via sampling, and trains classifiers for fast, accurate categorization.","FlaKat: A Machine Learning-Based Categorization Framework for Flaky  \nTests  \nby  \nShizhe Lin  \nA thesis  \npresented to the University of Waterloo  \nin fulfillment of the  \nthesis requirement for the degree of  \nMaster of Applied Science  \nin  \nComputer Engineering  \nWaterloo, Ontario, Canada, 2023  \n© Shizhe Lin 2023  \nAuthor’s Declaration  \nI hereby declare that I am the sole author of this thesis. This is a true copy of the thesis, including any required final revisions, as accepted by my examiners.  \nI understand that my thesis may be made electronically available to the public.  \nAbstract  \nFlaky tests can pass or fail non-deterministically, without alterations to a software system. Such tests are frequently encountered by developers and hinder the credibility of test suites. Thus, flaky tests have caught the attention of researchers in recent years. Numerous approaches have been published on defining, locating, and categorizing flaky tests, along with auto-repairing strategies for specific types of flakiness. Practitioners have developed several techniques to detect flaky tests automatically. The most traditional approaches adopt repeated execution of test suites accompanied by techniques such as shuffled execution order, and random distortion of environment. State-of-the-art research also incorporates machine learning solutions into flaky test detection and achieves reasonably good accuracy. Moreover, strategies for repairing flaky tests have also been published for specific flaky test categories and the process has been automated as well. However, there is a research gap between flaky test detection and category-specific flakiness repair.  \nTo address the aforementioned gap, this thesis proposes a novel categorization framework, called FlaKat, which uses machine-learning classifiers for fast and accurate categorization of a given flaky test case. FlaKat first parses and converts raw flaky tests into vector embeddings. The dimensionality of embeddings is reduced and then used for training machine learning classifiers. Sampling techniques are applied to address the imbalance between flaky test categories in the dataset.  \nThe evaluation of FlaKat was conducted to determine its performance with different combinations of configurations using known flaky tests from 108 open-source Java projects. Notably, Implementation-Dependent and Order-Dependent flaky tests, which represent almost 75% of the total dataset, achieved F1 scores (harmonic mean of precision and recall) of 0.94 and 0.90 respectively while the overall macro average (no weight difference between categories) is at 0.67 .  \nThis research work also proposes a new evaluation metric, called Flakiness Detection Capacity (FDC), for measuring the accuracy of classifiers from the perspective of information theory and provides proof for its effectiveness. The final obtained results for FDC also aligns with F1 score regarding which classifier yields the best flakiness classification.  \nAcknowledgements  \nI would like to first thank my supervisor Dr. Ladan Tahvildari. This thesis will not exist without her help and guidance along the way. Her endorsement and encouragement motivates me to take the path of research and explore my life from a different perspective. It is my pleasure to work with her and I am looking forward to upcoming collaborations.  \nI learnt many valuable lessons from Dr. Mark Crowley for data modeling and Dr. Patrick Lam for static analysis. They also spend great efforts providing feedback for my thesis which I am sincerely grateful. I also want to thank the members of the STAR group, especially Ryan Zheng He Liu, for the discussion and assistance at the early stage of the thesis development.  \nI would also like to thank my parents, Lin Gesheng and Pan Weiping, for their neverchanging support throughout the years. The pandemic makes it difficult to meet in person but I will always appreciate the experience and wisdom they shared through screen.  \nFinally,","cbCaigLdmgEq6CZ2","https://ap.wps.com/l/cbCaigLdmgEq6CZ2","pdf",1860068,1,80,"English","en",105,"# Introduction\n## Research Challenge\n## Thesis Contribution\n## Thesis Organization\n# Background and Related Work\n## Definition and Impact of Flaky Test\n## Cause and Categories of Flaky Test\n## Existing Tools for Flaky Test Detection\n## Repair and Mitigate Flaky Test\n## Summary\n# Framework Architecture\n## FlaKat: Workflow\n## Flakat: Research Goals\n## Summary\n# Framework Realization\n## Java Test Parser\n## Source Code Representation Algorithms\n## Dimensionality Reduction Techniques\n## Oversampling and Undersampling Techniques","[{\"question\":\"What problem does this thesis address?\",\"answer\":\"It addresses flaky tests that pass or fail nondeterministically without any software changes, undermining the reliability of test suites.\"},{\"question\":\"How does FlaKat categorize flaky tests?\",\"answer\":\"FlaKat converts flaky tests into vector embeddings, reduces embedding dimensionality, mitigates category imbalance with sampling techniques, and trains machine learning classifiers for categorization.\"},{\"question\":\"How was FlaKat evaluated and what were key results?\",\"answer\":\"Evaluation used known flaky tests from 108 open-source Java projects and compared classifier performance under different configuration combinations. Implementation-Dependent and Order-Dependent tests achieved strong F1 scores, while the overall macro average was lower due to category differences.\"}]","FlaKat - A Machine Learning-Based Categorization Framework for Flaky Tests | PDF",1785731230,202,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"flakat-a-machine-learning-based-categorization-framework-for-flaky-tests","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/flakat-a-machine-learning-based-categorization-framework-for-flaky-tests/120662/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does this thesis address?","Question",{"text":75,"@type":76},"It addresses flaky tests that pass or fail nondeterministically without any software changes, undermining the reliability of test suites.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does FlaKat categorize flaky tests?",{"text":80,"@type":76},"FlaKat converts flaky tests into vector embeddings, reduces embedding dimensionality, mitigates category imbalance with sampling techniques, and trains machine learning classifiers for categorization.",{"name":82,"@type":73,"acceptedAnswer":83},"How was FlaKat evaluated and what were key results?",{"text":84,"@type":76},"Evaluation used known flaky tests from 108 open-source Java projects and compared classifier performance under different configuration combinations. Implementation-Dependent and Order-Dependent tests achieved strong F1 scores, while the overall macro average was lower due to category differences.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,100,104,109,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":21,"slug":99},"Literature","literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":46,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":105,"slug":137},19,"General","general"]