[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-117799-en":3,"doc-seo-117799-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},117799,687197100911,"Himbo","https://ap-avatar.wpscdn.com/avatar/a000239b6f1da00475?x-image-process=image/resize,m_fixed,w_180,h_180&k=1785132997149421697",8,"Research & Report","Empirically evaluating flaky test detection techniques combining test case rerunning and machine learning models - Research paper","A flaky test is a test case whose outcome can change without any modification to the test case code or the program under test, undermining continuous integration and wasting developer time while reducing testing efficiency. Existing detection methods either rerun tests at high time cost or use machine learning for faster yet approximate detection with variable performance. This work presents CANNIER to cut rerunning-based time cost by combining it with machine learning, and an evaluation on 89,668 test cases from 30 Python projects shows an order-of-magnitude reduction versus baseline rerunning methods while keeping significantly better detection than machine learning alone.","This is a repository copy of Empirically evaluating flaky test detection techniques combining test case rerunning and machine learning models.  \nWhite Rose Research Online URL for this paper:  \n[https://eprints.whiterose.ac.uk/198846/](https://eprints.whiterose.ac.uk/198846/)  \nVersion: Published Version  \nArticle:  \nParry, [O. orcid.org/0000-0002-0917-1274](O. orcid.org/0000-0002-0917-1274) , Kapfhammer, G.M., Hilton, M. et al. (1 more author) (2023) Empirically evaluating flaky test detection techniques combining test case rerunning and machine learning models. Empirical Software Engineering, 28. 72. ISSN 1382-3256  \n[https://doi.org/10.1007/s10664-023-10307-w](https://doi.org/10.1007/s10664-023-10307-w)  \nReuse  \nThis article is distributed under the terms of the Creative Commons Attribution (CC BY) licence. This licence allows you to distribute, remix, tweak, and build upon the work, even commercially, as long as you credit the authors for the original work. More information and the full terms of the licence here: [https://creativecommons.org/licenses/](https://creativecommons.org/licenses/)  \nTakedown  \nIf you consider content in White Rose Research Online to be in breach of UK law, please notify us by  \nemailing [eprints@whiterose.ac.uk](eprints@whiterose.ac.uk) including the URL of the record and the reason for the withdrawal request.  \n[eprints@whiterose.ac.uk](eprints@whiterose.ac.uk)[ ](eprints@whiterose.ac.uk)[https://eprints.whiterose.ac.uk/](https://eprints.whiterose.ac.uk/)  \nEmpirically evaluating ﬂaky test detection techniques combining test case rerunning and machine learning models  \nOwain Parry1  · Gregory M. Kapfhammer2 · Michael Hilton3 · Phil McMinn1  \nAccepted: 9 February 2023 © The Author(s) 2023  \nAbstract  \nA flaky test is a test case whose outcome changes without modification to the code of the test case or the program under test. These tests disrupt continuous integration, cause a loss of developer productivity, and limit the efficiency of testing. Many flaky test detection techniques are rerunning-based, meaning they require repeated test case executions ata considerable time cost, or are machine learning-based, and thus they are fast but offer only an approximate solution with variable detection performance. These two extremes leave developers with a stark choice. This paper introduces CANNIER, an approach for reducing the time cost of rerunning-based detection techniques by combining them with machine learning models. The empirical evaluation involving 89,668 test cases from 30 Python projects demonstrates that CANNIER can reduce the time cost of existing rerunningbased techniques by an order of magnitude while maintaining a detection performance that is significantly better than machine learning models alone. Furthermore, the comprehensive study extends existing work on machine learning-based detection and reveals a number of additional findings, including (1) the performance of machine learning models for detecting polluter test cases; (2) using the mean values of dynamic test case features from repeated measurements can slightly improve the detection performance of machine learning models; and (3) correlations between various test case features and the probability of the test case being flaky.  \nKeywords Software testing · Flaky tests · Machine learning  \n1 Introduction  \nA flaky test is a test case that can exhibit both passing and failing behavior without changes to the code of the test case or the program under test (Parry et al. 2021) . They are a serious problem for software developers because they disrupt continuous integration, cause a loss of  \nCommunicated by: Dietmar Pfahl  \n􀀂 Owain Parry [oparry1@sheffield.ac.uk](oparry1@sheffield.ac.uk)  \nExtended author information available on the last page of the article.  \nproductivity, and limit the efficiency of testing. The pain of flaky tests is felt by developers in both the open-source domain (Durieux et al. 2020) and in large companies s","cbCaivLPfb3cRmzD","https://ap.wps.com/l/cbCaivLPfb3cRmzD","pdf",2655643,1,53,"English","en",105,"# Introduction\n## Flaky tests and their impact\n## Order-dependent flakes (victims and polluters)\n## Existing detection approaches and limitations\n# Proposed approach (CANNIER)\n## Combining rerunning and machine learning\n# Empirical evaluation\n## Dataset and research results\n## Findings on polluters and dynamic features\n## Feature correlations and flakiness probability","[{\"question\":\"What is a flaky test and why is it harmful?\",\"answer\":\"A flaky test is a test case whose outcome can flip between passing and failing without code changes to the test or the system under test. It disrupts continuous integration, reduces developer productivity, and limits testing efficiency.\"},{\"question\":\"How do existing flaky test detection techniques typically work, and what are their trade-offs?\",\"answer\":\"Rerunning-based techniques repeatedly execute test cases, which can be very expensive in time. Machine learning-based techniques are faster but provide approximate detection with varying performance.\"},{\"question\":\"What is CANNIER and what does the evaluation show?\",\"answer\":\"CANNIER reduces the time cost of rerunning-based detection by combining rerunning with machine learning models. The study over 89,668 test cases from 30 Python projects shows about an order-of-magnitude reduction in rerunning time cost while maintaining significantly better detection performance than machine learning alone.\"}]","Empirically evaluating flaky test detection techniques combining test case rerunning and machine learning models - Research paper | PDF",1785679632,134,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"empirically-evaluating-flaky-test-detection-techniques-combining-test-case-rerunning-and-machine-learning-models-research-paper","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/empirically-evaluating-flaky-test-detection-techniques-combining-test-case-rerunning-and-machine-learning-models-research-paper/117799/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is a flaky test and why is it harmful?","Question",{"text":75,"@type":76},"A flaky test is a test case whose outcome can flip between passing and failing without code changes to the test or the system under test. It disrupts continuous integration, reduces developer productivity, and limits testing efficiency.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How do existing flaky test detection techniques typically work, and what are their trade-offs?",{"text":80,"@type":76},"Rerunning-based techniques repeatedly execute test cases, which can be very expensive in time. Machine learning-based techniques are faster but provide approximate detection with varying performance.",{"name":82,"@type":73,"acceptedAnswer":83},"What is CANNIER and what does the evaluation show?",{"text":84,"@type":76},"CANNIER reduces the time cost of rerunning-based detection by combining rerunning with machine learning models. The study over 89,668 test cases from 30 Python projects shows about an order-of-magnitude reduction in rerunning time cost while maintaining significantly better detection performance than machine learning alone.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]