[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-124803-en":3,"doc-seo-124803-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},124803,4810365810221,"Aurora","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Revisiting Machine Learning based Test Case Prioritization for Continuous Integration - Comprehensive Study","To alleviate the cost of regression testing in continuous integration (CI), many machine learning–based test case prioritization techniques have been proposed, yet their performance under a consistent experimental setup remains unclear due to differing datasets and metrics. This paper conducts a first comprehensive study of 11 representative ML-based techniques across 11 open-source subjects. Results show performance shifts across CI cycles, driven mainly by changing training data. The rectified APFD (rAPFD) improves fair comparison. Pretraining with cross-subject data plus finetuning on within-subject data significantly boosts effectiveness, with pretrained MART achieving state-of-the-art results.","Revisiting Machine Learning based Test Case Prioritization for Continuous Integration  \n1st Yifan Zhao Key Laboratory of High Confidence Software Technologies (Peking University), Ministry of Education Beijing, China [zhaoyifan@stu.pku.edu.cn](zhaoyifan@stu.pku.edu.cn)  \n2nd Dan Hao∗  \nKey Laboratory of High Confidence Software Technologies (Peking University), Ministry of Education Beijing, China [haodan@pku.edu.cn](haodan@pku.edu.cn)  \n3rd Lu Zhang Key Laboratory of High Confidence Software Technologies (Peking University), Ministry of Education Beijing, China [zhanglucs@pku.edu.cn](zhanglucs@pku.edu.cn)  \narXiv :2311 . 13413v1 [ cs . SE] 22 Nov 2023  \nAbstract—To alleviate the cost of regression testing in continuous integration (CI), a large number of machine learningbased (ML-based) test case prioritization techniques have been proposed. However, it is yet unknown how they perform under the same experimental setup, because they are evaluated on different datasets with different metrics. To bridge this gap, we conduct the first comprehensive study on these ML-based techniques in this paper. We investigate the performance of 11 representative MLbased prioritization techniques for CI on 11 open-source subjects and obtain a series of findings. For example, the performance of the techniques changes across CI cycles, mainly resulting from the changing amount of training data, instead of code evolution and test removal/addition. Based on the findings, we give some actionable suggestions on enhancing the effectiveness of ML-based techniques, e.g., pretraining a prioritization technique with cross-subject data to get it thoroughly trained and then finetuning it with within-subject data dramatically improves its performance. In particular, the pretrained MART achieves stateof-the-art performance, producing the optimal sequence on 80% subjects, while the existing best technique, the original MART, only produces the optimal sequence on 50% subjects.  \nIndex Terms—test prioritization, machine learning, continuous integration  \nI. INTRODUCTION  \nContinuous Integration (CI) is a software development practice. It is widely used in industry, making software release more rapid and reliable [1] . In the CI environment, developers continuously commit code to modify software functionalities.  \nTo guarantee the quality of committed code, regression testing [2] tends to be conducted in each CI cycle, which is widely recognized to be time-consuming [3] and may hamper rapid software release. Many test case prioritization (TCP) techniques have been proposed to alleviate the cost of regression testing. Still, most of them target the general regression testing process. They cannot be directly applied to regression testing in CI, because a large amount of modification on code and tests occurs during frequent CI, which most existing TCP techniques cannot appropriately deal with [4], [5] .  \nTo address the specific problem of TCP in CI, heuristicbased and machine learning-based (ML-based) techniques are proposed. Some researchers give the first attempt by scheduling tests based on heuristics (e.g., time since the last  \n∗Dan Hao is the corresponding author.  \ntest failed [6] or test diversity [7]) . Other researchers harness the power of machine learning by using a large amount of historical data in CI, and propose numerous ML-based TCP techniques which have been demonstrated to be promising. These ML-based techniques build neural models to predict the optimal sequence of tests instead of human-defined strategies. In particular, these ML-based techniques can be categorized into supervised learning-based (SL-based) [8]–[10] and reinforcement learning-based (RL-based) techniques [11]–[14] . In training cycles, an SL-based technique trains a classification model based on tests and their labels. In testing cycles, the model is used to predict priority values for tests. That is, once the model completes training, it uses a fixed strategy to prioritize testi","cbCaitXefyp5Rwr7","https://ap.wps.com/l/cbCaitXefyp5Rwr7","pdf",499692,1,13,"English","en",105,"# Introduction\n## Continuous Integration and Regression Testing\n## Test Case Prioritization in CI\n## Heuristic vs. ML-based TCP\n## Research Gap and Contributions\n# Experimental Study Overview\n## Scope: 11 Open-Source Subjects\n## Evaluation Metrics and Fair Comparison\n## Proposed Rectified APFD (rAPFD)\n# Findings in ML-based TCP for CI\n## Data Imbalance Effects\n## Performance Changes Across CI Cycles","[{\"question\":\"Why is test case prioritization specifically hard in continuous integration (CI)?\",\"answer\":\"CI involves frequent code and test modifications, so many existing TCP techniques designed for general regression testing cannot cope well with CI dynamics.\"},{\"question\":\"What problem do the existing evaluation metrics face in this work?\",\"answer\":\"The paper argues that commonly used metrics used in prior studies create issues for fair comparison and discernment, motivating a new metric design.\"},{\"question\":\"How does the proposed approach improve ML-based test prioritization performance?\",\"answer\":\"It uses rectified APFD for evaluation and shows that pretraining on cross-subject data followed by finetuning with within-subject data substantially improves performance; pretrained MART achieves the strongest results.\"}]","Revisiting Machine Learning based Test Case Prioritization for Continuous Integration - Comprehensive Study | PDF",1785894744,33,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"revisiting-machine-learning-based-test-case-prioritization-for-continuous-integration-comprehensive-study","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/revisiting-machine-learning-based-test-case-prioritization-for-continuous-integration-comprehensive-study/124803/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is test case prioritization specifically hard in continuous integration (CI)?","Question",{"text":75,"@type":76},"CI involves frequent code and test modifications, so many existing TCP techniques designed for general regression testing cannot cope well with CI dynamics.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What problem do the existing evaluation metrics face in this work?",{"text":80,"@type":76},"The paper argues that commonly used metrics used in prior studies create issues for fair comparison and discernment, motivating a new metric design.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the proposed approach improve ML-based test prioritization performance?",{"text":84,"@type":76},"It uses rectified APFD for evaluation and shows that pretraining on cross-subject data followed by finetuning with within-subject data substantially improves performance; pretrained MART achieves the strongest results.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]