[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118467-en":3,"doc-seo-118467-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},118467,2336464648746,"Skyler","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Machine Learning Approaches to Code Similarity Measurement - A Systematic Review","Source code similarity measurement quantifies differences between code fragments and underpins critical software engineering activities, including code quality assurance, code review, code plagiarism detection, and security/vulnerability analysis. Although machine learning is increasingly applied in this area, a unified synthesis of existing methods is missing. This paper conducts a systematic review following established SLR protocols and analyzes 84 primary studies across application types, algorithms, code representations, datasets, and performance metrics. Results show 15 application areas using 51 ML algorithms, with AST as the dominant representation and BigCloneBench as the most common dataset, while limitations and future directions are highlighted.","Date of publication xxxx 00, 0000, date of current version xxxx 00, 0000.  \nDigital Object Identifier 10.1109/ACCESS.2024.xxxxxx  \nMachine Learning Approaches to Code Similarity Measurement: A Systematic Review  \nZIXIAN ZHANG 1, TAKFARINAS SABER2  \n1CRT-AI, School of Computer Science, University of Galway, Ireland,(Z.Zhang15@universityofgalway.ie)  \n2Lero, School of Computer Science, University of Galway, Ireland,([takfarinas.saber@universityofgalway.ie](takfarinas.saber@universityofgalway.ie))  \nCorresponding author: Zixian Zhang.  \n ABSTRACT Source code similarity measurement, which involves assessing the degree of difference between code segments, plays a crucial role in various aspects of the software development cycle. These include but are not limited to code quality assurance, code review processes, code plagiarism detection, security, and vulnerability analysis. Despite the increasing application of ML technique in this domain, a comprehensive synthesis of existing methodologies remains lacking. This paper presents a systematic review of Machine Learning techniques applied to code similarity measurement, aiming to illuminate current methodologies and contribute valuable insights to the research community. Following a rigorous systematic review protocol, we identified and analyzed 84 primary studies on a broad spectrum of dimensions covering application type, devised Machine Learning algorithms, used code representations, datasets, and performance metrics, as well as performance evaluations. A deep investigation reveals that 15 applications for code similarity measurement have utilized 51 different machine learning algorithms. Additionally, the most prevalent code representation is found to be the abstract syntax tree (AST) . Furthermore, the most frequently employed dataset across various code similarity research applications is BigCloneBench. Through this comprehensive analysis, the paper not only synthesizes existing research but also identifies prevailing limitations and challenges, shedding light on potential avenues for future work.  \n INDEX TERMS Code similarity, Code clone, Machine learning, Deep learning, Systematic literature review  \nI. INTRODUCTION  \nAs a cornerstone for many code-centric tasks within the software engineering field, source code similarity measurement has captivated the attention of many researchers. Several tasks encompassing a wide array of applications including, but not limited to, code clone detection [1]–[4], plagiarism detection [5], [6], code recommendation [7], [8], code prediction [9],[10], and code generation [11], [12] rely on code similarity measurement at their core. This burgeoning interest is largely driven by the critical role that effective similarity detection plays in enhancing code quality, fostering innovation, and maintaining academic integrity within the programming community. Accurate and efficient code similarity analysis not only aids in identifying duplicate or near-duplicate code fragments—thereby helping in the reduction of technical debt—but also supports the enforcement of copyright laws and academic standards by detecting instances of plagiarism [13], [14] . Consequently, the exploration of this area represents a critical direction in software engineering research.  \nIn recent years, the rapid advancement of Machine Learn-  \ning (ML) techniques, notably in the realms of natural language processing [15],[16], recommender systems [17],[18], and autonomous vehicles [19], [20], has sparked significant interest across various sectors (e.g. [14], [21]–[24]) . This trend has also positively impacted the field of code similarity measurement by bringing a large and increasing amount of studies that explore the application of ML methodologies to enhance the analysis and identification of similar code fragments.  \nDespite the considerable interest in code similarity, while several studies have concluded the application of ML technique in code clone, a comprehensive syste","cbCaia27wTVGUbEd","https://ap.wps.com/l/cbCaia27wTVGUbEd","pdf",2174090,1,37,"English","en",105,"# Introduction\n## Motivation and related applications\n## Machine learning impact and research gap\n## Review scope and methodology\n## Contributions\n# Overview of Findings (from abstract)\n## Application types and algorithm diversity\n## Code representations and datasets","[{\"question\":\"What is the main focus of the systematic review?\",\"answer\":\"The review synthesizes machine learning techniques used for code similarity measurement, covering applications, algorithms, representations, datasets, and evaluation metrics.\"},{\"question\":\"How many primary studies are analyzed in the review?\",\"answer\":\"The review identifies and analyzes 84 primary studies.\"},{\"question\":\"Which code representation and dataset appear most frequently?\",\"answer\":\"Abstract syntax tree (AST) is the most prevalent code representation, and BigCloneBench is the most frequently used dataset across the included studies.\"}]","Machine Learning Approaches to Code Similarity Measurement - A Systematic Review | PDF",1785683753,93,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"machine-learning-approaches-to-code-similarity-measurement-a-systematic-review","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/machine-learning-approaches-to-code-similarity-measurement-a-systematic-review/118467/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the main focus of the systematic review?","Question",{"text":75,"@type":76},"The review synthesizes machine learning techniques used for code similarity measurement, covering applications, algorithms, representations, datasets, and evaluation metrics.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How many primary studies are analyzed in the review?",{"text":80,"@type":76},"The review identifies and analyzes 84 primary studies.",{"name":82,"@type":73,"acceptedAnswer":83},"Which code representation and dataset appear most frequently?",{"text":84,"@type":76},"Abstract syntax tree (AST) is the most prevalent code representation, and BigCloneBench is the most frequently used dataset across the included studies.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]