[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81514-en":3,"doc-seo-81514-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81514,1099513958762,"Logic","https://ap-avatar.wpscdn.com/avatar/1000023916a998db790?x-image-process=image/resize,m_fixed,w_180,h_180&k=1784791008015729253",8,"Research & Report","Harnessing Large Language Models for Curated Code Reviews","Code review depends on structured, relevant review comments to surface defects and support accurate downstream code changes. Existing AI methods for automating comment generation are constrained by limited training data quality, since public datasets contain noise and lack refinement. The work presents an LLM-driven curation pipeline that improves the largest available code review dataset via an evaluation framework and category-based quality criteria, yielding clearer, more concise comments and better performance for comment generation and code refinement.","Harnessing Large Language Models for Curated  \nCode Reviews  \nOussama Ben Sghaier  \nUniversite´ de Montre´al Montral, Canada [oussama.ben.sghaier@umontreal.ca](oussama.ben.sghaier@umontreal.ca)  \nMartin Weyssow  \nSingapore Management University Singapore [mweyssow@smu.edu.sg](mweyssow@smu.edu.sg)  \nHouari Sahraoui  \nUniversite´ de Montre´al Montral, Canada [sahraouh@iro.umontreal.ca](sahraouh@iro.umontreal.ca)  \narXiv :2502 .03425v 1 [ cs . SE] 5 Feb 2025  \nAbstract—In code review, generating structured and relevant comments is crucial for identifying code issues and facilitating accurate code changes that ensure an efficient code review process. Well-crafted comments not only streamline the code review itself but are also essential for subsequent tasks like code refinement, where the code is modified to satisfy the input review comment. Although various AI-based approaches aimed to automate comment generation, their effectiveness remains limited by the quality of the training data. Existing code review datasets are often noisy and unrefined, posing limitations to the learning potential of AI models and hindering the automation process.  \nTo address these challenges, we propose a curation pipeline designed to enhance the quality of the largest publicly available code review dataset. We begin by establishing an evaluation framework, incorporating specific criteria and categories to empirically study the initial quality of the dataset. Using a large language model (LLM)-driven approach, we then apply our curation pipeline to refine the dataset. A comparative analysis of the newly curated dataset, based on the same evaluation framework, demonstrates substantial improvements in the clarity and conciseness of the comments. Additionally, we assess the impact of the curated dataset on automating downstream tasks, specifically comment generation and code refinement. Our findings show that the curated dataset leads to enhanced model performance in generating more accurate comments. Curated comments are also more useful as they lead to more accurate code refinement.  \nIndex Terms—Code review, large language models, software maintenance, empirical software engineering.  \nI. INTRODUTION  \nCode review is a critical component of the software development life cycle, aimed at identifying issues, detecting suboptimal code, and uncovering bugs [1], [2], while ensuring the overall quality and maintainability of the source code [3]–[5] . This process typically involves a manual inspection of code by one or more developers, reviewing code written by their peers [6],[7] . The code review process consists of several key tasks, with the most essential being the identification and documentation of potential issues through review comments, the subsequent code refinement to resolve these concerns, and the quality assessment of the submitted code to decide if it should be accepted or needs further review.  \nIssue identification and description (i.e., review comment generation) constitutes a foundational task in code review, focusing on the detection of defects or problems within the  \nThe data is available at [https://zenodo.org/records/14812107. The](https://zenodo.org/records/14812107. The) replication package is available at [https://github.com/OussamaSghaier/CuREV](https://github.com/OussamaSghaier/CuREV).  \ncode. This phase is pivotal, as it involves not only identifying specific issues but also offers potential solutions for resolving them [1], [2] . The significance of this task cannot be overstated, as subsequent stages of the review process heavily depend on its accuracy and thoroughness. Code refinement, for example, is an execution phase directly tied to insights gained from issue identification; developers revise and improve the code based on comments provided during this stage [8] . Thus, comment generation is essential to the entire code review process. Without precise execution of this task, the integrity and quality of later stages is com","cbCaik6omrSLDemV","https://ap.wps.com/l/cbCaik6omrSLDemV","pdf",522101,3,1,12,"English","en",105,"# Introduction\n## Code review as a core software engineering task\n## Challenges in automated comment generation\n## Importance of training data quality\n# Proposed Curation Pipeline\n## Evaluation framework and dataset quality criteria\n## LLM-driven dataset refinement\n## Impact on downstream tasks","[{\"question\":\"Why is generating structured review comments important in code review?\",\"answer\":\"Structured, relevant comments identify code issues and enable accurate subsequent code refinement. They also streamline the review process and improve later stages that modify code to address feedback.\"},{\"question\":\"What problem limits current AI-based approaches to automated code review comment generation?\",\"answer\":\"Their effectiveness is limited by training data quality. Existing datasets are often noisy and unrefined, which restricts learning and can lead models to produce irrelevant or incoherent feedback.\"},{\"question\":\"How does the proposed approach improve code review datasets and downstream tasks?\",\"answer\":\"It introduces an evaluation framework and an LLM-driven curation pipeline to refine the dataset using explicit quality criteria. The curated dataset improves comment clarity and conciseness and enhances model performance for generating accurate comments and performing code refinement.\"}]",1784173929,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"harnessing-large-language-models-for-curated-code-reviews","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/harnessing-large-language-models-for-curated-code-reviews/81514/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is generating structured review comments important in code review?","Question",{"text":75,"@type":76},"Structured, relevant comments identify code issues and enable accurate subsequent code refinement. They also streamline the review process and improve later stages that modify code to address feedback.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What problem limits current AI-based approaches to automated code review comment generation?",{"text":80,"@type":76},"Their effectiveness is limited by training data quality. Existing datasets are often noisy and unrefined, which restricts learning and can lead models to produce irrelevant or incoherent feedback.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the proposed approach improve code review datasets and downstream tasks?",{"text":84,"@type":76},"It introduces an evaluation framework and an LLM-driven curation pipeline to refine the dataset using explicit quality criteria. The curated dataset improves comment clarity and conciseness and enhances model performance for generating accurate comments and performing code refinement.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]