[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82385-en":3,"doc-seo-82385-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82385,1099514068365,"Aurelia","https://ap-avatar.wpscdn.com/avatar/10000253d8d9f28188e?_k=1776742907772140068",8,"Research & Report","Balancing Usefulness and Naturalness An LLM-based Curation Pipeline for Code Review Comments","Code review is central to software development, relying on written feedback to improve code quality, maintainability, and correctness. The promise of LLM-based automated review is constrained by the training data: existing datasets are often noisy, inconsistent, and poorly structured, reducing the accuracy and usefulness of generated comments. This work proposes two LLM curation pipelines that reformulate and enhance datasets, including CuREV and CuREV+, and evaluates them across multiple quality dimensions and downstream automation impact.","arXiv :2607 .09524v 1 [ cs . SE] 10 Jul 2026  \nEmpirical Software Engineering manuscript No.  \n(will be inserted by the editor)  \nBalancing Usefulness and Naturalness: An LLM-based Curation Pipeline for Code Review Comments  \nOussama Ben Sghaier · Martin Weyssow · Houari Sahraoui  \nAccepted: 6 July 2026  \nAbstract Code review is a cornerstone of software development, where reviewers provide feedback through written comments to ensure code quality, maintainability, and correctness. The effectiveness of this process hinges on the quality of review comments: they must be clear, concise, actionable, and realistic enough to guide accurate code changes. As large language models (LLMs) gain traction in automating code review tasks, the utility of these systems is directly limited by the quality of the datasets on which they are trained. Unfortunately, existing code review datasets are often noisy, inconsistent, or poorly structured, which hinders the ability of LLMs to learn to generate accurate, helpful, and human-like review comments.  \nTo overcome these limitations, we propose two different curation pipelines designed to improve both the quality and the utility of large-scale code review datasets. In the first pipeline, all review comments are systematically reformulated by an LLM to improve their clarity, conciseness, and civility while preserving their semantic intent. The curated dataset resulting from this approach, called CuREV, offers cleaner, higher-quality, and easier-to-learn-from comments that lead to measurable improvements in downstream automation tasks, namely review comment generation and code refinement. Building on this, we propose an improved pipeline, guided by high-quality exemplars, that enhances the realism and diversity of curated review comments. This method first separates the dataset into high-quality (“good”) and low-quality (“poor”) reviews, based on a systematic quality assessment using an evaluation framework. High-quality comments are preserved in their original form and further used as in-context exemplars to inspire the reformulation of low-quality comments. By varying the exemplars provided, the reformulated comments are not only clearer and more actionable but also exhibit a broader range of writing styles, making them more realistic and human-like. The resulting dataset, called CuREV+, thus combines improved quality and utility with enhanced diversity of review comments.  \nWe evaluate both curated datasets using a comprehensive evaluation framework that assesses review comments along multiple quality dimensions (e.g., clarity, conciseness, civility, nature) as well as their impact on downstream code review tasks compared to the original dataset. In addition, we analyze the diversity of the curated datasets, CuREV and CuREV+ . Our results show that while both approaches significantly enhance review comment quality and improve the performance of automated code review tasks, CuREV+ provides a more diverse dataset, enriched with broader vocabulary, varied lexical choices, and distinct writing styles. These findings demonstrate that curating datasets for code review requires not only refining quality but also balancing standardization with diversity.  \nKeywords Code review · dataset curation · large language models · comment generation · software maintenance · diversity  \nO. Ben Sghaier  \nSchool of Computing, Queen’s University, Kingston, ON, Canada E-mail: [oussama.sghaier@queensu.ca](oussama.sghaier@queensu.ca)  \n[M. Weyssow](M. Weyssow)  \nSingapore Management University, Singapore [E-mail: mweyssow@smu.edu.sg](E-mail: mweyssow@smu.edu.sg)  \nH. Sahraoui  \nDIRO, Universit´e de Montr´eal, Montr´eal, QC, Canada E-mail: [sahraouh@iro.umontreal.ca](sahraouh@iro.umontreal.ca)  \n1 Introduction  \nCode review is one of the most widely adopted practices in software development. It enables teams to detect defects early, improve maintainability, and transfer knowledge between developers (McIntosh et al. , 2014,","cbCaijl1vjKXJkdL","https://ap.wps.com/l/cbCaijl1vjKXJkdL","pdf",1214828,5,1,28,"English","en",105,"# Abstract\n# Introduction\n## Problem: dataset noise and inconsistency\n## Motivation: LLM-based comment automation limits\n# Proposed approach\n## CuREV: LLM reformulation with intent preservation\n## CuREV+: exemplar-guided separation and reformulation\n# Evaluation\n## Quality dimensions and downstream task impact\n## Diversity analysis","[{\"question\":\"Why is curating code review comment datasets necessary for LLM-based automation?\",\"answer\":\"Because publicly available datasets are often mined without curation and contain noise such as irrelevant, uncivil, vague, or unstructured comments. This noise limits how accurately and helpfully LLMs can generate review comments.\"},{\"question\":\"How does the CuREV pipeline improve code review comments?\",\"answer\":\"CuREV reformulates all review comments using an LLM to enhance clarity, conciseness, and civility while preserving the original semantic intent. The curated output is designed to be cleaner and easier for downstream learning and automation.\"},{\"question\":\"What distinguishes CuREV+ from CuREV in achieving naturalness and diversity?\",\"answer\":\"CuREV+ uses an evaluation framework to split reviews into high-quality and low-quality sets, preserves high-quality comments as exemplars, and reformulates low-quality comments with varying in-context exemplars. This increases realism through broader writing styles and vocabulary diversity.\"}]",1784180066,71,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"balancing-usefulness-and-naturalness-an-llm-based-curation-pipeline-for-code-review-comments","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/balancing-usefulness-and-naturalness-an-llm-based-curation-pipeline-for-code-review-comments/82385/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-22","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why is curating code review comment datasets necessary for LLM-based automation?","Question",{"text":76,"@type":77},"Because publicly available datasets are often mined without curation and contain noise such as irrelevant, uncivil, vague, or unstructured comments. This noise limits how accurately and helpfully LLMs can generate review comments.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does the CuREV pipeline improve code review comments?",{"text":81,"@type":77},"CuREV reformulates all review comments using an LLM to enhance clarity, conciseness, and civility while preserving the original semantic intent. The curated output is designed to be cleaner and easier for downstream learning and automation.",{"name":83,"@type":74,"acceptedAnswer":84},"What distinguishes CuREV+ from CuREV in achieving naturalness and diversity?",{"text":85,"@type":77},"CuREV+ uses an evaluation framework to split reviews into high-quality and low-quality sets, preserves high-quality comments as exemplars, and reformulates low-quality comments with varying in-context exemplars. This increases realism through broader writing styles and vocabulary diversity.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]