[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-120501-en":3,"doc-seo-120501-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},120501,13056703019662,"Evangeline","https://ap-avatar.wpscdn.com/avatar/be000253a8e92610077?_k=1778726343310543188",8,"Research & Report","Enhancing Plagiarism Detection Using Data Pre-processing and Machine Learning Approach","Modern technology and the internet have increased access to academic information, but they have also amplified concerns about plagiarism. This research investigates machine learning for plagiarism detection, emphasizing that robust data preprocessing is essential for strong model performance. Using 67 research papers and plagiarism rates with OCEAN-related factors, the study trains algorithms on an 80% subset and tests on 20% to evaluate generalization. The method integrates outlier detection, normalization, missing value imputation, and feature selection to improve precision and efficiency.","Enhancing plagiarism detection using data pre-processing and  \nmachine learning approach  \nVrushali Bhuyar1, Sachin N. Deshmukh2  \n1Department of Computer Applications, Maharashtra Institute of Technology, Dr. Babasaheb Ambedkar Marathwada University,  \nChh. Sambhajinagar, India  \n2Department of Computer Science and IT, Dr. Babasaheb Ambedkar Marathwada University, Chh. Sambhajinagar, India  \n\n| Article history:\u003Cbr>Received Feb 2, 2024 Revised Nov 19, 2024 Accepted Jan 27, 2025 | Modern technology and the internet have enhanced academic information accessibility, but this has led to a rising global concern about plagiarism. Researchers are actively exploring machine learning as a promising solution for detection. This study underscores the importance of robust data preprocessing for optimal machine learning algorithm performance. Using adataset of 67 research papers, big five factors (OCEAN), and plagiarism rates, the study employed machine learning to detect plagiarism. The training process involved exposing algorithms to an 80% training subset, followed by evaluating their performance on the remaining 20% in the testing phase, assessing generalization capabilities. For the random forest regressor, bagging regressor, gradient boosting regressor, XGB regressor, and AdaBoost regressor, corresponding root mean squared error (RMSE) are 9.48, 10.66, 11.79, 12.53, and 12.79, respectively. This research contributes novel insights to existing literature by introducing a plagiarism detection model that innovatively integrates outlier detection, normalization, missing value imputation, and feature selection. The unique aspect lies in the effective combination of feature selection and missing value imputation, surpassing previous benchmarks and optimizing precision and efficiency. The approach is metaphorically likened to assembling puzzle pieces, highlighting the distinctive methodology employed in enhancing the performance of the plagiarism detection model using data preprocessing.\u003Cbr>This is an open access article under the CC BY-SA license.\u003Cbr> |\n| --- | --- |\n| Keywords:\u003Cbr>Data preprocessing Machine learning Missing value imputation Plagiarism detection Regression |  |\n\nCorresponding Author:  \nVrushali Bhuyar  \nDepartment of Computer Applications, Maharashtra Institute of Technology Dr. Babasaheb Ambedkar Marathwada University  \nChh. Sambhajinagar, India  \n[Email: vrushali.bhuyar@gmail.com](Email: vrushali.bhuyar@gmail.com)  \nArticle Info ABSTRACT  \n1. INTRODUCTION  \nIn the ever-evolving landscape of academic and digital content, the issue of plagiarism remains a significant concern. As the volume of information available online continues to grow exponentially, traditional methods of detecting and preventing plagiarism are proving to be insufficient. This research aims to address this challenge by proposing an innovative approach that combines advanced data preprocessing techniques with state-of-the-art machine learning algorithms to enhance plagiarism detection.  \nTo extract relevant features from the raw textual data, our approach begins with thorough data preprocessing. To do this, the text must be transformed using methods like missing value imputation, removed unwanted attributes, outlier detection, and normalization into a format that is consistent and appropriate for  \nmachine learning analysis. This thorough preprocessing creates the foundation for machine learning models that are more precise and complex.  \nThe use of machine learning algorithms that can identify patterns and abnormalities in thepreprocessed data is the second essential element of our methodology. We train the model on labelled datasets using supervised learning approaches, including regression techniques, so that it can identify patterns linked to plagiarism. Our method seeks to greatly improve the effectiveness of plagiarism detection systems by merging the best features of machine learning and data preprocessing, thereby enhancing the integrity of acad","cbCaigsmNFnhNmkf","https://ap.wps.com/l/cbCaigsmNFnhNmkf","pdf",608188,1,11,"English","en",105,"# Abstract\n# Introduction\n## Background and problem of plagiarism\n## Proposed approach: data preprocessing + machine learning","[{\"question\":\"What problem does this study address in academic work?\",\"answer\":\"The study addresses the rising global concern about plagiarism as online access to academic and digital content grows rapidly.\"},{\"question\":\"How does the proposed method improve plagiarism detection performance?\",\"answer\":\"It combines thorough data preprocessing—such as normalization, outlier detection, missing value imputation, and feature selection—with supervised machine learning regression models.\"},{\"question\":\"How were the machine learning models evaluated?\",\"answer\":\"Models were trained on an 80% subset and evaluated on the remaining 20% test set to assess generalization performance using RMSE.\"}]","Enhancing Plagiarism Detection Using Data Pre-processing and Machine Learning Approach | PDF",1785730381,28,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"enhancing-plagiarism-detection-using-data-pre-processing-and-machine-learning-approach","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/enhancing-plagiarism-detection-using-data-pre-processing-and-machine-learning-approach/120501/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does this study address in academic work?","Question",{"text":75,"@type":76},"The study addresses the rising global concern about plagiarism as online access to academic and digital content grows rapidly.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed method improve plagiarism detection performance?",{"text":80,"@type":76},"It combines thorough data preprocessing—such as normalization, outlier detection, missing value imputation, and feature selection—with supervised machine learning regression models.",{"name":82,"@type":73,"acceptedAnswer":83},"How were the machine learning models evaluated?",{"text":84,"@type":76},"Models were trained on an 80% subset and evaluated on the remaining 20% test set to assess generalization performance using RMSE.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]