[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-122362-en":3,"doc-seo-122362-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},122362,1374391974468,"Eden","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Overview of Machine Learning Algorithms for Solving the Spam Detection Problem - paper summary","This research addresses spam filtering by applying machine learning techniques to spam identification. It reviews existing spam detection algorithms and develops a proposed classification framework grounded in underlying mathematical methodology. The work provides detailed mathematical formulations of the selected algorithms and reports empirical results that evaluate the accuracy of their current implementations. It also outlines directions for future research to further improve spam detection performance.","[https://gscjournal.com/IJLDI](https://gscjournal.com/IJLDI)  \nInternational Journal of Learning Development and Innovation  \nVol. 1, No. 2,October 2024, pp. 140–149 || EISSN 3057-0433  \nOverview of Machine Learning Algorithms for Solving the Spam Detection Problem  \nYuldasheva Khurshida  \nReceived: 2023 29, Aug  \nAccepted: 2023 27, Sep  \nPublished: 2024 31, Oct  \nCopyright © 2024 by author(s) and Scientific Research Publishing Inc. This work is licensed under the Creative Commons Attribution International License (CC BY 4.0) .  \n[http://creativecommons.org/licenses/](http://creativecommons.org/licenses/)[ ](http://creativecommons.org/licenses/)[by/4.0/](by/4.0/)  \nOpen Access  \nAnnotation  \nThis research focuses on addressing the challenge of spam filtering through the application of machine learning techniques. This research involved a comprehensive review of spam identification algorithms, resulting in a proposed classification framework. A detailed mathematical formulation of the algorithms is presented, accompanied by empirical results demonstrating the accuracy of their current implementations. Potential avenues for future research have been highlighted to enhance spam detection capabilities.  \nKeywords:  \nmachine learning, spam filtering, spam detection algorithms, classification framework, mathematical formulation, algorithm accuracy, research directions, spam identification, machine learning in spam filtering.  \n1 Introduction  \nSpam refers to unwanted bulk messages, distributed directly or indirectly, to recipients who have not opted in, regardless of measures implemented to curb such distribution [Cormack, 2008] . For decades, the battle against spam has been waged, and while significant progress has been made in developing and researching various solutions, a Kaspersky Lab study [Spam and Phishing in the Second Quarter of 2016] shows that spam continues to constitute a substantial portion of global and domestic email traffic (Fig. 1) .  \nFigure 1. Percentage of spam in email correspondence in 2023.  \nThe statistics presented highlight the necessity of developing novel spam filtering algorithms and enhancing existing ones. The goal of this paper is to examine algorithms for machine learning-based spam identification.  \n2. Algorithm Classification  \nA wide range of algorithms exist for identifying spam. Given the multitude of spam identification algorithms, it is necessary to classify them based on a certain criterion. A proposed classification categorizes these algorithms by their underlying mathematical approach (see Fig. 2) . This classification builds upon and extends the one presented in [Cormack, 2008] .  \n2.1 Notations and Abbreviations  \nLet С = {􀝏􀝌􀜽􀝉, 􀝊􀝋􀝊 − 􀝏􀝌􀜽􀝉} denote the set of classes, and 􀜶 = {􀝐} the set of terms, where |􀜶| = 􀝊 . For training algorithms, we assume a training set of documents D′⊆D, that is, a set of documents whose class is known in advance. Let us define a function class: D→C that correctly maps each document to a class. A document can be represented in various forms (for example, as a multiset of terms or as a vector in a term space) . To account for this, we introduce the concept of a document form d 􀝀 ∈􀜦 􀜴 (􀝀) . The document form and the definition of a term depend on the specific term extraction algorithm from the document [Cormack, 2008]. Based on this terminology,  \nlet us consider existing machine learning-based spam filtering algorithms.\" We will compare the classifiers using two metrics: accuracy, which is the percentage of documents that are correctly classified, and the 1-AUC statistic. The 1-AUC measures the area under the ROC curve, where a lower value indicates higher accuracy (e.g., 0.1% corresponds to an AUC of 0.999) [Hanley, McNeil, 1983] .  \nSpam filtering algorithms  \nProbabilistic-Bayesian classifiers-logistic regression-MRF  \nclassifier  \n-Perceptron  \n-Winnow -algorithm SVM  \nBased on similarity-k-nearest neighbors  \nDecision tree Rule-based inference  \nBased on d","cbCaimcsygwh6eG5","https://ap.wps.com/l/cbCaimcsygwh6eG5","pdf",540830,1,10,"English","en",105,"# Introduction\n# Algorithm Classification\n## Notations and Abbreviations\n## Probabilistic classification","[{\"question\":\"What problem does the document focus on?\",\"answer\":\"It focuses on improving spam filtering and spam identification using machine learning techniques.\"},{\"question\":\"How are spam filtering algorithms organized in the document?\",\"answer\":\"Algorithms are classified according to their underlying mathematical approach, extending prior classification work.\"},{\"question\":\"What evaluation aspects are used to compare classifiers?\",\"answer\":\"The document compares classifiers using accuracy and the 1-AUC statistic derived from the ROC curve.\"}]","Overview of Machine Learning Algorithms for Solving the Spam Detection Problem - paper summary | PDF",1785810244,25,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"overview-of-machine-learning-algorithms-for-solving-the-spam-detection-problem-paper-summary","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/overview-of-machine-learning-algorithms-for-solving-the-spam-detection-problem-paper-summary/122362/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the document focus on?","Question",{"text":75,"@type":76},"It focuses on improving spam filtering and spam identification using machine learning techniques.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How are spam filtering algorithms organized in the document?",{"text":80,"@type":76},"Algorithms are classified according to their underlying mathematical approach, extending prior classification work.",{"name":82,"@type":73,"acceptedAnswer":83},"What evaluation aspects are used to compare classifiers?",{"text":84,"@type":76},"The document compares classifiers using accuracy and the 1-AUC statistic derived from the ROC curve.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":21,"slug":133},"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]