[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-117091-en":3,"doc-seo-117091-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},117091,4810365810221,"Aurora","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Effective Spam Detection with Machine Learning - Research Article","The study reports empirical experiments comparing the accuracy of machine learning algorithms for detecting spam messages using a public spam dataset. A key focus is combining topic modeling—Latent Dirichlet Allocation (LDA)—with machine learning to extract hidden topics and reveal patterns distinguishing spam from non-spam. Classifiers are evaluated with F-score, showing Logistic Regression leading at 0.986, followed by SVM at 0.98 and Naive Bayes at 0.955. Results support more robust risk control to sustain digital communication ecosystems.","Effective Spam Detection with Machine Learning  \nGordana Borotić, Lara Granoša, Jurica Kovačević, Marina Bagić Babac  \nUniversity of Zagreb, Faculty of Electrical Engineering and Computing  \nAbstract  \nThis paper aims to provide results of empirical experiments on the accuracy of different machine learning algorithms for detecting spam messages, using a public dataset of spam messages. The originality of our study lies in the integration of topic modeling, specifically employing Latent Dirichlet Allocation (LDA) alongside machine learning algorithms for spam detection. By extracting hidden topics and uncovering patterns in spam and non-spam messages, we provide unique insights into the distinguishing characteristics of spam messages. Moreover, the integration of machine learning is a powerful tool in bolstering risk control measures ensuring the sustainability of digital platforms and communication channels. The research tests the accuracy of spam detection classifiers on an open-source dataset of spam messages. The key findings of this study reveal that the Logistic Regression classifier achieved the highest F score of 0.986, followed by the Support Vector Machine classifier with a score of 0.98 and the Naive Bayes classifier with a score of 0.955. The study concludes that Logistic Regression outperforms Naive Bayes and Support Vector Machine in text classification, particularly in spam detection, emphasizing the role of machine learning techniques in optimizing risk management strategies for sustained digital ecosystems. This capability stems from Logistic Regression's adeptness in modeling complex relationships, enabling it to achieve high accuracy on training and test datasets.  \nKeywords: spam, email, naive Bayes, logistic regression, support vector machine, risk, sustainability  \nPaper Type: Research article  \nReceived: 11 Oct 2023  \nAccepted: 28 Dec 2023  \nDOI: 10.2478/crdj-2023-0007  \nIntroduction  \nSpam messages are messages that are unsolicited and unwanted (Cranor & LaMacchia, 1998) . In August of 2022, 10.89 billion spam texts were sent. This significantly increased over eleven months compared to 1.227 million spam messages sent in September 2021 (uSMS[GH.com](GH.com), 2022). Most of these messages are product buying links, which would consume our personal data or could be some links and attachments. Such messages can be frustrating and dangerous simultaneously (Kudupudi and Nair, 2021) . Spam messages are estimated to cost Americans 10 billion dollars in 2021 (Orred, 2023) .  \nRecognizing the urgency of addressing this issue in the context of sustainability, this study delves into the application of machine learning to enhance risk control in spam detection. The exponential rise in spam messages poses a threat to individual privacy and demands innovative solutions to safeguard digital ecosystems, making the integration of machine learning crucial for sustainability.  \nSpam detection is a critical task in the context of digital transformation, where businesses and individuals rely heavily on email and other forms of electronic communication. Traditional rule-based methods have been widely used for spam detection, but they are limited due to the constantly evolving nature of spam messages. The current gap in spam detection with machine learning lies in the need for more robust and adaptive models that can effectively handle emerging spamming techniques and evolving spam patterns. With the increasing availability of large amounts of data and advances in machine learning techniques, Natural Language Processing (NLP)-based methods have emerged as a promising approach for spam detection. Specifically, the model based on generative Latent Dirichlet Allocation (LDA) topic modeling (Li et al., 2013) has successfully discerned subtle differences between deceptive and genuine reviews. This method could be valuable for identifying spam through content analysis and the thematic structure of messages.  \nThis research thus contr","cbCaitUaugOlWGx6","https://ap.wps.com/l/cbCaitUaugOlWGx6","pdf",931886,1,22,"English","en",105,"# Introduction\n## Spam growth and risks\n## Limits of rule-based filtering\n## Role of NLP and LDA topic modeling\n# Literature review\n## Email spam filtering research themes","[{\"question\":\"What is the main goal of the study on spam detection?\",\"answer\":\"To evaluate the accuracy of different machine learning algorithms for detecting spam messages and to understand how LDA topic modeling can improve insights into spam characteristics.\"},{\"question\":\"Which machine learning models achieved the best F-score results?\",\"answer\":\"Logistic Regression achieved the highest F-score (0.986), followed by Support Vector Machine (0.98), and Naive Bayes (0.955).\"},{\"question\":\"How does LDA contribute to spam detection in this research?\",\"answer\":\"LDA extracts hidden topics and uncovers thematic patterns in spam versus non-spam messages, extending analysis beyond traditional classification alone.\"}]","Effective Spam Detection with Machine Learning - Research Article | PDF",1785673709,55,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"effective-spam-detection-with-machine-learning-research-article","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/effective-spam-detection-with-machine-learning-research-article/117091/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the main goal of the study on spam detection?","Question",{"text":75,"@type":76},"To evaluate the accuracy of different machine learning algorithms for detecting spam messages and to understand how LDA topic modeling can improve insights into spam characteristics.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which machine learning models achieved the best F-score results?",{"text":80,"@type":76},"Logistic Regression achieved the highest F-score (0.986), followed by Support Vector Machine (0.98), and Naive Bayes (0.955).",{"name":82,"@type":73,"acceptedAnswer":83},"How does LDA contribute to spam detection in this research?",{"text":84,"@type":76},"LDA extracts hidden topics and uncovers thematic patterns in spam versus non-spam messages, extending analysis beyond traditional classification alone.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]