[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-121896-en":3,"doc-seo-121896-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},121896,4398048949847,"Eliana","https://ap-avatar.wpscdn.com/avatar/400002536579ef2da7f?_k=1778318612642679267",8,"Research & Report","Classification of Spam URLs Using Machine Learning Approachs - Abstract","The Internet enables fast and free communication, but the growth in usage also produces massive spam volumes that waste network resources and users’ time. This study evaluates machine learning models for classifying URLs as spam or non-spam. URL-derived features are extracted and several classifiers are compared, including k-nearest neighbors, bagging, random forest, logistic regression, and others. Results show bagging achieves the highest accuracy of 98.64% and outperforms prior state-of-the-art approaches, indicating strong promise for URL spam detection.","Classification of Spam URLs Using Machine  \nLearning Approachs  \nOmar Husni Odeh 1 , Anas Arram2 , and Murad Njoum2  \n1An-Najah National University  \n2Department of Computer Science, Birzeit University  \narXiv :2310 .05953v2 [ cs .CR] 3 Dec 2023  \nAbstract—The Internet is used by billions of users every day because it offers fast and free communication tools and platforms. Nevertheless, with this significant increase in usage, huge amounts of spam are generated every second, which wastes internet resources and, more importantly, users’ time. This study investigates the use of machine learning models to classify URLs as spam or non-spam. We first extract the features from the URL as it has only one feature, and then we compare the performance of several models, including k-nearest neighbors, bagging, random forest, logistic regression, and others. Experimental results demonstrate that bagging outperformed other models and achieved the highest accuracy of 98.64% . In addition, bagging outperformed the current state-of-the-art approaches which emphasize its effectiveness in addressing spam-related challenges on the Internet. This suggests that bagging is a promising approach for URL spam classification.  \nKeywords— Spam, URL, dataset, machine learning, model, KNeighbors, bagging, random forest, logistic regression, classifier  \nI. INTRODUCTION  \nThe Internet is an open space for everyone to freely create content, publish it, and share it with others. In the last decade, internet access has increased tremendously. This increase in audience came with some side effects. Too many adsand spam are being shared everywhere, whether it is email or any other type of social media. Platforms like email clients which are used by hundreds of millions of users every day struggle to effectively filter the content to the end user,[1][2] .  \nBlacklist is one of the common methods used to identify malicious URLs. Although blacklisting has been effective for many URLs, the rapid increase in these URLs makes it an insufficient method. Therefore, Machine learning techniques have been proposed to address this issue [3] . These techniques can detect  \nmalicious websites, even if they have never been encountered before. In the context of this discussion, machine learning (ML) refers to the field of computer algorithms that autonomously enhance their performance through experiential learning and data analysis [4], [5] . Because ML models possess the capability to comprehend the underlying structural patterns within URLs, they provide more insightful methods for classifying URLs [6], [7], [8] .  \nIn this study, we want to investigate the use of machine learning models to classify URLs that are most likely spam. The model’s input and output are pretty straightforward; it will take a URL and it will classify it as spam or not. Since we are taking only one feature as an input, we will need to analyze and extend this feature to extract more information about the URL to determine its characteristics, [9], [10] . In addition, we will build multiple machine learning models, with different parameters to evaluate different options, and finally to come up with the best model that can result in the highest possible accuracy.  \nThe paper is organized into five sections. Section 2 covers related works in spam URLs. Section 3 presents the proposed ML models, while Section 4 discusses the obtained results. Finally section 5 concludes the paper.  \nII. LITERATURE REVIEW  \nURL spam detection is a modernistic field that form a solid interest for both organizations and researchers. It started to receive more and more attention from researchers due to the major evolving and exposing that happened on the internet over the past few years. Previous work  \n2  \non this topic has involved analysis of the URL and the page itself.  \nOshingbesan et al, 2023, [11] examined different ML models to classify malicious websites across different datasets. From their results, K-Nearest neighbo","cbCailhDiR724TTK","https://ap.wps.com/l/cbCailhDiR724TTK","pdf",946184,1,9,"English","en",105,"# Introduction\n## Background and motivation\n## Proposed study and organization\n# Literature Review\n## Prior URL spam detection methods\n# (Implied) Proposed Models\n# (Implied) Results\n# Conclusion","[{\"question\":\"What is the main goal of this study?\",\"answer\":\"Classify URLs as spam or non-spam using machine learning models by extracting URL features and evaluating multiple classifiers.\"},{\"question\":\"Which machine learning models are compared in the study?\",\"answer\":\"The study compares k-nearest neighbors, bagging, random forest, logistic regression, and additional models with different parameters.\"},{\"question\":\"What performance does the best model achieve?\",\"answer\":\"Bagging outperforms other models and achieves the highest accuracy of 98.64%.\"}]","Classification of Spam URLs Using Machine Learning Approachs - Abstract | PDF",1785807626,23,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"classification-of-spam-urls-using-machine-learning-approaches-abstract","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/classification-of-spam-urls-using-machine-learning-approaches-abstract/121896/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the main goal of this study?","Question",{"text":75,"@type":76},"Classify URLs as spam or non-spam using machine learning models by extracting URL features and evaluating multiple classifiers.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which machine learning models are compared in the study?",{"text":80,"@type":76},"The study compares k-nearest neighbors, bagging, random forest, logistic regression, and additional models with different parameters.",{"name":82,"@type":73,"acceptedAnswer":83},"What performance does the best model achieve?",{"text":84,"@type":76},"Bagging outperforms other models and achieves the highest accuracy of 98.64%.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]