[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-122682-en":3,"doc-seo-122682-105":30,"detail-sidebar-cat-0-en-105":84},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},122682,549758146520,"Patrick","https://ap-avatar.wpscdn.com/avatar/80002397d8c0411e94?_k=1775819394049821470",8,"Research & Report","Feature Selection to Enhance Phishing Website Detection Based On URL - Using Machine Learning Techniques","Phishing website detection based on machine learning receives strong attention for identifying newly generated phishing URLs. Many approaches combine URL, web page content, and external features, but page content and external signals introduce high time costs and computing requirements, limiting deployment on resource-constrained devices. This study applies URL-based feature selection to streamline detection. The workflow includes preprocessing, dataset splitting, feature selection, 10-fold cross-validation, model validation, and performance evaluation on two public datasets.","Feature Selection to Enhance Phishing Website Detection Based On URL Using Machine Learning Techniques Liyana Mat Rani1,2, Cik Feresa Mohd Foozy1* Siti Noor Baini Mustafa3  \n1Faculty of Computer Science and Information Technology,  \nUniversiti Tun Hussein Onn Malaysia, Parit Raja, Batu Pahat, 86400, Johor, MALAYSIA  \n2Department of Information Technology and Communication,  \nPoliteknik Sultan Mizan Zainal Abidin, KM 8, Jalan Paka, 23000 Dungun, Terengganu, MALAYSIA  \n3Book Hack Enterprise,  \nBandar Baru Nilai,71800 Nilai, Negeri Sembilan, MALAYSIA  \nDOI: [https://doi.org/10.30880/jscdm.2023.04.01.003](https://doi.org/10.30880/jscdm.2023.04.01.003)  \nReceived 02 February 2023; Accepted 12 April 2023; Available online 25 May 2023  \nAbstract: The detection of phishing websites based on machine learning has gained much attention due to its ability to detect newly generated phishing URLs. To detect phishing websites, most techniques combine URLs, web page content, and external features. However, the content of the web page and external features are time-consuming, require large computing power, and are not suitable for resource-constrained devices. To overcome this problem, this study applies feature selection techniques based on the URL to improve the detection process. The methodology for this study consists of seven stages, including data preparation, preprocessing, splitting the dataset into training and validation, feature selection, 10-fold cross-validation, validating the model, and finally performance evaluation. Two public datasets were used to validate the method. TreeSHAP and Information Gain were used to rank features and select the top 10, 15, and 20. These features are fed into three machine learning classifiers which are Naïve Bayes, Random Forest, and XGBoost. Their performance is evaluated based on accuracy, precision, and recall. As a result, the features ranked by TreeSHAP contributed most to improving detection accuracy. The highest accuracy of 98.59 percent was achieved by XGBoost for the first dataset with 15 features. For the second dataset, the highest accuracy is 90.21 percent using 20 features and Random Forest. As for Naïve Bayes, the highest accuracy recorded is 98.49 percent using the first dataset.  \nKeywords: Machine learning, phishing detection, feature selection, information gain, TreeSHAP  \n1. Introduction  \nPhishing is a social engineering attack that aims to steal a user’s identity data and financial account credentials [1] . Attackers exploit the lack of cybersecurity awareness among users and the insecure Internet protocol to trick users into visiting phishing websites. The current problem in phishing website detection is the inability of the list-based and visual similarity technique to detect zero-day phishing attacks due to the nature of the phishing websites, which only appear fora short time [2], [3] . Recent phishing techniques use hybrid features, which combine URL, content, and external-based [4]-[6] . However, web page content and external-based features are time-consuming [7], [8] . Many past works of literature used a variety of features but did not include any information about feature selection [9], [10], which is crucial since using many features requires a lot of computing power and is not suitable for resource-constrained devices such as mobile phones [11] . Therefore, this study aims to employ a feature selection technique based on URLs using machine learning to improve the detection of phishing websites based on URLs.  \nThe objectives ofthis research are (1) to study feature profiles for phishing website detection based on URL, (2) to develop a feature selection model for phishing website detection based on URL using machine learning techniques,(3)  \nto validate the website phishing detection model in terms of accuracy, precision, and recall. This study will make three contributions. The first is feature profiles that can be applied to phishing website detection. Second, the d","cbCaisgEe3cymY42","https://ap.wps.com/l/cbCaisgEe3cymY42","pdf",675444,1,12,"English","en",105,"# Introduction\n## Objectives and contributions\n## Paper organization\n# Related Works\n## Phishing\n## Phishing technique\n# Feature Selection to Enhance Phishing Website Detection\n## URL-based methodology overview\n## Classifiers and evaluation metrics","[{\"question\":\"What are the best reported accuracies for the classifiers?\",\"answer\":\"XGBoost achieves the highest accuracy of 98.59% with 15 features on the first dataset. For the second dataset, Random Forest reaches 90.21% with 20 features, while Naïve Bayes records 98.49% on the first dataset.\"}]","Feature Selection to Enhance Phishing Website Detection Based On URL - Using Machine Learning Techniques | PDF",1785812165,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":79,"head_meta":81,"extra_data":83,"updated_unix":28},"feature-selection-to-enhance-phishing-website-detection-based-on-url-using-machine-learning-techniques","",{"@graph":36,"@context":78},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/feature-selection-to-enhance-phishing-website-detection-based-on-url-using-machine-learning-techniques/122682/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05","2026-08-04",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72],{"name":73,"@type":74,"acceptedAnswer":75},"What are the best reported accuracies for the classifiers?","Question",{"text":76,"@type":77},"XGBoost achieves the highest accuracy of 98.59% with 15 features on the first dataset. For the second dataset, Random Forest reaches 90.21% with 20 features, while Naïve Bayes records 98.49% on the first dataset.","Answer","https://schema.org",{"og:url":52,"og:type":80,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":82,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":85},[86,90,94,98,103,108,113,115,120,123,127],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":87,"show_sort_weight":88,"slug":89},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":91,"show_sort_weight":92,"slug":93},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Exam",70,"exam",{"id":99,"doc_module":4,"doc_module_name":46,"category_name":100,"show_sort_weight":101,"slug":102},5,"Comic",60,"comic",{"id":104,"doc_module":4,"doc_module_name":46,"category_name":105,"show_sort_weight":106,"slug":107},6,"Technology",50,"technology",{"id":109,"doc_module":4,"doc_module_name":46,"category_name":110,"show_sort_weight":111,"slug":112},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":114},"research-report",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":128,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":99,"slug":130},19,"General","general"]