[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-117810-en":3,"doc-seo-117810-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},117810,549758146520,"Patrick","https://ap-avatar.wpscdn.com/avatar/80002397d8c0411e94?_k=1775819394049821470",8,"Research & Report","Machine Learning Based Phishing Attacks Detection Using Multiple Datasets","Phishing attacks increasingly target individuals and organizations, making reliable detection systems essential for protecting credentials and payment-related information. This study trains and evaluates machine learning classifiers using three publicly archived datasets from Phish-Tank and the UCI Machine Learning Repository. An Information Gain (IG) approach performs feature reduction and selection, then compares model behavior across two stages: all-features learning versus IG-selected features. Results show RandomForest performs best overall before selection, while after feature selection it remains strongest on Dataset-1 and Dataset-2, and ANN leads on Dataset-3, with performance affected by Dataset-3’s limited instances and features.","Machine Learning Based Phishing Attacks Detection Using Multiple Datasets  \n[https://doi.org/10.3991/ijim.v17i05.37575](https://doi.org/10.3991/ijim.v17i05.37575)  \nAshraf H. Aljammal1(􀀍 ), Salah taamneh 1, Ahmad Qawasmeh1, Hani Bani Salameh2 1 Department of Computer Science and Applications, The Hashemite University, Zarqa, Jordan  \n2 Department of Software Engineering, The Hashemite University, Zarqa, Jordan [ashrafj@hu.edu.jo](ashrafj@hu.edu.jo)  \nAbstract—Nowadays, individuals and organizations are increasingly targeted by phishing attacks, so an accurate phishing detection system is required. Therefore, many phishing detection techniques have been proposed as well as phishing datasets have been collected. In this paper, three datasets have been used to train and test machine learning classifiers. The datasets have been archived by Phish-Tank and UCI Machine Learning Repository. Furthermore, Information Gain algorithm have been used for features reduction and selection purpose. In addition, six machine learning classifiers have been evaluated, namely NaiveBayes, ANN, DecisionStump, KNN, J48 and RandomForest. However, the classifiers have been trained and tested over the three datasets in two stages. The first stage is using all features included in each dataset while the second stage using selected features by IG algorithm. At the first stage RandomForest classifier has shown the best performance over Dataset-1 and Dataset-2, while J48 has shown the best performance over Dataset-3. On the other hand, after features selection, the RandomForest classifier was the superior among the other five classifiers over Dataset-1 and Dataset-2 with accuracy of 98% and 93.66% respectively. While ANN classifier has shown the best performance with accuracy of 88.92% over Dataset-3. Because of the few number of instances as well as features in Dataset-3 comparing to the other two dataset; the performance of the classifiers has been affected.  \nKeywords—phishing attack, phishing attack detection, cybersecurity, machine learning, web security, network security  \n1 Introduction  \nPhishing is considered a cybersecurity threat which is used to commit fraudulent actions; such as stealing users sensitive information which includes user accounts credentials and banking cards information [1-3] . The phishing attack occurs when attackers pretend to be a trusted entity and attracting the victim to open the socially engineered sent message (Email, IM message or text message)[4, 5] . Sometimes, these messages ask the user to input a critical information or even entreating the user for financial gain. Furthermore, the content of the message could be a URL of a rogue version of legitimate webpage created by the attacker and hosted on their own servers  \nluring user to click the URL. However, it will be hard for the internet user to differentiate between the legitimate and phishing (mimic) webpages. When exploring these webpages it eventually leads the user to reveal her/his own information to the attacker. However, attackers always attract user to explore the phishing website. Therefore, the success of phishing attacks rely on the weaknesses of the user [6] . Individual users as well as organizations are vulnerable to many types of attacks including phishing attacks[7, 8] . According to Anti-Phishing Working Group (APWG)[9] report, the phishing attacks ratio is doubled since early 2020. The report statistics show that the highest number of phishing attacks was in July 2021 with 260,642 attacks. In addition, the most effected sectors in phishing attacks were software-as-a-service and webmail with 29. 1% of total attacks. Furthermore, the combined attacks against financial institutions and payment providers with 34.9% of all attacks. Figure 1 shows the most targeted industries in the 3rd quarter of 2021.  \nMOST TARGETED INDUSTRIES 3Q 2021  \nSocial Media,  \n11.00%  \nTelecom, Crypto,  \n3.50% 5.60%  \neCommerce / Retail,  \n13.10%  \nFinancial Institution  \n17.8","cbCaimlOo6Grr3sk","https://ap.wps.com/l/cbCaimlOo6Grr3sk","pdf",1022239,1,13,"English","en",105,"# Introduction\n## Literature review\n# Proposed approach\n## Implementation and results analysis\n# Conclusion","[{\"question\":\"How many datasets and where are they sourced from?\",\"answer\":\"The study uses three datasets. They are archived by Phish-Tank and the UCI Machine Learning Repository.\"},{\"question\":\"What role does the Information Gain algorithm play in the workflow?\",\"answer\":\"Information Gain is applied for feature reduction and selection. The models are trained and tested in two stages: with all features and with IG-selected features.\"},{\"question\":\"Which classifiers perform best on each dataset?\",\"answer\":\"In the first stage, RandomForest achieves the best performance on Dataset-1 and Dataset-2, while J48 is best on Dataset-3. After feature selection, RandomForest is superior on Dataset-1 and Dataset-2, and ANN performs best on Dataset-3.\"}]","Machine Learning Based Phishing Attacks Detection Using Multiple Datasets | PDF",1785679678,33,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"machine-learning-based-phishing-attacks-detection-using-multiple-datasets","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/machine-learning-based-phishing-attacks-detection-using-multiple-datasets/117810/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05","2026-08-02",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"How many datasets and where are they sourced from?","Question",{"text":76,"@type":77},"The study uses three datasets. They are archived by Phish-Tank and the UCI Machine Learning Repository.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What role does the Information Gain algorithm play in the workflow?",{"text":81,"@type":77},"Information Gain is applied for feature reduction and selection. The models are trained and tested in two stages: with all features and with IG-selected features.",{"name":83,"@type":74,"acceptedAnswer":84},"Which classifiers perform best on each dataset?",{"text":85,"@type":77},"In the first stage, RandomForest achieves the best performance on Dataset-1 and Dataset-2, while J48 is best on Dataset-3. After feature selection, RandomForest is superior on Dataset-1 and Dataset-2, and ANN performs best on Dataset-3.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]