[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-121332-en":3,"doc-seo-121332-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},121332,687197207057,"Sage","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Detecting Fake News on Social Media - A Data Mining Perspective - Exploring Machine Learning","False information spread on social media undermines information quality and can destabilize society. This thesis examines how machine learning can identify fraudulent content in online channels using a balanced Kaggle dataset with real and fake news categories. Supervised models, including logistic regression, random forest, and gradient boosting, are trained with TF-IDF feature extraction after text normalisation and tokenisation. Model quality is evaluated via accuracy, precision, recall, F1-score, and AUC-ROC, with logistic regression achieving the strongest F1 performance.","DETECTING FAKE NEWS ON SOCIAL MEDIA: A DATA MINING PERSPECTIVE-EXPLORING MACHINE  \nLEARNING  \nObajimi George Adewunmi  \nThesis  \nInformation and Communication Technology Machine Learning and Data Engineering (LapinAMK)  \n2025  \nStudy Programme in Information and Communication Technology Bachelor of Engineering  \n\n| Author Supervisor Commissioned by Title of Thesis\u003Cbr>Number of pages | Adewunmi Obajimi Year 2025\u003Cbr>Kenneth Karlson\u003Cbr>LapinAMK\u003Cbr>Detecting Fake News on Social Media: A Data Mining Perspective-Exploring Machine Learning 54 |\n| --- | --- |\n\nThe widespread spread of false information seriously threatens the quality of information and the stability of society. This paper looked at how machine learning techniques could be used to find fraudulent information on online channels. A balanced dataset from Kaggle comprises 51,063 entries, categorised as real news (24,563) and fake news (26,500) . Many supervised machine learning algorithms—including logistic regression, random forest, and gradient boosting—were used to build and evaluate the model. Text normalisation, tokenisation, and feature extraction using the Term Frequency-Inverse Document Frequency (TFIDF) technique constituted the data preparation stage. In the data there is class balance therefore guaranteeing strong model performance. The assessment of the model used several measures: accuracy, precision, recall, F1-score, and AUC-ROC. According to the results, Logistic Regression was the bestperforming model with an F1 score of 0.959 and a precision of 96.1 percent. Among other visualisation techniques, confusion matrices, bar charts, and metric comparisons improved the clarity of the model projections. By offering scalable solutions appropriate for real-world situations, this study underlined the possibility of machine learning to solve the growing problem of disinformation. Future studies should concentrate on combining transformers with deep learning architectures to improve contextual analysis and expand the range to cover multilingual datasets, hence increasing applicability.  \nKeywords  \nFake News Detection, Machine Learning, Gradient Boosting, Supervised Learning, Natural Language Processing, TF-IDF, Digital Misinformation, Text Classification, Data Preprocessing.  \nTABLE OF CONTENTS  \n1. INTRODUCTION ............................................................................................5  \n1.1 Statement of Problem ............................................................................9  \n1.2 Research Aims and Objectives ..............................................................9  \n1.3 Research Questions ............................................................................ 10  \n1.4 Justification of the Study........................................................................ 10  \n1.5 Scope of Study .................................................................................... 10  \n2.0 LITERATURE REVIEW .............................................................................. 11  \n2.1 Definition of Fake News.......................................................................... 12  \n2.2. Differentiating Misinformation from Fake News ...................................... 13  \n2.3. Fake News on Social Media ................................................................... 14  \n2.4. The 2016 U.S. Presidential Election ....................................................... 16  \n2.5. The COVID-19 Pandemic ....................................................................... 17  \n2.6. The 2020 U.S. Presidential Election ....................................................... 17  \n2.7. Review of Previous Studies Addressing Fake News Detection on Social Media ............................................................................................................ 18  \n2.8. The Rise of Fake News and Its Implications and their Economic Consequences ........................................................................","cbCainl6jnBRpAKq","https://ap.wps.com/l/cbCainl6jnBRpAKq","pdf",761512,1,58,"English","en",105,"# Introduction\n## Statement of Problem\n## Research Aims and Objectives\n## Research Questions\n## Justification of the Study\n## Scope of Study\n# Literature Review\n## Definition of Fake News\n## Differentiating Misinformation from Fake News\n## Fake News on Social Media\n## The 2016 U.S. Presidential Election\n## The COVID-19 Pandemic\n## The 2020 U.S. Presidential Election\n## Review of Previous Studies Addressing Fake News Detection\n## Gaps in the Literature\n# Research Methodology\n## Research Design\n## Data Collection\n## Machine Learning Approaches\n## Feature Extraction and Natural Language Processing\n## Model Evaluation and Validation\n## Tools, Software, and Libraries\n# Results and Analysis\n## Text Preprocessing\n## Feature Extraction\n## Model Performance\n## Confusion Matrix Analysis","[{\"question\":\"What dataset and class balance are used for fake news detection?\",\"answer\":\"A balanced Kaggle dataset with 51,063 entries is used, categorised into real news (24,563) and fake news (26,500). The class balance supports strong model performance.\"},{\"question\":\"Which models and feature extraction method are applied in the study?\",\"answer\":\"Supervised learning algorithms such as logistic regression, random forest, and gradient boosting are applied. Text normalisation, tokenisation, and TF-IDF feature extraction are used during data preparation.\"},{\"question\":\"How is model performance evaluated and what is the best-performing model?\",\"answer\":\"Performance is assessed using accuracy, precision, recall, F1-score, and AUC-ROC. Logistic regression achieves the best results with an F1-score of 0.959 and precision of 96.1%.\"}]","Detecting Fake News on Social Media - A Data Mining Perspective - Exploring Machine Learning | PDF",1785735113,146,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"detecting-fake-news-on-social-media-a-data-mining-perspective-exploring-machine-learning","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/detecting-fake-news-on-social-media-a-data-mining-perspective-exploring-machine-learning/121332/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What dataset and class balance are used for fake news detection?","Question",{"text":75,"@type":76},"A balanced Kaggle dataset with 51,063 entries is used, categorised into real news (24,563) and fake news (26,500). The class balance supports strong model performance.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which models and feature extraction method are applied in the study?",{"text":80,"@type":76},"Supervised learning algorithms such as logistic regression, random forest, and gradient boosting are applied. Text normalisation, tokenisation, and TF-IDF feature extraction are used during data preparation.",{"name":82,"@type":73,"acceptedAnswer":83},"How is model performance evaluated and what is the best-performing model?",{"text":84,"@type":76},"Performance is assessed using accuracy, precision, recall, F1-score, and AUC-ROC. Logistic regression achieves the best results with an F1-score of 0.959 and precision of 96.1%.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]