[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-117878-en":3,"doc-seo-117878-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},117878,8796095360427,"Lucas Martin","https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d",8,"Research & Report","A STUDY OF VARIOUS DATA SIZES USING MACHINE LEARNING","Social media enables low-cost, user-friendly news consumption, yet it also accelerates the spread of fake news, harming society, businesses, and consumers. Fake news detection remains an emerging research area, but progress toward a universally effective machine learning model is constrained by limited resources, especially large-scale datasets. This project examines how varying dataset sizes influence a model’s accuracy. Naïve Bayes models trained with Kaggle data are evaluated per dataset, then combined to test whether accuracy improves with more data.","California State University, San Bernardino  \nCSUSB ScholarWorks  \n\n| Electronic Theses, Projects, and Dissertations | Office of Graduate Studies |\n| --- | --- |\n| 5-2023\u003Cbr>A STUDY OF VARIOUS DATA SIZES USING MACHINE LEARNING Sochaeta Koeum\u003Cbr>Follow this and additional works at: [https://scholarworks.lib.csusb.edu/etd](https://scholarworks.lib.csusb.edu/etd)\u003Cbr> Part of the Databases and Information Systems Commons, and the Data Science Commons |  |\n\nRecommended Citation  \nKoeum, Sochaeta, \"A STUDY OF VARIOUS DATA SIZES USING MACHINE LEARNING\" (2023) . Electronic Theses, Projects, and Dissertations. 1694.  \n[https://scholarworks.lib.csusb.edu/etd/1694](https://scholarworks.lib.csusb.edu/etd/1694)  \nThis Project is brought to you for free and open access by the Office of Graduate Studies at CSUSB ScholarWorks. It has been accepted for inclusion in Electronic Theses, Projects, and Dissertations by an authorized administrator of CSUSB ScholarWorks. For more information, please contact [scholarworks@csusb.edu](scholarworks@csusb.edu).  \nA STUDY OF VARIOUS DATA  \nSIZES USING MACHINE LEARNING  \nA Project Presented to the Faculty of  \nCalifornia State University, San Bernardino  \nIn Partial Fulfillment of the Requirements for the Degree Master of Science in  \nInformation Systems and Technology  \nby Sochaeta Koeum  \nMay 2023  \nA STUDY OF VARIOUS DATA  \nSIZES USING MACHINE LEARNING  \nA Project Presented to the Faculty of  \nCalifornia State University, San Bernardino  \nby  \nSochaeta Koeum  \nMay 2023  \nApproved by:  \nDr. Conrad Shayo , Committee Chair  \nDr. Bailey Benedict , Committee Member  \nDr. Conrad Shayo , Department Chair, Information Decision Sciences  \n© 2023 Sochaeta Koeum  \nABSTRACT  \nSocial media is a great domain for news consumption; however, it is referred to as a double-edged sword. While it is user-friendly and low-cost, social media is the reason why fake news can spread rapidly, which is detrimental to society, businesses, and many consumers. Therefore, fake news detection is an emerging field. However, some challenges have restricted other researchers from developing a universal machine learning model that is fast , efficient , and reliable to stop the proliferation because of the lack of resources available , such as large-sized datasets. The goal of this culminating experience project is to explore how varying datasets sizes affect the accuracy percentage of a machine learning model. The research questions are: Q1) How do large volumes of fake news datasets affect the accuracy percentage of a machine learning model? Q2) As one increases the volume of data fed into a machine learning model from small to large datasets, what will the cutoff accuracy percentage point be? Various data sizes collected from Kaggle were fed into the machine learning model, Naïve Bayes, to help answer the two questions. Then, all three datasets were combined together to see if the accuracy of the model improves as more data is fed into the model. The results and findings for each question are; 1) Larger dataset sizes do increase the accuracy percentage because there is more data to train and test on. 2) The cutoff accuracy is dependent on the number of unique values within the dataset. Since it is not finite, we can expect that large dataset sizes to have a cutoff accuracy of above 90%, given that the data is  \ncleaned and pre-processed. Compared to a data size ranging from small to medium, it will achieve an accuracy score of around 70%-90% . An accuracy score below 70% means that the model is highly unreliable and that the datasetsize is too small. For instance , Dataset 1 achieved an accuracy score of 66%, Dataset 2 was 83%, and Dataset 3 was 92% . To effectively study and experiment on how to build an optimized model, one must use a large datasetsize for analysis. Furthermore, other areas for future studies that appeared from this study are building a new and improved fact-checking website that quickly and accurately processes large d","cbCailjXFmjQf95C","https://ap.wps.com/l/cbCailjXFmjQf95C","pdf",546950,1,51,"English","en",105,"# TABLE OF CONTENTS\n## ABSTRACT\n## LIST OF FIGURES\n## CHAPTER ONE: INTRODUCTION\n## Brief Background\n## Fake News\n## Machine Learning and Algorithms\n## Problem Statement\n## Research Questions\n## Objective\n## Organization of this Project","[{\"question\":\"What problem does the project address?\",\"answer\":\"The project addresses how fake news can spread rapidly through social media and how limited resources, particularly lack of large datasets, restrict development of reliable machine learning models.\"},{\"question\":\"How does the project evaluate the effect of dataset size on accuracy?\",\"answer\":\"Various dataset sizes collected from Kaggle are used to train and test a machine learning model (Naïve Bayes). The datasets are then combined to check whether accuracy improves as more data is provided.\"},{\"question\":\"What does the project conclude about larger datasets?\",\"answer\":\"Larger dataset sizes increase accuracy because they provide more data to train and test. The study also notes that the accuracy cutoff depends on the number of unique values in the dataset, with cleaned and pre-processed large datasets reaching above 90%.\"}]","A STUDY OF VARIOUS DATA SIZES USING MACHINE LEARNING | PDF",1785680114,129,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"a-study-of-various-data-sizes-using-machine-learning","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/a-study-of-various-data-sizes-using-machine-learning/117878/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the project address?","Question",{"text":75,"@type":76},"The project addresses how fake news can spread rapidly through social media and how limited resources, particularly lack of large datasets, restrict development of reliable machine learning models.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the project evaluate the effect of dataset size on accuracy?",{"text":80,"@type":76},"Various dataset sizes collected from Kaggle are used to train and test a machine learning model (Naïve Bayes). The datasets are then combined to check whether accuracy improves as more data is provided.",{"name":82,"@type":73,"acceptedAnswer":83},"What does the project conclude about larger datasets?",{"text":84,"@type":76},"Larger dataset sizes increase accuracy because they provide more data to train and test. The study also notes that the accuracy cutoff depends on the number of unique values in the dataset, with cleaned and pre-processed large datasets reaching above 90%.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]