[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-120402-en":3,"doc-seo-120402-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},120402,2336464648746,"Skyler","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Detecting Real and Forged Tweets Using Machine Learning - A Thesis - University Honors Program Requirements","Twitter’s influence makes information reliability dependent on whether tweets originate from legitimate owners or from forged posts created after account hacking. This thesis applies machine learning models to detect real versus forged tweets using features extracted from tweet content alone. A dataset combining forged tweets and tweets from legitimate users is constructed and cleaned, then engineered features such as sentiment polarity, part-of-speech counts, and URL presence are used for exploratory analysis and model evaluation. Logistic regression achieves the highest accuracy, 61.57%, though results indicate tweet-only features are limited and should be augmented with profile data for improved detection.","DETECTING REAL AND FORGED TWEETS USING MACHINE LEARNING  \nA THESIS  \nPresented to the University Honors Program  \nCalifornia State University, Long Beach  \nIn Partial Fulfillment  \nof the Requirements for the  \nUniversity Honors Program Certificate  \nNyssa Yota  \nMay 2023  \nI, THE UNDERSIGNED MEMBER OF THE COMMITTEE,  \nHAVE APPROVED THIS THESIS  \nDETECTING REAL AND FORGED TWEETS USING MACHINE LEARNING  \nBY  \nNyssa Yota  \n_____________________________________________________________  \nXiyue Liao, Ph.D. (Thesis Advisor) Statistics  \nCalifornia State University, Long Beach  \nABSTRACT  \nDETECTING REAL AND FORGED TWEETS USING MACHINE LEARNING  \nBy  \nNyssa Yota  \nMay 2023  \nOne of the most popular social media platforms in which information spreads is Twitter. Reliability of this information can lie in whether the tweet was real and written by the actual owner of the account or whether the tweet was forged and written by someone who hacked into the account. This paper will use machine learning models to detect real and forged tweets on Twitter, using features extracted from only tweets.  \nWe started with creating a dataset by combining a set of forged tweets with a set of tweets from legitimate users. Once the dataset was cleaned, we created features based on the tweets, such as the polarity of each tweet, number of nouns, and whether URLs were included. Exploratory data analysis was conducted, and different machine learning models were run on the dataset to see which model performed best.  \nThe model that resulted in the best accuracy in detecting the forged and real tweets was the logistic regression, with an accuracy of 61 .57%. This accuracy is quite low, meaning that using only the features extracted from only the tweet might not be sufficient to provide accurate detection. However, the methods used in this paper may be used conjointly with profile data to further improve accuracy in detecting impersonators.  \nACKNOWLEDGEMENTS  \nI would firstly like to thank my thesis advisor, Dr. Xiyue Liao, for sparking my interest in statistical machine learning and for her patience and guidance throughout this process. During the completion of this thesis, there were a few challenges, so I would also like to thank the people who lent an ear while I let out my frustrations, including my parents and friends. Lastly, I would like to thank myself for having the strength to push through this whole process.  \nTABLE OF CONTENTS  \nPage  \nACKNOWLEDGEMENTS ......................................................................................... iii  \n[LIST OF TABLES....................................................................................................... vi](LIST OF TABLES....................................................................................................... vi)  \n[LIST OF FIGURES ...............................](LIST OF FIGURES ...............................)...................................................................... vii  \nCHAPTER  \n1. INTRODUCTION ............................................................................................ 1  \n2. LITERATURE REVEW................................................................................... 3  \nBot and Gender Detecting using Tweet Content ....................................... 3  \nFake News Detection using Tweet Content ............................................... 4  \nBot Detection using Tweet Content and User’s Profile Data .................... 4  \nFake Twitter Identity Detection using User’s Profile Data ....................... 5  \nA Gap in the Literature: The Present Study............................................... 6  \n3. METHODOLOGY ........................................................................................... 7  \nThe Dataset ................................................................................................ 7  \nFeature Engineering ................................................................................... 8  ","cbCaiovJ3ALEDFfW","https://ap.wps.com/l/cbCaiovJ3ALEDFfW","pdf",1311321,1,32,"English","en",105,"# Acknowledgements\n# List of Tables\n# List of Figures\n# Chapter 1. Introduction\n# Chapter 2. Literature Review\n## Bot and Gender Detecting using Tweet Content\n## Fake News Detection using Tweet Content\n## Bot Detection using Tweet Content and User’s Profile Data\n## Fake Twitter Identity Detection using User’s Profile Data\n## A Gap in the Literature: The Present Study\n# Chapter 3. Methodology\n## The Dataset\n## Feature Engineering\n## Exploratory Data Analysis\n## Machine Learning Models\n## Classification Models\n## Choosing the Best Model\n# Chapter 4. Results\n# Chapter 5. Conclusion\n# References","[{\"question\":\"How does the thesis detect real versus forged tweets?\",\"answer\":\"It trains machine learning models using features extracted from tweet content, including sentiment polarity, part-of-speech counts such as number of nouns, and whether URLs are present.\"},{\"question\":\"What dataset is used for model training and evaluation?\",\"answer\":\"The dataset combines forged tweets with tweets from legitimate users, and it is cleaned before feature engineering and exploratory data analysis.\"},{\"question\":\"Which model performs best, and what accuracy is reported?\",\"answer\":\"Logistic regression yields the best accuracy for detecting forged versus real tweets, with an accuracy of 61.57%.\"}]","Detecting Real and Forged Tweets Using Machine Learning - A Thesis - University Honors Program Requirements | PDF",1785729842,81,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"detecting-real-and-forged-tweets-using-machine-learning-a-thesis-university-honors-program-requirements","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/detecting-real-and-forged-tweets-using-machine-learning-a-thesis-university-honors-program-requirements/120402/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"How does the thesis detect real versus forged tweets?","Question",{"text":75,"@type":76},"It trains machine learning models using features extracted from tweet content, including sentiment polarity, part-of-speech counts such as number of nouns, and whether URLs are present.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What dataset is used for model training and evaluation?",{"text":80,"@type":76},"The dataset combines forged tweets with tweets from legitimate users, and it is cleaned before feature engineering and exploratory data analysis.",{"name":82,"@type":73,"acceptedAnswer":83},"Which model performs best, and what accuracy is reported?",{"text":84,"@type":76},"Logistic regression yields the best accuracy for detecting forged versus real tweets, with an accuracy of 61.57%.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]