[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118266-en":3,"doc-seo-118266-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},118266,687197207639,"Asher","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Cyberbullying Detection on Twitter Data Using Machine Learning Classifiers - Abstract","This study compares Logistic Regression, Multinomial Naive Bayes, K-Nearest Neighbor, and Extreme Gradient Boosting to classify tweets into three categories: cyberbullying based on religion, cyberbullying based on ethnicity, and no cyberbullying. Tweet data is first cleaned through preprocessing to improve quality. After cleaning, word embedding methods—Bag of Words and Term Frequency-Inverse Document Frequency—convert text into numerical vectors. Models are then trained using each embedding-method plus classifier combination and evaluated to support accurate cyberbullying word detection.","University of Central Florida  \nSTARS  \nData Science and Data Mining  \nMay 2024  \nCyberbullying Detection on Twitter Data Using Machine Learning Classifiers  \nPradip Dhakal  \nUniversity of Central Florida, [pradip.dhakal@ucf.edu](pradip.dhakal@ucf.edu)  \n Part of the Data Science Commons  \nFind similar works at: [https://stars.library.ucf.edu/data-science-mining](https://stars.library.ucf.edu/data-science-mining)  \nUniversity of Central Florida Libraries [http://library.ucf.edu](http://library.ucf.edu)  \nThis Article is brought to you for free and open access by STARS. It has been accepted for inclusion in Data Science and Data Mining by an authorized administrator of STARS. For more information, please [contact STARS@ucf.edu](contact STARS@ucf.edu).  \nSTARS Citation  \nDhakal, Pradip, \"Cyberbullying Detection on Twitter Data Using Machine Learning Classifiers\" (2024) . Data Science and Data Mining. 23.  \n[https://stars.library.ucf.edu/data-science-mining/23](https://stars.library.ucf.edu/data-science-mining/23)  \nCyberbullying Detection on Twitter Data Using Machine Learning Classifers  \nPradip Dhakal  \nStatistics and Data Science Department  \nUniversity of Central Florida  \nAbstract—This study compares some of the popular machine learning techniques like Logistic Regression, Multinomial Naive Bayes, K-Nearest Neighbor, and Extreme Gradient Boosting to classify the tweets into three different categories: cyberbullying based on religion, cyberbullying based on ethnicity, or no cyberbullying. First, various data-cleaning approaches are used to clean the tweet data. After the data is clean and ready, the word embedding techniques, such as a bag of words and term frequency-Inverse document frequency, are used to convert the words into mathematical vectors. Finally, the model will be ftted using the combination of the above-mentioned word embedding techniques and machine learning algorithms.  \nKeywords—Logistic Regression, Multinomial Naive Bayes, KNearest Neighbor, Extreme Gradient Boosting, Bag of Words, Term Frequency-Inverse Document Frequency  \nI. INTRODUCTION  \nSocial media has been popular for quite a long time. People from all around the world are able to communicate with eachother, share their knowledge and thoughts, and know what’s happening on another side of the world. Some of the popular social platforms are Facebook, Instagram, Twitter, YouTube, and Snapchat. Alongside these advantages, people argue with each other, show aggressive behaviors, leave negative and racist comments, and bully other people on different social platforms. Many people have been victims of these kinds of activities, leading to an adverse psychological impact on the victim’s emotions. Many research studies have been proposed to mitigate these kinds of activities. Most researchers formulated this problem as a classifcation problem. The study by Dinakar et al. [1] performed binary classifcation to see whether the comments on YouTube could be classifed as sensitive or not; they also performed multi-label classifcation to see what comments belong to what classes. Another study from Chavanand Shylaja [2] performed the binary classifcation, where they classifed the texts as bullying texts and non-bullying texts. Similarly, the study from Dadvar et al. [3] incorporated the user’s age and gender to improve the accuracy of cyberbullying detection. These articles motivated me to work in the area of cyberbullying detection.  \nIn addition to these foundational works, recent studies have further advanced our understanding and methodologies. For instance, Wang et al. [4] introduced the triangular user-activitycontent view, emphasizing the importance of understanding the defning features of online bullying users. This perspective has signifcantly informed our approach to cyberbullying detection. Alipour Yengejeh [5] explored machine learning algo-  \nrithms to detect cyberbullying through text analysis on Twitter, demonstrating the effectiveness of logistic reg","cbCaie09R8lKKXoG","https://ap.wps.com/l/cbCaie09R8lKKXoG","pdf",849274,1,7,"English","en",105,"# Introduction\n# Design Overview\n# Data\n## Data Information","[{\"question\":\"Which machine learning classifiers are compared in the study?\",\"answer\":\"The study compares Logistic Regression, Multinomial Naive Bayes, K-Nearest Neighbor, and Extreme Gradient Boosting for tweet classification.\"},{\"question\":\"How are tweets transformed into machine-learning inputs?\",\"answer\":\"Tweets are cleaned and then converted into vectors using word embedding techniques such as Bag of Words and Term Frequency-Inverse Document Frequency (TF-IDF).\"},{\"question\":\"What are the three classification categories and labels used?\",\"answer\":\"The study uses three classes: ethnicity-based cyberbullying (label 1), religion-based cyberbullying (label 2), and no cyberbullying (label 0).\"}]","Cyberbullying Detection on Twitter Data Using Machine Learning Classifiers - Abstract | PDF",1785682720,18,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"cyberbullying-detection-on-twitter-data-using-machine-learning-classifiers-abstract","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/cyberbullying-detection-on-twitter-data-using-machine-learning-classifiers-abstract/118266/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Which machine learning classifiers are compared in the study?","Question",{"text":75,"@type":76},"The study compares Logistic Regression, Multinomial Naive Bayes, K-Nearest Neighbor, and Extreme Gradient Boosting for tweet classification.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How are tweets transformed into machine-learning inputs?",{"text":80,"@type":76},"Tweets are cleaned and then converted into vectors using word embedding techniques such as Bag of Words and Term Frequency-Inverse Document Frequency (TF-IDF).",{"name":82,"@type":73,"acceptedAnswer":83},"What are the three classification categories and labels used?",{"text":84,"@type":76},"The study uses three classes: ethnicity-based cyberbullying (label 1), religion-based cyberbullying (label 2), and no cyberbullying (label 0).","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]