[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-124203-en":3,"doc-seo-124203-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},124203,1099514067415,"Rowan","https://ap-avatar.wpscdn.com/avatar/100002539d78ffe74a7?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779092875211072502",8,"Research & Report","Combating Cyberbullying on Social Media - A Machine Learning Approach with Text Analysis on Twitter","The work addresses the rapid growth of cyberbullying, toxicity, and harassment on social media and its harmful effects on individuals and communities, motivating the need for automatic detection. It applies active machine learning models—Logistic Regression, Multinomial Naive Bayes, K-Nearest Neighbour, and Extreme Gradient Boosting—to Twitter textual data. Using Bag of Words and TF-IDF feature extraction, the study converts text to numerical representations and evaluates performance, reporting the strongest results with XGBoost and Logistic Regression.","University of Central Florida  \nSTARS  \nData Science and Data Mining  \nFall 2024  \nCombating Cyberbullying on Social Media: A Machine Learning Approach with Text Analysis on Twitter  \nAmir Alipour Yengejeh  \nUniversity of Central Florida, [amir.alipouryengejeh@ucf.edu](amir.alipouryengejeh@ucf.edu)  \n Part of the Data Science Commons  \nFind similar works at: [https://stars.library.ucf.edu/data-science-mining](https://stars.library.ucf.edu/data-science-mining)  \nUniversity of Central Florida Libraries [http://library.ucf.edu](http://library.ucf.edu)  \nThis Article is brought to you for free and open access by STARS. It has been accepted for inclusion in Data Science and Data Mining by an authorized administrator of STARS. For more information, please [contact STARS@ucf.edu](contact STARS@ucf.edu).  \nSTARS Citation  \nAlipour Yengejeh, Amir, \"Combating Cyberbullying on Social Media: A Machine Learning Approach with Text Analysis on Twitter\" (2024) . Data Science and Data Mining. 15.  \n[https://stars.library.ucf.edu/data-science-mining/15](https://stars.library.ucf.edu/data-science-mining/15)  \nCombating Cyberbullying on Social Media: A Machine Learning Approach with Text Analysis on  \nTwitter  \nAmir Alipour Yengejeh  \ndept. Statistics and Data Science  \nUniversity of Central Florida  \nOrlando, United States  \n[amir.alipouryengejeh@ucf.edu](amir.alipouryengejeh@ucf.edu)  \nAbstract—The popularity of the electronic mobile devices along with social media as well as networking websites have been tremendously increased in the recent year. Most people around the world daily engage in the variety of cyberspace additives. Even though the users can take most advantages of these system such as exchange the idea and information, being sociable, and enjoyments, they might be faced with such adverse behaviours such as toxicity, bullying, extremism, and cruelty. The recent statistics reports that such mentioned behaviours has been noticeably grown on the cyberspace such that can threaten the individuals and even any community. Thus, it is drastically demand to invent a device to detect cyberbullying automatically. To do so, most studies are using the idea of the classifcation and then machine learning algorithms to build such a device. In this study, therefore, we employed some active machine learning models like Logistic Regression(LR), Multinomial Naive Bayes(MNB), K-Nearest Neighbour (KNN), and Extreme Gradient Boosting(XGbost) on Twitter textual dataset to detect the quality of cyberbullying related to ethnicity and religion. Since the data is contextual, we used some feature extraction techniques like Bag of Words and TFIDF to convert the texts into numerical sets. According the computational results, we saw XGbost and LR achieves the highest performance.  \nIndex Terms—cyberbullying, social media, machine learning, classifcation, feature extraction  \nI. INTRODUCTION  \nBullying is considered as an intentional and most frequent adverse and aggressive behavior can be conducted by an individual or group of people to attack and insult another one as a victim [1] [2] . There are various types of bullying like verbally, physically, psychologically etc. One evolution form of the bullying is called cyberbullying can be carried out on a cyberspace [2] . Recently, the prevalence of cyberbullying has been a complicated and signifcant phenomenon in the online social medias like Twitter, Facebook, Instagram, and to name but a few. This is because that statistics shows that cyberbullying can lead to adverse consequences such as anxiety, depression, harassment, and even suicide among the individuals in particular youths. Therefore, researches has been conducted to cope with this horrible incident in these platforms so much so that to control users’ activities and detect bullying or harassment language automatically. To do so, studies try to employ statistical and computational algorithms such machine  \nlearning. To apply these methods, however, the collec","cbCairVR2jswopWq","https://ap.wps.com/l/cbCairVR2jswopWq","pdf",694559,1,6,"English","en",105,"# Introduction\n## Related Works\n## Data Sets and Preparation\n## Feature Extraction Methods\n## Model Implementation (Traditional Machine Learning)\n## Conclusion","[{\"question\":\"What is the main goal of the study?\",\"answer\":\"The study aims to detect cyberbullying automatically from Twitter text and compare machine learning approaches for performance.\"},{\"question\":\"Which machine learning models are used?\",\"answer\":\"It evaluates Logistic Regression, Multinomial Naive Bayes, K-Nearest Neighbour, and Extreme Gradient Boosting on a Twitter textual dataset.\"},{\"question\":\"How are tweets converted into model inputs?\",\"answer\":\"Tweets are transformed into numerical features using feature extraction methods such as Bag of Words and TF-IDF, enabling traditional classifiers to learn patterns from text.\"}]","Combating Cyberbullying on Social Media - A Machine Learning Approach with Text Analysis on Twitter | PDF",1785820999,15,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"combating-cyberbullying-on-social-media-a-machine-learning-approach-with-text-analysis-on-twitter","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/combating-cyberbullying-on-social-media-a-machine-learning-approach-with-text-analysis-on-twitter/124203/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the main goal of the study?","Question",{"text":75,"@type":76},"The study aims to detect cyberbullying automatically from Twitter text and compare machine learning approaches for performance.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which machine learning models are used?",{"text":80,"@type":76},"It evaluates Logistic Regression, Multinomial Naive Bayes, K-Nearest Neighbour, and Extreme Gradient Boosting on a Twitter textual dataset.",{"name":82,"@type":73,"acceptedAnswer":83},"How are tweets converted into model inputs?",{"text":84,"@type":76},"Tweets are transformed into numerical features using feature extraction methods such as Bag of Words and TF-IDF, enabling traditional classifiers to learn patterns from text.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]