[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-122668-en":3,"doc-seo-122668-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},122668,549758146520,"Patrick","https://ap-avatar.wpscdn.com/avatar/80002397d8c0411e94?_k=1775819394049821470",8,"Research & Report","Multi-modal Hate Speech Detection using Machine Learning","Internet users and rich media growth make tracking hateful speech in audio and video difficult. Simple audio/video-to-text conversion can miss hate, since people may use hateful words humorously or pleasantly and vary voice tones and visible actions. Most existing hate-speech systems rely on a single modality. This work proposes a multimodal machine-learning and NLP approach that extracts features from video images, audio values, and accompanying text, then combines model outputs with a hard-voting ensemble for a final decision.","Multi-modal Hate Speech Detection using Machine  \nLearning  \nFariha Tahosin Boishakhi  \nComputer Science and Engineering BRAC University Dhaka , Bangladesh [fariha.tahosin.boishakhi@g.bracu.ac.bd](fariha.tahosin.boishakhi@g.bracu.ac.bd)  \nPonkoj Chandra Shill  \nComputer Science and Engineering BRAC University Dhaka, Bangladesh [ponkoj.chandra.shill@g.bracu.ac.bd](ponkoj.chandra.shill@g.bracu.ac.bd)  \nMd. Golam Rabiul Alam  \nComputer Science and Engineering BRAC University Dhaka , Bangladesh [rabiul.alam@bracu.ac.bd](rabiul.alam@bracu.ac.bd)  \narXiv :2307 . 1 15 19v 1 [ cs .AI] 15 Jun 2023  \nAbstract—With the continuous growth of internet users and media content, it is very hard to track down hateful speech in audio and video. Converting video or audio into text does not detect hate speech accurately as human sometimes uses hateful words as humorous or pleasant in sense and also uses different voice tones or show different action in the video. The state-ofthe-art hate speech detection models were mostly developed on a single modality. In this research, a combined approach of multimodal system has been proposed to detect hate speech from video contents by extracting feature images, feature values extracted from the audio, text and used machine learning and Natural language processing.  \nIndex Terms—Audio hate Speech, Video hate Speech, Hate Speech detection, Machine Learning, Multi-modal Hate Speech detection.  \nI. INTRODUCTION  \nIn this era of digital communications, hatred data is not only in social media comments and posts or text messages but also in voice messages and video content as well [1] . Content like this causes cyberbullying, rioting, fraud, loss of respect, and even murder. (Hate Crime Statistics, 2019) shows 15588 law enforcement agencies reported crimes, suspects, offenders, and hate crime zones. These groups reported 7314 hate crimes with 8,559 offenses. The ﬁndings include 57.6%race/ethnicity/ancestry/bias, 20.1 % religion, and 16.7 % sexual orientation. [2] In another research, online hate crimes frequently start online and affect us ofﬂine. They were also mistreated online, according to the report's victims. The research shows that hate data is linked to voice tone and facial emotions. Voice tones and facial expressions are crucial aspects to discern hatred. In some cases, only the text data, facial expressions or vocal information as audio data is not enough individually to detect the hateful conversations. Almost all current research on hate speech identiﬁcation uses text data. This study suggests integrating audio, video, and text elements to detect hate speech. The following research is implemented by taking all the modes available to deliver hate speech into account and designed to detect hate speech more precisely. The ﬁnal conclusion of hate speech is determined by merging the results of a hard voting ensemble or majority voting where we combine all model result image, audio, and text to determine the ﬁnal output.  \nThe major contributions of this research can be summarized as follows -  \n􀀏 Data has been prepared with both hateful and non-hateful speech video from different sources such as YouTube, EMBY mostly taken from movies or series. Afterwards, image, audio and text data has been extracted from the video contents.  \n􀀏 Feature Extracted from images, audio and text individually. For feature selection, Recursive Feature Selection (RFE), Maximum Relevance - Minimum Redundancy (MRMR) is used to take the most relevant features.  \n􀀏 A hate speech detection model has been developed with images, audio and text separately and then combining the results with a hard voting ensemble model to determine the ﬁnal outcome of hate speech.  \n􀀏 The performance of seven different classiﬁers has been studied on the hate speech data set to show the comparative study of different classical machine learning models in hate speech detection.  \nII. LITERATURE REVIEW  \nWarner and Hirschberg [3], for detecting hate speech fr","cbCaihhSeXQ9vMA7","https://ap.wps.com/l/cbCaihhSeXQ9vMA7","pdf",1343972,1,5,"English","en",105,"# Introduction\n# Literature Review\n# Methodology","[{\"question\":\"Why is hate speech harder to detect in audio and video than in text?\",\"answer\":\"Audio and video require interpreting voice tone, facial emotions, and actions. Converting audio/video to text alone may miss hate when words are used humorously or in pleasant contexts with changing delivery.\"},{\"question\":\"What is the core idea of the proposed system?\",\"answer\":\"The approach extracts features from multiple modalities—video images, audio, and text—then uses machine learning and NLP to determine whether the content contains hate speech.\"},{\"question\":\"How is the final hate-speech decision produced?\",\"answer\":\"Results from image, audio, and text models are merged using a hard voting ensemble (majority voting) to produce the final output.\"}]","Multi-modal Hate Speech Detection using Machine Learning | PDF",1785812077,13,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"multi-modal-hate-speech-detection-using-machine-learning","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/multi-modal-hate-speech-detection-using-machine-learning/122668/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is hate speech harder to detect in audio and video than in text?","Question",{"text":75,"@type":76},"Audio and video require interpreting voice tone, facial emotions, and actions. Converting audio/video to text alone may miss hate when words are used humorously or in pleasant contexts with changing delivery.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the core idea of the proposed system?",{"text":80,"@type":76},"The approach extracts features from multiple modalities—video images, audio, and text—then uses machine learning and NLP to determine whether the content contains hate speech.",{"name":82,"@type":73,"acceptedAnswer":83},"How is the final hate-speech decision produced?",{"text":84,"@type":76},"Results from image, audio, and text models are merged using a hard voting ensemble (majority voting) to produce the final output.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,109,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":106,"show_sort_weight":107,"slug":108},"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":21,"slug":137},19,"General","general"]