[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86277-en":3,"doc-seo-86277-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86277,1099514068365,"Aurelia","https://ap-avatar.wpscdn.com/avatar/10000253d8d9f28188e?_k=1776742907772140068",8,"Research & Report","Beyond Benchmarks: Exposing the Hidden Crisis in Bangla Hate Speech Detection","The spread of hate speech across social media platforms creates a major barrier to online safety and ethical moderation, especially for under-resourced languages like Bangla. Existing benchmark-trained systems often fail on implicit, context-dependent HS shaped by cultural nuance, informal phrasing, and emoji usage. Six deep learning models are trained on benchmark and merged multi-source datasets, then externally validated on a newly annotated real-world set to measure explicit versus implicit performance drops.","Beyond Benchmarks: Exposing the Hidden Crisis in Bangla Hate  \nSpeech Detection  \nFaria Afrin Tisha1, Fariya Tabassum1, Hafsa Binte Kibria1, Md. Nahiduzzaman1 and Mominul Ahsan2*  \n1Department of Electrical and Computer Engineering, Rajshahi University of Engineering and Technology, Rajshahi, 6204, Bangladesh.  \nemail: [1810017@student.ruet.ac.bd](1810017@student.ruet.ac.bd) (F.A.T); [tabassum.fariya27@gmail.com](tabassum.fariya27@gmail.com) (F.T); [hafsabintekibria@ece.ruet.ac.bd](hafsabintekibria@ece.ruet.ac.bd) (H.B.K);  \n[nahiduzzaman@ece.ruet.ac.bd](nahiduzzaman@ece.ruet.ac.bd) (M.N.).  \n2Department of Computer Science, University of York, Deramore Lane, Heslington, York YO10 5GH, UK; email:  \n[mominul.ahsan2@gmail.com](mominul.ahsan2@gmail.com) (M.A.).  \n*[Corresponding author: mominul.ahsan2@gmail.com](Corresponding author: mominul.ahsan2@gmail.com) (M.A.)  \nAbstract  \nThe spread of hate speech (HS) across different social media platforms (SMPs) poses a major concern for ensuring online safety and ethical moderation. Automatic detection of HS remains a challenging task, especially in under-resourced languages like Bangla, due to cultural context, implicit expressions, and informal linguistic patterns. This study aimed to expose the underlying crisis of existing Bangla HS detection systems by diagnosing how and why benchmark-trained models fail to identify implicit, context-dependent HS. Six deep learning architectures (FastText + CNN, FastText + LSTM, FastText + BiLSTM, BanglaBERT, BanglaBERT + CNN, and BanglaBERT + BiLSTM) were trained on benchmark datasets (≈75,000 posts) and a merged multi-source dataset (≈120,000 posts), then externally validated on a newly annotated real-world dataset (≈200 posts) collected from Facebook, Twitter, and YouTube, labeled as HS and non-HS, where HS was further categorized as explicit and implicit. BanglaBERT achieved an F1-score of 91.4% on benchmark datasets but declined to 75.3% on the external set and 63.4% for implicit HS involving sarcasm and emojis. The accuracy of FastText + CNN dropped from 78.0% to 51.2% under similar conditions. Emoji-aware preprocessing improved implicit HS detection by up to 12%, whereas emoji removal caused a notable decline in performance (F1: 0.75 → 0.63) . Frequent misclassifications in politically charged or satirical comments revealed over-policing risks. This study not only exposes the generalization crisis due to implicit, culturally embedded, and emoji-laden expressions but also underscores the need for developing adaptive, emoji-aware, and culturally grounded frameworks that ensure ethical moderation while preserving freedom of expression. Findings ofthis study provide actionable insights for researchers, SMPs, and policymakers to design more context-sensitive HS detection systems for low-resource languages.  \nKeywords Bangla HS, Low-Resource Language, Social Media, Deep Learning, Transformer Models, Emoji Interpretation, Context-Aware Detection  \n1 Introduction  \nOver the past decade, SMPs such as Facebook, YouTube, Twitter, and TikTok have changed the game of public communication by providing unprecedented room for freedom of expression. However, this freedom has also allowed cyberbullying (Ray et al., 2024) in the form of HS. HS refers to a direct or indirect attack on individuals based on protected qualities such as race, ethnicity, nationality, disabilities, religious affiliation, caste, gender, and serious disease, which is severe enough to cause harm by dehumanizing words or negative stereotypes (Hateful Conduct | Transparency Centre, n.d.) People or groups are targeted based on these qualities. The prevalence of social media services in Bangladesh has resulted in a disturbing rise in HS, which poses a serious threat to social harmony, personal security, and internet wellbeing (Atikuzzaman & Akter, 2023), While languages with high resources, such as English, and Arabic, have experienced tremendous success in detecting HS through machine lea","cbCaitDdxGmkx0QZ","https://ap.wps.com/l/cbCaitDdxGmkx0QZ","pdf",1777886,4,1,40,"English","en",105,"# Abstract\n# Keywords\n# 1 Introduction\n## Hate speech on social media\n## Challenges in Bangla hate speech detection\n## Benchmark generalization and implicit HS","[{\"question\":\"Why is hate speech detection especially difficult in Bangla?\",\"answer\":\"Detection is difficult because HS can be culturally subtle, highly contextual, and often implicit, expressed through sarcasm, memes, coded language, or emojis that evade keyword-based approaches.\"},{\"question\":\"How were the models evaluated in this study?\",\"answer\":\"Models were trained on benchmark datasets (and a merged multi-source dataset) and then externally validated using a newly annotated real-world dataset labeled as hate/non-hate, with hate further split into explicit and implicit.\"},{\"question\":\"What impact did emoji-aware preprocessing have on implicit hate speech detection?\",\"answer\":\"Emoji-aware preprocessing improved implicit HS detection by up to 12%, while emoji removal caused a noticeable performance decline (F1 decreased from about 0.75 to 0.63).\"}]",1784209976,101,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"beyond-benchmarks-exposing-the-hidden-crisis-in-bangla-hate-speech-detection","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/beyond-benchmarks-exposing-the-hidden-crisis-in-bangla-hate-speech-detection/86277/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is hate speech detection especially difficult in Bangla?","Question",{"text":75,"@type":76},"Detection is difficult because HS can be culturally subtle, highly contextual, and often implicit, expressed through sarcasm, memes, coded language, or emojis that evade keyword-based approaches.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How were the models evaluated in this study?",{"text":80,"@type":76},"Models were trained on benchmark datasets (and a merged multi-source dataset) and then externally validated using a newly annotated real-world dataset labeled as hate/non-hate, with hate further split into explicit and implicit.",{"name":82,"@type":73,"acceptedAnswer":83},"What impact did emoji-aware preprocessing have on implicit hate speech detection?",{"text":84,"@type":76},"Emoji-aware preprocessing improved implicit HS detection by up to 12%, while emoji removal caused a noticeable performance decline (F1 decreased from about 0.75 to 0.63).","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":22,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]