[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83447-en":3,"doc-seo-83447-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83447,1099513958607,"Jiven","https://ap-avatar.wpscdn.com/avatar/100002390cf8733938c?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778829742770036399",8,"Research & Report","Hate Speech Detection in Turkish and Arabic: A Comprehensive Study","Online hate speech is tied to growing violence against minorities, creating a persistent need to reconcile freedom of expression with effective moderation on widely used social platforms. The study introduces a comprehensive hate speech dataset spanning five Turkish topics (refugees, Israel–Palestine conflict, anti-Greek sentiment, multiple ethnic/religious groups, and LGBTI+) and one Arabic topic (refugees). It also presents state-of-the-art BERT-based models supporting hate category classification, intensity prediction, target identification, and hate span detection for deeper analysis of online discourse.","arXiv :2607 .00 143v2 [ cs .CL] 5 Jul 2026  \nHate Speech Detection in Turkish and Arabic: A Comprehensive Study  \nSomaiyeh DEHGHAN 1 ,2 ∗ , G¨ok¸ce ULUDO˘GAN3, Mehmet Umut S¸EN 1 ,2,  \nElif EROL4, Arzucan ¨OZG¨UR3, Berrin YANIKOGLU 1 ,2  \n1 Department of Computer Engineering, Sabanci University, Istanbul, Turkey 34956  \n2 Center of Excellence in Data Analytics (VERIM), Sabanci University, Istanbul, Turkey 34956  \n3 Department of Computer Engineering, Bogazici University, Istanbul, Turkey 34342  \n4 Hrant Dink Foundation, Istanbul, Turkey 34373  \nA PREPRINT  \nAbstract: Online hate speech has been linked to a global rise in violence against minorities, including incidents such as mass shootings, lynchings, and ethnic cleansing. Societies grappling with this issue, particularly when hate speech targets specific groups based on religion, race, ethnicity, culture, nationality, or migration status, face the challenge of balancing freedom of expression with the need for effective content moderation on widely used online platforms. In response to this challenge, we introduce a comprehensive hate speech dataset covering five distinct topics in Turkish: refugees, the Israel-Palestine conflict, anti-Greek sentiment in Turkey, ethnic or religious communities (Alevis, Armenians, Arabs, Jews, and Kurds), and LGBTI+, alongside one topic in Arabic (refugees) . In addition, we develop state-of-the-art BERT-based models to address multiple dimensions of hate speech analysis, including hate category classification, hate intensity prediction, target identification, and hate speech span detection, enabling a comprehensive understanding of hateful content in online discourse.  \nKey words: Hate Speech Detection, Hate Intensity Prediction, Hate Speech Target Identification, Hate Speech Span Detection, NLP, BERT, ChatGPT  \nDisclaimer:  \nThis study contains examples of offensive language and hate speech for research purposes. These examples do not reflect the authors’ views and are included solely to support the detection and prevention of harmful content targeting vulnerable communities.  \n1. Introduction  \nWith the widespread use of social media, online platforms have increasingly become spaces where hate speech can spread rapidly [1] . Such content fosters hostility and intolerance and, in some cases, contributes to realworld violence targeting religious, racial, ethnic, and gender-based groups [2] . As a result, detecting and mitigating hate speech has become a critical challenge for both online platforms and policymakers. To address this issue, researchers have increasingly turned to Natural Language Processing (NLP) techniques, which enable the automatic identification and analysis of hateful content in large volumes of social media text [3–5] .  \n∗ Correspondence: [so.dehghan87@gmail.com](so.dehghan87@gmail.com)  \n1  \nDEHGHAN et al. / A PREPRINT  \nHate speech detection is a challenging task due to both the complexity of defining hate speech and the nuances of language. In recent years, numerous studies have focused on developing automatic methods for detecting hate speech in social media. However, there is a limited amount of research on hate speech detection in Turkish and Arabic. Early approaches to hate speech detection focused on using lexicons of manually chosen hateful keywords [6, 7] . However, this technique tends to be limited in effectiveness, as hate speech is not always overt. More recent research in the field, particularly for English, has shifted toward n-grams, TF-IDF, and word embedding techniques (e.g. Word2Vec, GLoVe) [8–10] .  \nMore recent advancements leverage large language models (e.g. , BERT, RoBERTa, ConvBERT, mBERT, and XLM-R) for hate speech detection [11–18] . Additionally, Most research in this domain adopts a binary approach to hate speech classification. However, contemporary studies have recognized the limitations of this approach, prompting a shift towards multi-class classification to gain a better understanding o","cbCaihm10GX76vIT","https://ap.wps.com/l/cbCaihm10GX76vIT","pdf",480731,3,1,16,"English","en",105,"# Introduction\n## Motivation and challenge of hate speech detection\n## NLP approaches and language-specific limitations\n## Shift from binary to multi-class analysis\n## Hate intensity prediction\n## Target identification and dataset gaps\n## Hate span detection and moderation needs","[{\"question\":\"Why is hate speech detection important for online platforms and policymakers?\",\"answer\":\"Online platforms enable hate speech to spread rapidly, fostering hostility and intolerance and sometimes contributing to real-world violence against religious, racial, ethnic, and gender-based groups.\"},{\"question\":\"What dataset does the study introduce, and which topics does it cover?\",\"answer\":\"It introduces a comprehensive hate speech dataset with five topics in Turkish (refugees, Israel–Palestine conflict, anti-Greek sentiment in Turkey, ethnic or religious communities, and LGBTI+) and one topic in Arabic (refugees).\"},{\"question\":\"Which tasks do the proposed BERT-based models support beyond basic hate/non-hate classification?\",\"answer\":\"They support hate category classification, hate intensity prediction, hate speech target identification, and hate speech span detection to provide more granular analysis of hateful content.\"}]",1784187889,40,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"hate-speech-detection-in-turkish-and-arabic-a-comprehensive-study","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/hate-speech-detection-in-turkish-and-arabic-a-comprehensive-study/83447/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-21","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is hate speech detection important for online platforms and policymakers?","Question",{"text":75,"@type":76},"Online platforms enable hate speech to spread rapidly, fostering hostility and intolerance and sometimes contributing to real-world violence against religious, racial, ethnic, and gender-based groups.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What dataset does the study introduce, and which topics does it cover?",{"text":80,"@type":76},"It introduces a comprehensive hate speech dataset with five topics in Turkish (refugees, Israel–Palestine conflict, anti-Greek sentiment in Turkey, ethnic or religious communities, and LGBTI+) and one topic in Arabic (refugees).",{"name":82,"@type":73,"acceptedAnswer":83},"Which tasks do the proposed BERT-based models support beyond basic hate/non-hate classification?",{"text":84,"@type":76},"They support hate category classification, hate intensity prediction, hate speech target identification, and hate speech span detection to provide more granular analysis of hateful content.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":29,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]