[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84770-en":3,"doc-seo-84770-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84770,4398048950312,"Violet","https://ap-avatar.wpscdn.com/avatar/400002538284de19e3c?_k=1778320343897328908",8,"Research & Report","Evaluating Large Language Models for Antisemitic Incident Classification","Automated detection of hateful events from public reports remains underexplored despite the need for timely identification to support monitoring and protection of targeted communities. This work introduces hateful event detection with fine-grained labels and evaluates large language models on antisemitic event descriptions drawn from news articles, civil society reports, and official records. Experiments compare GPT-4o and Llama-3 3B-Instruct, showing promise but requiring substantial improvement. Clear term definitions and in-context examples in prompts enhance performance, with definitions benefiting rhetoric-focused cases and examples aiding action-focused incidents. A college-news case study indicates LLMs can surface relevant real-world events for early monitoring and intervention, while revealing critical evaluation and policy gaps.","arXiv :2607 .04890v 1 [ cs .CL] 6 Jul 2026  \nEvaluating Large Language Models for Antisemitic Incident Classification  \nKarina Halevy1 , Julia Mendelsohn2 , Chan Young Park3 , Yulia Tsvetkov4 , Maarten Sap1,5  \n1 Carnegie Mellon University, 2University of Maryland, 3Microsoft Research,  \n4University of Washington, 5Allen Institute for Artificial Intelligence  \nCorrespondence: [khalevy@andrew.cmu.edu](khalevy@andrew.cmu.edu)  \nAbstract  \nAddressing hate and violence in society requires timely detection of hateful events from public reporting, but automated identification of hateful events remains underexplored. We introduce the task of hateful event detection and investigate the ability of AI systems, specifically large language models (LLMs), to discover and classify reports of antisemitic events with fine-grained labels. We evaluate OpenAI’s GPT-4o and Meta’s Llama-3 .2-3B-Instruct on multiple expert-annotated datasets containing antisemitic event descriptions from news articles, civil society reports, and official records. We show that LLMs, particularly GPT-4o, have potential for this task, but substantial improvement is needed. Providing clear term definitions and in-context examples in prompts can improve performance: definitions are most helpful for rhetoric-oriented events (e.g. classical antisemitic tropes), while examples help label action-oriented events (e.g. physical assault) . A case study of college newspapers demonstrates that LLMs can help surface relevant real-world events, supporting early monitoring and intervention. Overall, our findings highlight both opportunities and critical gaps in AI’s ability to recognize complex harms and underscore the need for collaborative efforts among AI developers, policymakers, and civil society to design models, implement robust evaluation, and develop policy frameworks for defining and combating hate efficiently and effectively.  \n1 Introduction  \nHate and violence in society are not only individual tragedies but also indicators of broader social harm. Detecting hateful events—broadly defined as crimes, threats, or encouragement of crimes motivated by bias—in descriptions from news articles, civil society reports, and official records is crucial for monitoring societal trends and protecting targeted communities (U.S. Department of Justice, 2024) . However, the growing volume of such reports makes comprehensive, timely monitoring increasingly difficult to carryout by human analysts alone, and there is a clear need for automated tools.  \nPrior computational approaches for analyzing real-world hate have largely focused on detecting hateful or toxic language, typically in social media posts. While valuable, these approaches are limited for jointly monitoring online and offline harm—they center on speech rather than events. This excludes many forms of bias-motivated harm such as physical violence, vandalism, and discrimination, which may contain no explicitly hateful language.  \nThis work addresses these limitations by introducing the novel task of hateful event detection—characterizing both online and offline hate incidents as described in textual reports (e.g. distilling mentions of hateful events from streams of social media posts, press releases, news articles) . This task is conceptually distinct from, yet complementary to, hate speech detection, which focuses on explicitly hateful language. This task moreover emphasizes fine-grained classification, requiring computational models to distinguish among specific forms of harm rather than solely assigning overly broad labels—coarse labels obscure important distinctions among different types of targets and harm, limiting their utility for practitioners who must decide when, where, and how to intervene. This task facilitates a structured and actionable representation of real-world hate that better aligns with the goals of policymakers, educators, and civil society organizations. Stakeholders such as NGOs, journalists, law enforce","cbCaijZn7Sila13H","https://ap.wps.com/l/cbCaijZn7Sila13H","pdf",1467157,1,25,"English","en",105,"# Introduction\n## Hate and violence monitoring\n## Limits of toxicity and hate-speech approaches\n## Novel task: hateful event detection\n## Case study focus on antisemitism","[{\"question\":\"What problem does the document address in detecting hate?\",\"answer\":\"It targets the underexplored task of automatically identifying hateful events from public reports in a timely way, since current approaches often focus on hateful or toxic language rather than events.\"},{\"question\":\"What models and datasets are evaluated for antisemitic incident classification?\",\"answer\":\"The evaluation compares OpenAI GPT-4o and Meta Llama-3 3B-Instruct on multiple expert-annotated datasets built from antisemitic event descriptions in news articles, civil society reports, and official records.\"},{\"question\":\"How can prompting improve performance for fine-grained event labeling?\",\"answer\":\"Providing clear term definitions and in-context examples improves results: definitions are most helpful for rhetoric-oriented events, while examples are especially helpful for action-oriented events like physical assault.\"}]",1784198128,63,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"evaluating-large-language-models-for-antisemitic-incident-classification","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/evaluating-large-language-models-for-antisemitic-incident-classification/84770/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the document address in detecting hate?","Question",{"text":75,"@type":76},"It targets the underexplored task of automatically identifying hateful events from public reports in a timely way, since current approaches often focus on hateful or toxic language rather than events.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What models and datasets are evaluated for antisemitic incident classification?",{"text":80,"@type":76},"The evaluation compares OpenAI GPT-4o and Meta Llama-3 3B-Instruct on multiple expert-annotated datasets built from antisemitic event descriptions in news articles, civil society reports, and official records.",{"name":82,"@type":73,"acceptedAnswer":83},"How can prompting improve performance for fine-grained event labeling?",{"text":84,"@type":76},"Providing clear term definitions and in-context examples improves results: definitions are most helpful for rhetoric-oriented events, while examples are especially helpful for action-oriented events like physical assault.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]