[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85879-en":3,"doc-seo-85879-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},85879,687197207639,"Asher","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Structured Evidence Selection for Weakly Supervised Video Anomaly Detection","Weakly supervised video anomaly detection depends only on video-level labels, which limits accurate localization of anomalous events in complex scenes. Real-world anomalies vary greatly in appearance and temporal duration, while scene appearance and action dynamics are tightly intertwined, causing models to exploit scene statistical cues rather than true behavioral deviations. SESAD reformulates detection as structured reasoning over clip-level visual evidence by selecting candidate evidence under scene and action constraints. It also adds a lightweight geometric discrimination module with a dual-prototype embedding design, improving stability and achieving 67.92/97.99/88.46 AUC on UBnormal, ShanghaiTech, and UCF-Crime.","Structured Evidence Selection for Weakly Supervised Video  \nAnomaly Detection  \nChenglizhao Chen  \nChina University of Petroleum (East China) Qingdao, China  \nTianxiang Nan  \nChina University of Petroleum (East China) Qingdao, China  \nWen Li  \nChina University of Petroleum (East China) Qingdao, China  \nXinyu Liu∗ China University of Petroleum (East China) Qingdao, China [liuxy001005@163.com](liuxy001005@163.com)  \nGuisheng Zhang  \nChina University of Petroleum (East China) Qingdao, China  \nMengke Song  \nChina University of Petroleum (East China) Qingdao, China  \nXiaomin Yu  \nThe Hong Kong University of Science and Technology (Guangzhou) Guangzhou, China  \narXiv :2607 . 10298v1 [ cs .CV] 11 Jul 2026  \nAbstract  \nWeakly supervised video anomaly detection relies solely on videolevel labels for training, making it difficult to accurately localize anomalous events in complex scenes. In real-world videos, anomalous behaviors exhibit large variations in appearance and temporal duration, while scene appearance and action dynamics are often tightly entangled. Consequently, existing models tend to rely on scene-related statistical cues rather than true behavioral deviations, resulting in unstable detection performance. To address this challenge, we propose a Structured Evidence Selection framework (SESAD) that reformulates anomaly detection as a structured reasoning process over clip-level visual evidence. Instead of directly mapping aggregated features to anomaly scores, SESAD reorganizes clip representations into semantically structured candidate evidence and performs context-conditioned selection under scene and action constraints. This mechanism adaptively emphasizes anomaly-relevant semantics while suppressing scene interference, thereby alleviating semantic entanglement under weak supervision. Furthermore, we introduce a lightweight geometric discrimination module that constructs a dual-prototype structure in the embedding space, enabling anomaly decisions through relative geometric relations. Extensive experiments on UBnormal, ShanghaiTech, and UCF-Crime show that SESAD achieves 67.92, 97.99, and 88.46 AUC, respectively, while maintaining high computational efficiency and overall consistently stable anomaly discrimination.  \nKeywords  \nWeakly Supervised Video Anomaly Detection, Structured Evidence Selection, Multi-Perspective Representation  \n1 Introduction  \nVideo anomaly detection aims to identify events that deviate from normal behavior patterns in complex real-world environments [19, 25, 28] . It plays an important role in intelligent surveillance, public safety, and video understanding [27] . In practice, anomalous events are rare and diverse, making precise temporal annotations very costly [33] . To reduce the annotation burden, weakly supervised  \n∗ Corresponding author.  \nFigure 1: Motivation and overview of our method. Existing WS-VAD methods compress heterogeneous cues into a single embedding for anomaly regression, leading to scene bias and semantic entanglement. Our approach instead performs structured evidence selection guided by scene and action queries to achieve more robust anomaly detection.  \nvideo anomaly detection (WS-VAD) has attracted increasing attention. Under this setting, models are trained using only video-level labels without clip-level annotations, which significantly increases the difficulty of reliable anomaly localization [58] .  \nAs illustrated in Figure 1(A), existing weakly supervised video anomaly detection methods can be broadly categorized into two research directions. The first direction introduces additional modalities, such as optical flow and skeleton information, and leverages multimodal fusion to enhance motion cues or structural information, thereby improving the discriminability of anomalous behaviors [13, 31, 39, 43] . The second direction focuses on representation modeling within RGB videos, aiming to improve feature expressiveness by designing stronger backbone networks, temporal m","cbCairy8H35fvX0n","https://ap.wps.com/l/cbCairy8H35fvX0n","pdf",3636315,1,10,"English","en",105,"# 1 Introduction\n## Motivation and problem setting\n## Related directions and modeling paradigm\n## Limitations of single-embedding regression\n# Proposed method overview\n## Structured evidence reasoning concept\n## SESAD framework components","[{\"question\":\"Why is temporal localization difficult in weakly supervised video anomaly detection?\",\"answer\":\"Training uses only video-level labels without clip-level annotations, making it hard to reliably pinpoint anomalous events in time.\"},{\"question\":\"What causes unstable anomaly predictions in existing weakly supervised methods?\",\"answer\":\"Scene appearance and motion dynamics, together with anomaly cues, are compressed into a single embedding, leading models to depend on scene statistics and mix anomaly signals with irrelevant patterns.\"},{\"question\":\"How does SESAD address scene bias and semantic entanglement?\",\"answer\":\"SESAD reorganizes clip features into semantically structured, multi-perspective evidence and performs context-conditioned selection under scene and action constraints, then suppresses scene interference before the final decision.\"}]",1784206911,25,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"structured-evidence-selection-for-weakly-supervised-video-anomaly-detection","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/structured-evidence-selection-for-weakly-supervised-video-anomaly-detection/85879/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is temporal localization difficult in weakly supervised video anomaly detection?","Question",{"text":75,"@type":76},"Training uses only video-level labels without clip-level annotations, making it hard to reliably pinpoint anomalous events in time.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What causes unstable anomaly predictions in existing weakly supervised methods?",{"text":80,"@type":76},"Scene appearance and motion dynamics, together with anomaly cues, are compressed into a single embedding, leading models to depend on scene statistics and mix anomaly signals with irrelevant patterns.",{"name":82,"@type":73,"acceptedAnswer":83},"How does SESAD address scene bias and semantic entanglement?",{"text":84,"@type":76},"SESAD reorganizes clip features into semantically structured, multi-perspective evidence and performs context-conditioned selection under scene and action constraints, then suppresses scene interference before the final decision.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":21,"slug":133},"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]