[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82202-en":3,"doc-seo-82202-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82202,2336464648746,"Skyler","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Event Stream Based Multi-Modal Video Anomaly Detection Benchmark Dataset and Algorithms","Video anomaly detection (VAD) for surveillance remains unreliable under illumination changes, fast motion, and complex backgrounds when it depends only on visible-light video. E-VAD introduces an event-enhanced framework that fuses conventional video with bio-inspired event-camera streams. Event sensors asynchronously record brightness changes with high temporal resolution, improving robustness to motion blur and extreme lighting while supplying motion-salient cues. A large visible–event benchmark (6.3B events, 376,368 frames) supports scalable research, and contrastive multi-modal pretraining plus adaptive fusion improves performance. Experiments show E-VAD consistently outperforms prior methods, validating event sensing for real-world VAD.","Event Stream based Multi-Modal Video Anomaly Detection: A Benchmark Dataset and Algorithms  \nPeipei Zhu*, Yueqing Niu*, Lin Zhu, Guanchong Niu, Yang Yu, Zheng Li  \narXiv :2607 .09114v1 [ cs .CV] 10 Jul 2026  \nAbstract—Video anomaly detection (VAD) is critical for automated surveillance but remains fragile under challenging conditions such as illumination variations, fast motion, and complex backgrounds when relying solely on visible-light videos. To address these limitations, we propose E-VAD, an event-enhanced VAD framework that jointly exploits conventional video and event streams captured by bio-inspired event cameras. Event sensors asynchronously capture brightness changes with high temporal resolution, offering robustness to motion blur and extreme lighting, and providing motion-salient cues complementary to videobased visual information. To support multi-modal VAD research, we construct a large-scale visible–event benchmark comprising 6.3 billion events and 376,368 video frames collected under diverse illumination levels, motion patterns, and background complexities, filling the gap of realistic and scalable datasets for event-based anomaly detection. Building upon this dataset, we design a contrastive multi-modal pretraining framework to learn discriminative event representations by aligning semantic embeddings across event streams, visible videos, and textual descriptions. An adaptive fusion module then dynamically integrates event-based temporal cues with video-based spatial semantics, improving robustness to environmental disturbances. Experiments on benchmarks and the proposed TJUTCM Pha dataset demonstrate that E-VAD consistently outperforms methods, validating the effectiveness of event-based sensing for VADin real-world scenarios.  \nIndex Terms—Video anomaly detection, Event camera, Multimodal learning, Feature fusion  \nI. INTRODUCTION  \nVIdeo anomaly detection (VAD) plays a central role  \nin intelligent surveillance, autonomous perception, and industrial safety, enabling automatic identification of irregular patterns in continuous video streams [1], [2] . As VAD systems are increasingly deployed in manufacturing [3]–[5], transportation [6], and public security infrastructures [7], they are expected to remain reliable under fluctuating illumination, rapid motion, and operational complexity. Benefiting from  \nPeipei Zhu, Yueqing Niu, and Yang Yu are with the College of Pharmaceutical Engineering of Traditional Chinese Medicine, Tianjin University of Traditional Chinese Medicine, Tianjin 301617, China (e-mail: [zhupp0527@tjutcm.edu.cn](zhupp0527@tjutcm.edu.cn), [1035415060@qq.com](1035415060@qq.com), andyu [yang@tjutcm.edu.cn](yang@tjutcm.edu.cn)).  \nLin Zhu is with the School of Artificial Intelligence, Beijing Normal University, Beijing 100875, China. (e-mail: [linzhu@pku.edu.cn](linzhu@pku.edu.cn)).  \nGuanchong Niu is with the Guangzhou Institute of Technology, Xidian University, Guangzhou 510555, China. (e-mails: [niuguanchong@xidian.edu.cn](niuguanchong@xidian.edu.cn)).  \nZheng Li is with the College of Pharmaceutical Engineering of Traditional Chinese Medicine, Tianjin University of Traditional Chinese Medicine, Tianjin 301617, China, and also with State Key Laboratory of Component-based Chinese Medicine, Tianjin University of Traditional Chinese Medicine, Tianjin 301617, China (e-mail: [lizheng@tjutcm.edu.cn](lizheng@tjutcm.edu.cn)).  \nCorresponding author: Zheng Li. *The authors have contributed equally to this work.  \nFig. 1. Comparison of sensing characteristics between visible-light cameras and event cameras in VAD. Subfigures (a–d) illustrate scenarios where event streams provide clearer anomaly cues than conventional video, including extreme illumination, high-speed motion, and both simple and cluttered environments. In these settings, the sparse, change-driven responses of event sensors suppress static background interference and enhance the signal-tonoise ratio of anomalous behaviors, enabling more r","cbCaihimVyApEsF9","https://ap.wps.com/l/cbCaihimVyApEsF9","pdf",6787721,2,1,14,"English","en",105,"# Introduction\n## Problem: visible-only VAD limitations\n## Motivation: event cameras for robust sensing\n## Proposed approach and contributions","[{\"question\":\"What challenge does event-enhanced VAD address compared with visible-only video?\",\"answer\":\"Visible-only VAD degrades under illumination fluctuations, fast motion blur, and low signal-to-noise conditions, which harms precise anomaly perception. E-VAD adds event streams that remain robust under these sensing constraints.\"},{\"question\":\"What is E-VAD’s core idea for multi-modal anomaly detection?\",\"answer\":\"E-VAD jointly exploits visible video and event streams from bio-inspired event cameras. It uses contrastive multi-modal pretraining to align event/video/text semantics and an adaptive fusion module to combine event temporal cues with video spatial semantics.\"},{\"question\":\"What resources does the paper contribute for research on event-based VAD?\",\"answer\":\"It constructs a large visible–event benchmark dataset with 6.3 billion events and 376,368 video frames collected under diverse illumination, motion, and background conditions, plus associated algorithms and evaluation on benchmarks including the TJUTCM Pha dataset.\"}]",1784178791,35,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"event-stream-based-multi-modal-video-anomaly-detection-benchmark-dataset-and-algorithms","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/event-stream-based-multi-modal-video-anomaly-detection-benchmark-dataset-and-algorithms/82202/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-21","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What challenge does event-enhanced VAD address compared with visible-only video?","Question",{"text":75,"@type":76},"Visible-only VAD degrades under illumination fluctuations, fast motion blur, and low signal-to-noise conditions, which harms precise anomaly perception. E-VAD adds event streams that remain robust under these sensing constraints.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is E-VAD’s core idea for multi-modal anomaly detection?",{"text":80,"@type":76},"E-VAD jointly exploits visible video and event streams from bio-inspired event cameras. It uses contrastive multi-modal pretraining to align event/video/text semantics and an adaptive fusion module to combine event temporal cues with video spatial semantics.",{"name":82,"@type":73,"acceptedAnswer":83},"What resources does the paper contribute for research on event-based VAD?",{"text":84,"@type":76},"It constructs a large visible–event benchmark dataset with 6.3 billion events and 376,368 video frames collected under diverse illumination, motion, and background conditions, plus associated algorithms and evaluation on benchmarks including the TJUTCM Pha dataset.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]