[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-124214-en":3,"doc-seo-124214-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},124214,1099514067415,"Rowan","https://ap-avatar.wpscdn.com/avatar/100002539d78ffe74a7?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779092875211072502",6,"Technology","ALERT-Transformer - Bridging Asynchronous and Synchronous Machine Learning for Real-Time Event-based Spatio-Temporal Data","This work enables classic processing of continuous ultra-sparse spatiotemporal streams from event-based sensors using dense machine learning. The proposed hybrid pipeline combines asynchronous sensing with synchronous feature extraction, including an ALERT PointNet-style embedding that continuously updates via a leakage mechanism, a flexible readout that provides always up-to-date features at any sampling rate, and a patch-based Vision Transformer inspired method to exploit sparsity efficiently. A transformer trained for object and gesture recognition attains state-of-the-art performance with lower latency, while the asynchronous model operates at arbitrary desired sampling rates.","ALERT-Transformer: Bridging Asynchronous and Synchronous Machine Learning for Real-Time Event-based Spatio-Temporal Data  \nCarmen Martin-Turrero * 1 2 Maxence Bouvier * 1 Manuel Breitenstein 1 Pietro Zanuttigh 2 Vincent Parret 1  \narXiv :2402 .01393v3 [ cs .CV] 30 Jul 2024  \nAbstract  \nWe seek to enable classic processing of continuous ultra-sparse spatiotemporal data generated by event-based sensors with dense machine learning models. We propose a novel hybrid pipeline composed of asynchronous sensing and synchronous processing that combines several ideas: (1) an embedding based on PointNet models – the ALERT module – that can continuously integrate new and dismiss old events thanks to a leakage mechanism,(2) a flexible readout of the embedded data that allows to feed any downstream model with always up-to-date features at any sampling rate,(3) exploiting the input sparsity in a patch-based approach inspired by Vision Transformer to optimize the efficiency of the method. These embeddings are then processed by a transformer model trained for object and gesture recognition. Using this approach, we achieve performances at the state-of-the-art with a lower latency than competitors. We also demonstrate that our asynchronous model can operate at any desired sampling rate.  \n1. Introduction  \nEvent-based sensors capture visual information in an eventdriven, asynchronous manner (Finateu et al., 2020 ; Gallego et al., 2020) . Efficiently exploiting their data has proven challenging as the vast majority of approaches published in the literature consist of either converting event-based data to dense representations, or deploying spiking neural networks (SNNs) on streams of events. The former allows to exploit standard machine learning (ML) frameworks such as PyTorch and Tensorflow, but does not leverage the inherent sparsity and other properties of event-based data  \n*Equal contribution 1 Sony Semiconductor Solutions Europe, Sony Europe B.V, Stuttgart Laboratory 1, Zurich, Switzerland 2University of Padova, MEDIA Lab, Veneto, Italy. Correspondence to: Vincent Parret \u003C[vincent.parret@sony.com](vincent.parret@sony.com) >.  \nProceedings of the 41 st International Conference on Machine Learning, Vienna, Austria. PMLR 235, 2024 . Copyright 2024 by the author(s) .  \n(Gehrig et al., 2019) . The latter relies on SNNs, which are hard to train and usually exhibit lower accuracy than an equivalent dense neural network. Furthermore, while the neuromorphic community has argued in favor of their higher energy efficiency for decades, recent research and breakthroughs in edge AI accelerators indicate this is still an open question (Dampfhoffer et al., 2023 ; Garrett et al., 2023 ; Moosmann et al., 2023 ; Caccavella et al., 2023) .  \nNevertheless, considering the inherent advantages of eventbased vision sensors, namely high dynamic range (HDR) and high temporal resolution – simultaneously, without any tradeoffs between the two –, we aim to find a way to leverage this sparse and low-latency data for real-world situations.  \nStandard ML relies on tensor-based processing. Converting the stream of events – represented as tuples of values (x andy pixel coordinates, polarities and timestamp)– to a multidimensional tensor is thus a crucial step. The challenge involves (1) representing time in a reliable and continuous manner, allowing it to be processed similarly to the finite spatial and polarity dimensions,(2) continuously incorporating new events in the feature tensors which also requires forgetting previous events,(3) using limited computational resources to allow real-time processing. Our main contributions towards Event-Based ML are the following:  \n• The ALERT module, an embedding based on PointNet which continuously integrates new events dismissing old ones via a leakage mechanism. This module introduces novel asynchronous embedding updates.  \n• A flexible readout of the embedded data that can feed any downstream model with up-to-date features at diffe","cbCainvXG9pd9Xj0","https://ap.wps.com/l/cbCainvXG9pd9Xj0","pdf",4279044,1,18,"English","en",105,"# Introduction\n## Related Works\n### Event-Based Data Representations","[{\"question\":\"What problem does ALERT-Transformer address for event-based sensors?\",\"answer\":\"It targets efficient processing of continuous ultra-sparse spatiotemporal data produced asynchronously by event-based sensors, avoiding limitations of converting to dense tensors or relying only on spiking neural networks.\"},{\"question\":\"What are the key components of the proposed hybrid pipeline?\",\"answer\":\"The pipeline includes an ALERT embedding with continuous asynchronous updates via a leakage mechanism, a flexible readout for up-to-date features at any sampling rate, and a patch-based approach inspired by Vision Transformer to leverage input sparsity efficiently.\"},{\"question\":\"How does the method support both low latency and high accuracy?\",\"answer\":\"It uses an ALERT-Transformer framework that can run synchronously for high-accuracy gesture recognition or asynchronously for ultra-low-latency decision-making, with a seamless interface between the two regimes.\"}]","ALERT-Transformer - Bridging Asynchronous and Synchronous Machine Learning for Real-Time Event-based Spatio-Temporal Data | PDF",1785821051,45,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"alert-transformer-bridging-asynchronous-and-synchronous-machine-learning-for-real-time-event-based-spatio-temporal-data","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/technology/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/alert-transformer-bridging-asynchronous-and-synchronous-machine-learning-for-real-time-event-based-spatio-temporal-data/124214/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does ALERT-Transformer address for event-based sensors?","Question",{"text":75,"@type":76},"It targets efficient processing of continuous ultra-sparse spatiotemporal data produced asynchronously by event-based sensors, avoiding limitations of converting to dense tensors or relying only on spiking neural networks.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What are the key components of the proposed hybrid pipeline?",{"text":80,"@type":76},"The pipeline includes an ALERT embedding with continuous asynchronous updates via a leakage mechanism, a flexible readout for up-to-date features at any sampling rate, and a patch-based approach inspired by Vision Transformer to leverage input sparsity efficiently.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the method support both low latency and high accuracy?",{"text":84,"@type":76},"It uses an ALERT-Transformer framework that can run synchronously for high-accuracy gesture recognition or asynchronously for ultra-low-latency decision-making, with a seamless interface between the two regimes.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,113,118,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":111,"slug":112},50,"technology",{"id":114,"doc_module":4,"doc_module_name":46,"category_name":115,"show_sort_weight":116,"slug":117},7,"Healthcare",40,"healthcare",{"id":119,"doc_module":4,"doc_module_name":46,"category_name":120,"show_sort_weight":121,"slug":122},8,"Research & Report",30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]