[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83429-en":3,"doc-seo-83429-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83429,7971461741311,"Ophelia","https://ap-avatar.wpscdn.com/avatar/74000253aff267980c6?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779345379180704826",8,"Research & Report","LongE2V Long-Horizon Event-based Video Reconstruction Prediction and Frame Interpolation with Video Diffusion Models","LongE2V tackles long-horizon event-based video generation from sparse event streams by unifying video reconstruction, video prediction, and frame interpolation within a single conditional video diffusion framework. The method leverages pre-trained video diffusion priors and fine-tunes a foundational model to recover photometric textures, generate long-term sequences with autoregressive unrolling to minimize temporal drift, and synthesize intermediate frames with zero-shot adaptation using event dynamics as temporal guidance. Dedicated mechanisms further enhance temporal coherence and bidirectional interpolation consistency, while event voxel density augmentation improves robustness across sensor resolutions.","arXiv :2607 .08770v 1 [ cs .CV] 9 Jul 2026  \nLongE2V: Long-Horizon Event-based Video Reconstruction, Prediction, and Frame Interpolation with Video Diffusion Models  \nCHENG-DE FAN, National Yang Ming Chiao Tung University, Taiwan CHUN-WEI TUAN MU, National Yang Ming Chiao Tung University, Taiwan CHEN-WEI CHANG, National Yang Ming Chiao Tung University, Taiwan CHIN-YANG LIN, National Yang Ming Chiao Tung University, Taiwan KUN-RU WU, National Yang Ming Chiao Tung University, Taiwan  \nYU-CHEE TSENG, National Yang Ming Chiao Tung University, Taiwan YU-LUN LIU, National Yang Ming Chiao Tung University, Taiwan  \nEvent Stream Only  \nEvent Stream + Start Frame  \nEvent Stream + Start/End Frames  \n(a) Video Reconstruction  \n(b) Video Prediction  \n(c) Video Frame Interpolation  \nHigh Fidelity  \nMinimal Drift  \nZero-Shot Adaptation  \nFig. 1. Event-based video generation. We leverage pre-trained video diffusion priors to address three distinct inverse problems within a single architecture. Depending on the input condition, our model performs: (a) Video Reconstruction, recovering high-fidelity textures from sparse event streams,(b) Video Prediction, generating long-term sequences from a single start frame with minimal drift via our autoregressive unrolling strategy, and (c) Video Frame Interpolation, achieving zero-shot adaptation to synthesize intermediate frames by leveraging event dynamics as temporal guidance.  \nRecovering high-quality video from sparse event streams is a challenging task. Regression methods often blur textures, while existing generative models struggle with long-term stability. We propose LongE2V, a novel approach that leverages pre-trained video diffusion priors to jointly handle event-based video reconstruction, prediction, and frame interpolation. By fine-tuning a foundational video model, our approach achieves high data efficiency and  \nSIGGRAPH Conference Papers ’26, Los Angeles, CA, USA © 2026 Copyright held by the owner/author(s) .  \nACM ISBN 979-8-4007-2554-8/2026/07  \n[https://doi.org/10.1145/3799902.3811151](https://doi.org/10.1145/3799902.3811151)  \nSIGGRAPH Conference Papers ’26, July 19–23, 2026, Los Angeles, CA, USA.  \nsuperior perceptual quality. We introduce Autoregressive Unrolling and Adaptive Context Switching to mitigate temporal drift in extremely long sequences. We also propose Reencoding Alignment with Cross Residual Correction to ensure precise bidirectional consistency during frame interpolation. Furthermore, Event Voxel Density Augmentation ensures robustness across varying sensor resolutions. Extensive experiments on real-world benchmarks demonstrate that LongE2V outperforms state-of-the-art methods across all three tasks, exhibiting exceptional temporal coherence and zero-shot generalization. Project page: [https://cdfan0627.github.io/LongE2V-page/](https://cdfan0627.github.io/LongE2V-page/)  \n[CCS Concepts:](CCS Concepts:) • Computing methodologies → Reconstruction; Computational photography; Neural networks; • Hardware → Sensors and actuators.  \nACM Reference Format:  \nCheng-De Fan, Chun-Wei Tuan Mu, Chen-Wei Chang, Chin-Yang Lin, KunRu Wu, Yu-Chee Tseng, and Yu-Lun Liu. 2026. LongE2V: Long-Horizon Event-based Video Reconstruction, Prediction, and Frame Interpolation with Video Diffusion Models. In Special Interest Group on Computer Graphics and Interactive Techniques Conference Conference Papers (SIGGRAPH Conference Papers ’26), July 19–23, 2026, Los Angeles, CA, USA. ACM, New York, NY, USA, 14 pages. [https://doi.org/10.1145/3799902.3811151](https://doi.org/10.1145/3799902.3811151)  \n1 Introduction  \nEvent cameras are bio-inspired sensors capturing asynchronous brightness changes with microsecond resolution and high dynamic range (HDR). Unlike standard cameras prone to motion blur, they excel in high-speed dynamics. However, their sparse, intensity-free output is incompatible with standard vision algorithms, making high-fidelity video recovery an inherently ill-posed problem. We address","cbCaivAe0co4Nxqu","https://ap.wps.com/l/cbCaivAe0co4Nxqu","pdf",4992737,2,1,14,"English","en",105,"# Introduction\n## Event cameras and the ill-posed nature of event-to-video recovery\n## Limitations of existing task-specific and diffusion-based approaches\n## LongE2V overview and unified conditional framework","[{\"question\":\"What three tasks does LongE2V handle in a unified architecture?\",\"answer\":\"LongE2V jointly performs video reconstruction, long-term video prediction, and video frame interpolation within one architecture conditioned on event voxels.\"},{\"question\":\"How does LongE2V reduce temporal drift in long-horizon prediction?\",\"answer\":\"It introduces an autoregressive unrolling strategy designed to minimize drift during long-term sequence generation.\"},{\"question\":\"How does LongE2V perform zero-shot frame interpolation?\",\"answer\":\"It synthesizes intermediate frames by adapting to event dynamics as temporal guidance, enabling zero-shot behavior for interpolation.\"}]",1784187589,35,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"longe2v-long-horizon-event-based-video-reconstruction-prediction-and-frame-interpolation-with-video-diffusion-models","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/longe2v-long-horizon-event-based-video-reconstruction-prediction-and-frame-interpolation-with-video-diffusion-models/83429/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What three tasks does LongE2V handle in a unified architecture?","Question",{"text":75,"@type":76},"LongE2V jointly performs video reconstruction, long-term video prediction, and video frame interpolation within one architecture conditioned on event voxels.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does LongE2V reduce temporal drift in long-horizon prediction?",{"text":80,"@type":76},"It introduces an autoregressive unrolling strategy designed to minimize drift during long-term sequence generation.",{"name":82,"@type":73,"acceptedAnswer":83},"How does LongE2V perform zero-shot frame interpolation?",{"text":84,"@type":76},"It synthesizes intermediate frames by adapting to event dynamics as temporal guidance, enabling zero-shot behavior for interpolation.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]