[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82951-en":3,"doc-seo-82951-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82951,1099514068035,"Ezra","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","MV-Forcing Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Self-Forcing","Video diffusion advances enable either long single-view generation via temporal autoregression or short multi-view synthesis using bidirectional attention. Long, multi-view consistent video generation for dynamic scenes remains unresolved. MV-Forcing introduces a single diffusion framework that composes temporal and view-wise autoregression using a 4D geometric bridge between sequentially generated views. A 3D reconstruction module builds a geometric prior for the next viewpoint, refined into high-quality video, while joint denoising enables temporally unbounded generation. Distribution Matching Distillation reduces train–inference exposure bias for both temporal and view autoregression.","arXiv :2607 .05376v 1 [ cs .CV] 6 Jul 2026  \nMV-Forcing: Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Self-Forcing  \nGal Fiebelman 1 , Hadar Averbuch-Elor2 , and Sagie Benaim 1  \n1 The Hebrew University of Jerusalem  \n2 Cornell University  \n[https://galfiebelman.github.io/mv-forcing/](https://galfiebelman.github.io/mv-forcing/)  \nAbstract. Recent advances in video diffusion models have enabled either long single-view generation through temporal autoregression, or short multi-view synthesis through bidirectional attention. However, generating long, multi-view consistent videos of dynamic scenes remains unsolved. In this work, we present MV-Forcing, a framework that composes temporal and view-wise autoregression within a single diffusion model by introducing a 4D geometric bridge between sequentially generated views. Our key insight is that an autoregressive 3D reconstruction model naturally interfaces between autoregressively generated views.  \nGiven a completed source view, we reconstruct its 3D structure and render a geometric prior of the next target viewpoint, which the diffusion model refines into a high-quality video. To extend generation beyond the teacher’s fixed temporal window, we introduce a joint denoising regime where both view slots are initialized from noise during training, enabling temporally unbounded generation. We distill the model via Distribution Matching Distillation with Spatio-Temporal Self-Forcing, closing the train-inference exposure bias gap for both temporal and view-sequential autoregression. Extensive experiments on both synthetic and real-world data demonstrate that MV-Forcing produces geometrically consistent multi-view videos of dynamic scenes at arbitrary lengths and viewpoint counts using a single few-step student model.  \nKeywords: Video Diffusion Models · Multi-View Generation · Autoregressive Generation  \n1 Introduction  \nGenerating a long, temporally coherent video from multiple viewpoints simultaneously is a long-standing challenge in computer vision. This task, long multiview video generation, requires a model to synthesize the continuous temporal dynamics of a scene while maintaining strict 3D geometric consistency across arbitrary camera trajectories. Unlocking this capability is critical for a wide array of downstream applications, ranging from immersive virtual reality and advanced cinematic content creation to building dynamic, interactive simulations. However, jointly modeling the complex distribution of photorealistic video across  \n2 G. Fiebelman et al.  \nFig. 1: Given a text prompt and camera sequences, MV-Forcing generates coherent video across an arbitrary number of viewpoints and unbounded temporal horizons. For each prompt, we show 3 views at increasing camera displacements (rows) across 160 timesteps (columns) . Appearance, motion, and scene geometry remain consistent across all views and timesteps, demonstrating the effectiveness of our 4D-grounded geometric prior in maintaining cross-view consistency over long-horizon generation.  \nboth unbounded temporal horizons and an arbitrary number of viewpoints demands an intricate balance of spatial, temporal, and geometric reasoning.  \nRecent years have witnessed remarkable progress along two orthogonal directions in video synthesis. On one hand, continuous advances in autoregressive video diffusion have pushed the boundaries of temporal duration, enabling singleview generation over minute-long horizons by sequentially conditioning on previously generated context [16,23,33] . On the other hand, multi-view generation has seen exciting progress, with models now capable of synchronizing dynamic openworld scenes across multiple camera viewpoints [2, 20] . However, despite these parallel successes, unifying these paradigms to achieve long, multi-view generation remains a formidable bottleneck. Current multi-view approaches heavily rely on bidirectional attention across the entire time-view grid to enforce 3D consi","cbCaidZzCboYF68g","https://ap.wps.com/l/cbCaidZzCboYF68g","pdf",37793146,3,1,29,"English","en",105,"# Introduction\n## Long multi-view video generation challenges\n## Related directions: temporal autoregression vs multi-view diffusion\n## MV-Forcing framework overview","[{\"question\":\"What problem does MV-Forcing address?\",\"answer\":\"MV-Forcing targets long, multi-view video generation that remains geometrically consistent across arbitrary camera trajectories and unbounded time horizons.\"},{\"question\":\"How does MV-Forcing enforce 3D consistency across views?\",\"answer\":\"It reconstructs a 3D structure from a completed source view and renders a 4D-grounded geometric prior for the next viewpoint, integrating this prior within the diffusion model rather than relying on dense bidirectional attention.\"},{\"question\":\"How does MV-Forcing enable generation beyond a fixed temporal window?\",\"answer\":\"It uses a joint denoising regime where view slots are initialized from noise during training, allowing temporally unbounded generation during inference.\"}]",1784184280,73,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"mv-forcing-long-multi-view-video-generation-via-4d-grounded-spatio-temporal-self-forcing","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/mv-forcing-long-multi-view-video-generation-via-4d-grounded-spatio-temporal-self-forcing/82951/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does MV-Forcing address?","Question",{"text":75,"@type":76},"MV-Forcing targets long, multi-view video generation that remains geometrically consistent across arbitrary camera trajectories and unbounded time horizons.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does MV-Forcing enforce 3D consistency across views?",{"text":80,"@type":76},"It reconstructs a 3D structure from a completed source view and renders a 4D-grounded geometric prior for the next viewpoint, integrating this prior within the diffusion model rather than relying on dense bidirectional attention.",{"name":82,"@type":73,"acceptedAnswer":83},"How does MV-Forcing enable generation beyond a fixed temporal window?",{"text":84,"@type":76},"It uses a joint denoising regime where view slots are initialized from noise during training, allowing temporally unbounded generation during inference.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]