[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82453-en":3,"doc-seo-82453-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82453,7971461741311,"Ophelia","https://ap-avatar.wpscdn.com/avatar/74000253aff267980c6?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779345379180704826",8,"Research & Report","PanoWorld: Real-World Panoramic Generation","PanoWorld targets long-range memory in panoramic world models by leveraging the rotation-equivariant nature of omnidirectional representations, treating rotation as an implicit geometric transformation. The approach simplifies camera trajectories into translations through fixed headings and introduces Dense Panoramic Ray-Conditioning (DPRC) and Geometry-aware Memory Augmentation (GMA) for both current-action modeling and long-range memory. A three-stage training pipeline progressively optimizes each component. To evaluate physical consistency under large spatial variation and diverse illumination, World360 is built from 70K real videos and 50K AirSim360 simulations.","PanoWorld: Real-World Panoramic Generation  \nHaoyuan Li1 Dizhe Zhang1 \\# ∗ Yuemei Zhou1 Xiangkai Zhang1, 2 Haoran Feng1, 3  \nXiaofan Lin1 Wenjie Jiang1 Bo Du4 Ming-Hsuan Yang5 Lu Qi1,4 \\#  \n1Insta360 Research 2Institute of Automation Chinese Academy of Sciences  \n3Tsinghua University 4Wuhan University 5UC Merced  \narXiv :2607 .09661v1 [ cs .CV] 10 Jul 2026  \n. . .  \nMatrix3D  \nOmniRoam  \nOurs  \nFigure 1: PanoWorld is a novel framework for high-fidelity and controllable panoramic video generation. Our approach achieves precise trajectory control across complex movements while maintaining high-fidelity visual synthesis with physical consistency in diverse real-world environments.  \nAbstract  \nIn this work, we aim to address the challenge of long-range memory in panoramic world models by exploiting the rotation-equivariant property of omnidirectional representations, where rotation can be treated as an implicit geometric transformation. Building on this insight, we propose PanoWorld, which simplifies camera trajectories into translations via fixed headings for both current-action modeling and long-range memory through Dense Panoramic Ray-Conditioning (DPRC) and Geometry-aware Memory Augmentation (GMA). Then, a three-stage training pipeline is introduced to progressively optimize each component. To better evaluate physical consistency under large-scale spatial variations and diverse illumination conditions, where existing datasets are relatively stable, we construct World360, a large-scale dataset consisting of both real-world video clips collected via panoramic unmanned aerial vehicles and high-quality simulated clips generated by AirSim360 .  \nExtensive experiments on World360 demonstrate the effectiveness of PanoWorld, outperforming alternative methods by a large margin. Our models, training code, and dataset will be publicly available. More information can be found on our project page: [https://lihaoy-ux.github.io/panoworld-page/](https://lihaoy-ux.github.io/panoworld-page/) .  \n∗ Project Lead \\# Corresponding Author  \nPreprint.  \n1 Introduction  \nRecently, world models have attracted significant attention for modeling dynamic environments and enabling controllable generation [Tang et al., 2025, Team et al., 2026] . They have shown strong potential in various robotic applications, including autonomous driving and unmanned aerial vehicles [Wu et al., 2026] .  \nAmong various world models, panoramic representations, particularly equirectangular projections (ERP), have emerged as a mainstream paradigm by capturing the full 360◦ field of view (FoV) for each frame in a specific trajectory [Lin et al., 2025] . Despite recent advancements, achieving physical consistency, such as geometry and illumination, across space and time remains challenging for panoramic world models, where such problem can be amplified by the full-view nature.  \nMost existing work, including 3DGS-based and video-generation paradigms, adopts memory mechanisms (e.g., 3D points or KV caches) to retrieve past information for spatiotemporal consistency [Yang et al., 2025, Chou et al., 2025, Schwarz et al., 2025] . However, these solutions inherit perspective assumptions and ignore the unique properties of panoramic data, leading to misaligned memory retrieval under severe distortion and rotation-induced viewpoint shifts. Thus, one question raised: how can memory mechanisms better adapt panoramic representations?  \nTo address this issue, we begin by analyzing the properties of the equirectangular projection (ERP), which is rotation-equivariant, where rotations mainly alter the distortion pattern while preserving the underlying scene content. This observation inspires us to simplify camera motion by treating rotation as an geometric transformation and modeling only translation explicitly. Building on this insight, we propose PanoWorld, a diffusion-based framework that incorporates Dense Panoramic Ray-Conditioning (DPRC) and Geometry-aware Memory Augmentation (GMA) mod","cbCaimwh54GKzurU","https://ap.wps.com/l/cbCaimwh54GKzurU","pdf",17286800,1,23,"English","en",105,"# Introduction\n# Related Work","[{\"question\":\"What problem does PanoWorld address in panoramic world models?\",\"answer\":\"PanoWorld addresses long-range memory, where panoramic representations suffer from misaligned memory retrieval due to distortion and rotation-induced viewpoint shifts.\"},{\"question\":\"How does PanoWorld simplify camera motion for training and generation?\",\"answer\":\"It decouples translation and rotation by treating rotation as an implicit geometric transformation, then modeling only translation explicitly using fixed headings.\"},{\"question\":\"What are the main modules and training strategy used by PanoWorld?\",\"answer\":\"PanoWorld uses Dense Panoramic Ray-Conditioning (DPRC) and Geometry-aware Memory Augmentation (GMA), optimized through a three-stage training pipeline that progressively refines each component.\"}]",1784180501,58,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"panoworld-real-world-panoramic-generation","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/panoworld-real-world-panoramic-generation/82453/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-21","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does PanoWorld address in panoramic world models?","Question",{"text":75,"@type":76},"PanoWorld addresses long-range memory, where panoramic representations suffer from misaligned memory retrieval due to distortion and rotation-induced viewpoint shifts.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does PanoWorld simplify camera motion for training and generation?",{"text":80,"@type":76},"It decouples translation and rotation by treating rotation as an implicit geometric transformation, then modeling only translation explicitly using fixed headings.",{"name":82,"@type":73,"acceptedAnswer":83},"What are the main modules and training strategy used by PanoWorld?",{"text":84,"@type":76},"PanoWorld uses Dense Panoramic Ray-Conditioning (DPRC) and Geometry-aware Memory Augmentation (GMA), optimized through a three-stage training pipeline that progressively refines each component.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]