[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83431-en":3,"doc-seo-83431-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83431,7971461741311,"Ophelia","https://ap-avatar.wpscdn.com/avatar/74000253aff267980c6?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779345379180704826",8,"Research & Report","Enhancing In-context Panoramic Generation via Geometric-aware Pretraining","Canvas360 introduces a two-stage framework for in-context panoramic generation, combining geometry-aware pretraining with downstream task-specific fine-tuning to overcome limited, high-quality training data. Canvas360Dataset provides 1M paired panoramic samples for style transfer, inpainting, outpainting, and editing. Geometry-aware modeling uses parallel depth generation, velocity circular padding, and similarity loss to learn geometry-aware representations and improve geometric consistency. A unified token-level concatenation scheme enables flexible multi-task in-context generation, delivering higher panorama fidelity with strong FAED results.","arXiv :2607 .08765v2 [ cs .CV] 12 Jul 2026  \nEnhancing In-context Panoramic Generation via Geometric-aware Pretraining  \nHaoran Feng1,2 ∗ Ruiyang Zhang1,3 ∗ Longyi Zhang2 Dizhe Zhang1B† Lu Qi1,4 B 1 Insta360 Research 2 Tsinghua University 3 Beihang University 4 Wuhan University  \nText-to-Panorama Generation  \nStyle Transfer Panorama Editing Inpainting & Outpainting  \nFigure 1: Visualization of Canvas360’s results. The examples cover text-to-panorama generation, inpainting, outpainting, panorama editing, and style transfer. These results demonstrate that Canvas360 achieves strong generative performance, captures a rich panoramic prior, and supports a wide range of downstream applications. Additional results are provided in Sec. G.  \nAbstract  \nIn this work, we present Canvas360, a two-stage framework for in-context panoramic generation that combines geometry-aware pretraining with downstream task-specific fine-tuning. To address the lack of large-scale, high-quality training data tailored to in-context panoramic tasks, we propose Canvas360Dataset, a collection of 1M high-quality paired panoramic samples for style transfer, inpainting, outpainting, and editing, enabling effective supervision across diverse in-context generation scenarios. On the modeling side, Canvas360 enhances text-to-panorama generation through parallel depth generation, velocity circular padding, and similarity loss regularization, enabling the model to learn geometry-aware representations, capture object distortion details, and improve geometric consistency and global coherence. Furthermore, empowered by strong panoramic priors, Canvas360 enables a unified in-context panoramic generation framework that supports diverse downstream tasks via token-level concatenation, surpassing prior methods in both task coverage and modeling flexibility. Extensive experiments show that Canvas360 improves panoramic image fidelity, achieving particularly strong performance on the panorama-specific FAED metric and competitive or leading results across the reported quantitative evaluations. More information can be found on our project  \n page: [https://zry000.g](https://zry000.github.io/Canvas360/)[ithub.io/Canvas360/](https://zry000.github.io/Canvas360/) . 0 ∗ Equal Contribution † Project Lead B Corresponding Author  \nPreprint.  \n1 Introduction  \nWith the rapid progress of panoramic text-to-image generation models [Ye et al., 2024, Zhang et al., 2024, Xie, 2025, Sun et al., 2025, Ni et al., 2025, Bar-Tal et al., 2023, Li and Bansal, 2023, Shi et al., 2023, Tang et al., 2023], in-context editing has emerged as a natural extension beyond basic text-to-panorama, enabling image generation conditioned jointly on user-provided images and textual prompts [Brooks et al., 2023, Labs et al., 2025, Liu et al., 2025, Suvorov et al., 2022] . This capability underpins a wide range of interactive applications, including content-aware editing [Google, 2026a, ByteDance Seed, 2026, Wu et al., 2025b] and immersive scene manipulation [Deng et al., 2025, Yu et al., 2025b] .  \nDespite these advances, the dominant equirectangular projection (ERP) representation for panoramic images inherently exhibits latitude-dependent distortions, posing challenges for geometry-consistent editing. Existing panoramic image editing methods [Yang et al., 2025a, Zhong et al., 2025] attempt to mitigate this issue through distortion-aware designs, such as cube-map-based editing [Yang et al., 2025a] or 3D spherical positional embeddings [Zhong et al., 2025] . Nevertheless, we empirically observe that these approaches still struggle to preserve geometric consistency in the underlying 3D scene structure when operating on ERP panoramas.  \nInspired by common practices in perspective visual generation, prior works often introduce depth constraints as explicit geometric priors during training [Huang et al., 2025a, Bai et al., 2025b, Bhat et al., 2024, Zhang et al., 2023a, Yu et al., 2025b] . However, the geometric formulation ","cbCaislmAFTZ70Fe","https://ap.wps.com/l/cbCaislmAFTZ70Fe","pdf",20248559,5,1,26,"English","en",105,"# Abstract\n# Introduction\n## Motivation: geometry-consistent in-context editing on ERP panoramas\n## Canvas360: geometry-aware pretraining and unified fine-tuning\n## Data: Canvas360Dataset and scalable synthesis pipeline","[{\"question\":\"What is Canvas360 and what problem does it address?\",\"answer\":\"Canvas360 is a two-stage framework for in-context panoramic generation that combines geometry-aware pretraining with downstream fine-tuning. It targets the lack of large-scale, high-quality training data tailored to in-context panoramic tasks and the difficulty of maintaining geometry consistency on equirectangular panoramas.\"},{\"question\":\"How does Canvas360 improve text-to-panorama generation geometrically?\",\"answer\":\"During pretraining, the model performs parallel depth generation, applies velocity circular padding, and uses similarity loss regularization. These components help the model learn geometry-aware representations, preserve object distortion details, and enhance geometric consistency and global coherence.\"},{\"question\":\"Which downstream tasks does Canvas360 support in its unified in-context generation stage?\",\"answer\":\"Canvas360 supports style transfer, inpainting, outpainting, and panorama editing. In fine-tuning, depth is discarded and a unified model is trained using token-level concatenation for heterogeneous contextual conditions.\"}]",1784187628,66,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"enhancing-in-context-panoramic-generation-via-geometric-aware-pretraining","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/enhancing-in-context-panoramic-generation-via-geometric-aware-pretraining/83431/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What is Canvas360 and what problem does it address?","Question",{"text":76,"@type":77},"Canvas360 is a two-stage framework for in-context panoramic generation that combines geometry-aware pretraining with downstream fine-tuning. It targets the lack of large-scale, high-quality training data tailored to in-context panoramic tasks and the difficulty of maintaining geometry consistency on equirectangular panoramas.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does Canvas360 improve text-to-panorama generation geometrically?",{"text":81,"@type":77},"During pretraining, the model performs parallel depth generation, applies velocity circular padding, and uses similarity loss regularization. These components help the model learn geometry-aware representations, preserve object distortion details, and enhance geometric consistency and global coherence.",{"name":83,"@type":74,"acceptedAnswer":84},"Which downstream tasks does Canvas360 support in its unified in-context generation stage?",{"text":85,"@type":77},"Canvas360 supports style transfer, inpainting, outpainting, and panorama editing. In fine-tuning, depth is discarded and a unified model is trained using token-level concatenation for heterogeneous contextual conditions.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]