[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86383-en":3,"doc-seo-86383-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86383,34359740700684,"Finn","https://ap-avatar.wpscdn.com/avatar/1f400023980c374ae676?_k=1777273430885731487",8,"Research & Report","SalFormer360: A Transformer-Based Saliency Estimation Model for 360-Degree Videos","SalFormer360 presents a transformer-based saliency estimation model tailored to 360-degree videos, targeting key VR tasks such as viewport prediction and immersive content optimization. The method combines a transformer-driven encoder based on SegFormer with a custom decoder, adapting 2D segmentation design to spherical video frames. A viewing-center bias is integrated to better match user attention patterns. Experiments on three major saliency benchmarks show consistent improvements, including gains in Pearson correlation over prior state of the art.","SalFormer360: a transformer-based saliency estimation model for 360-degree videos  \nMahmoud Z. A. Wahba, Graduate Student Member, IEEE, Francesco Barbato, Member, IEEE, Sara Baldoni, Member, IEEE, and Federica Battisti, Senior Member, IEEE  \narXiv :2602 .04584v2 [ cs .CV] 13 Jul 2026  \nAbstract—Saliency estimation has received growing attention in recent years due to its importance in a wide range of applications. In the context of 360-degree video, it has been particularly valuable for tasks such as viewport prediction and immersive content optimization. In this paper, we propose SalFormer360, a novel saliency estimation model for 360-degree videos built on a transformer-based architecture. Our approach is based on the combination of an existing encoder architecture, SegFormer, anda custom decoder. The SegFormer model was originally developed for 2D segmentation tasks, and it has been fine-tuned to adapt it to 360-degree content. To further enhance prediction accuracy in our model, we incorporated a viewing center bias to reflect user attention in 360-degree environments. Extensive experiments on the three largest benchmark datasets for saliency estimation demonstrate that SalFormer360 outperforms existing state-ofthe-art methods. In terms of Pearson correlation coefficient, our model achieves 8.4% higher performance on Sport360, 2.5% on PVS-HM, and 18.6% on VR-EyeTracking compared to previous state-of-the-art.  \nIndex Terms—Saliency estimation, Omni-directional video, Viewing bias, Transformers.  \nI. INTRODUCTION  \nIn recent years, Virtual Reality (VR) has gained widespread popularity, providing users with highly immersive experiencesand allowing them to feel as if they were in a virtual world distinct from reality.  \nOmnidirectional images and videos have become one of the most popular content types for VR, thanks to the availability of user-friendly and low-cost 360-degree cameras. This content allows users to be placed at the center of a sphere and freely explore the environment in any direction by simply moving their heads. Although 360-degree media allows increased user immersion, their processing and transmission still entail many challenges. One major issue is the high memory and bandwidth requirements with respect to standard 2D content. For example, streaming a 4K 2D video requires approximately 25 Mb/s, while delivering a 4K resolution for each eye to provide a full 360-degree viewing experience requires around 400 Mb/s [1] .  \nOne effective way to address the transmission challenges of 360-degree videos is to implement a user-centered streaming  \nThis work was partially supported by the European Union under the Italian National Recovery and Resilience Plan (NRRP) Mission 4, Component 2, Investment 1.3, CUP C93C22005250001, partnership on “Telecommunications of the Future” (PE00000001 - program “RESTART”)” and by the European Union’s Horizon Europe Program under Agreement 101135637 (HEAT Project) .  \nM. Z. A. Wahba, F. Barbato, S. Baldoni, and F. Battisti are with the Department of Information Engineering, University of Padova, Via Gradenigo 6b, 35131, Padua, Italy. (Corresponding author e-mail: [sara.baldoni@unipd.it](sara.baldoni@unipd.it)).  \nFig. 1: SalFormer360 can estimate future salient points in 360-degree videos using only a single previous frame and with limited computational resources.  \nparadigm. To this aim, human attention mechanisms have been studied to design saliency estimation methods. These algorithms compute 2D probability maps which highlight the regions inside a 360-degree scene most likely to draw users attention [2] . These maps can then be used to transmit the salient regions at higher quality while encoding at lower quality (or discarding) less relevant areas [3] .  \nAnother possibility to reduce the transmission burden of 360-degree content considers the limitations of the human visual system, reflected by the Head-Mounted Displays (HMDs) . Indeed, despite the availability of 360-degree c","cbCaioMVG2VRtrE6","https://ap.wps.com/l/cbCaioMVG2VRtrE6","pdf",9331348,6,1,12,"English","en",105,"# Introduction\n## Motivation from VR and 360-degree streaming\n## Related work: attention, saliency maps, and viewport prediction\n## Proposed approach and model design","[{\"question\":\"What problem does SalFormer360 address in 360-degree video systems?\",\"answer\":\"It estimates visual saliency for 360-degree videos, supporting applications like viewport prediction and immersive content optimization to improve streaming efficiency and effectiveness.\"},{\"question\":\"How is the SalFormer360 model built?\",\"answer\":\"It uses a SegFormer encoder adapted from 2D segmentation and a custom decoder to generate saliency outputs for 360-degree frames.\"},{\"question\":\"What additional component improves prediction accuracy?\",\"answer\":\"A viewing center bias is incorporated to reflect where users tend to focus in 360-degree environments.\"}]",1784211392,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"salformer360-a-transformer-based-saliency-estimation-model-for-360-degree-videos","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/salformer360-a-transformer-based-saliency-estimation-model-for-360-degree-videos/86383/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does SalFormer360 address in 360-degree video systems?","Question",{"text":76,"@type":77},"It estimates visual saliency for 360-degree videos, supporting applications like viewport prediction and immersive content optimization to improve streaming efficiency and effectiveness.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How is the SalFormer360 model built?",{"text":81,"@type":77},"It uses a SegFormer encoder adapted from 2D segmentation and a custom decoder to generate saliency outputs for 360-degree frames.",{"name":83,"@type":74,"acceptedAnswer":84},"What additional component improves prediction accuracy?",{"text":85,"@type":77},"A viewing center bias is incorporated to reflect where users tend to focus in 360-degree environments.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,115,120,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":107,"slug":137},19,"General","general"]