[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83476-en":3,"doc-seo-83476-105":30,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83476,1099513958762,"Logic","https://ap-avatar.wpscdn.com/avatar/1000023916a998db790?x-image-process=image/resize,m_fixed,w_180,h_180&k=1784791008015729253",8,"Research & Report","RetailSMV Exocentric vs Egocentric Adaptation of Foundation Video World Models in Retail","Foundation video diffusion models are used as world simulators for embodied agents, but their internet-scale pretraining can mismatch real retail deployment domains. The study develops parameter-efficient adaptation of a pretrained foundation video world model to retail using Low-Rank Adaptation (LoRA). RetailSMV introduces 32,105 captioned clips with synchronized egocentric and exocentric views from five supermarkets, and trains matched LoRA variants for viewpoint-specific learning. Exocentric-only adaptation matches or exceeds combined adaptation on most metrics, with the largest gap at the shortest rollout horizon.","arXiv :2607 .003 10v 1 [ cs .CV] 1 Jul 2026  \nRetailSMV: Exocentric vs . Egocentric Adaptation of Foundation Video World Models in Retail  \nDreamVu 1  \nFoundation video diffusion models are increasingly viewed as world simulators for embodied agents, yet their pretraining on internet-scale generic video leaves them poorly aligned with real-world deployment domains. We study parameter-efficient adaptation of a pretrained foundation video world model to retail scenes: when synchronized egocentric and exocentric video of the same activity are available, which viewpoint of training data produces the strongest adapted model?  \nWe introduce RetailSMV (Retail Synchronized Multi-View), a corpus of 32 , 105 captioned retail clips from five supermarkets with synchronized ego/exo capture from the store-staff perspective (stocking, arranging, weighing, managing supply carts, scanning at checkout), rather than the customer-centric framing of prior retail video corpora, and train three matched Low-Rank Adaptation (LoRA) configurations of Cosmos3-Nano (egocentric-only, exocentric-only, combined) under identical hyperparameters. On a 200-clip held-out test set evaluated with seven complementary metrics under a strict paired statistical protocol, exocentric-only adaptation matches or exceeds combined adaptation on six of seven point estimates and is significantly better on LPIPS, PSNR, and DreamSim, despite training on only 15 ,985 exocentric clips (versus 32 , 105 for combined) . A symmetric paired comparison further shows that adding exocentric data to egocentric-only training helps while adding egocentric data toexocentric-only training hurts. The absolute adaptation gap is largest at the shortest rollout time, identifying the near-horizon prediction window as the regime in which adaptation is most beneficial.  \nDate: July 2, 2026  \nFigure 1 Base vs. RetailSMV-adapted video world model. The same retail prompts are continued by the pretrained Cosmos3-Nano foundation model (top, red) and by our RetailSMV-adapted LoRA (bottom, green) under identical inference settings. RetailSMV adaptation preserves scene layout (hand-off watermelon), action grounding (weigh tomato crate), and physical geometry (open fridge) where the pretrained baseline drifts.  \n1A detailed list of contributors and acknowledgments can be found in section E of this paper.  \n1 Introduction  \nA video world model takes an observation history (typically a text description, an image, or a short video clip) and predicts a plausible video continuation. Modern video diffusion systems such as Sora OpenAI (2024), Stable Video Diffusion Blattmann et al. (2023), Movie Gen Meta GenAI (2024), and NVIDIA Cosmos3-Nano and Cosmos-Predict 2.5 NVIDIA (2025) now produce coherent multi-second video from text prompts and have been framed as world simulators for embodied agents Ha and Schmidhuber (2018); Gao et al. (2025) . The promise is that a faithful world model could let an embodied agent plan, evaluate counterfactuals, generate training data, and verify policies before acting in the physical world. We discuss the broader landscape of video diffusion, world models, and embodied deployment in section 2 .  \nThat promise depends on domain alignment. Foundation video models are pretrained on internet-scale generic video (entertainment content, vlogs, dashcam footage, and robot demonstrations) which may underrepresent the structured visual vocabulary of specific deployment domains. Retail environments are a clear example: dense product shelving, narrow parallel aisles, repetitive geometry at multiple scales, distinctive signage and end-caps, and multi-person dynamics around carts and checkout counters. A retail world model must render scenes built from this vocabulary and continue shopping behaviors consistent with the physical and behavioral regularities of supermarket environments. Without explicit adaptation, generated retail scenes drift toward generic interiors, and predicted human motion drif","cbCailJm65lsFudF","https://ap.wps.com/l/cbCailJm65lsFudF","pdf",38351696,4,1,31,"English","en",105,"# Introduction\n## Problem: Domain alignment for retail video world models\n## RetailSMV dataset and synchronized ego/exo capture\n## Research question and experimental setup","[{\"question\":\"Which training viewpoint produces the best adapted model for retail scenes?\",\"answer\":\"Exocentric-only adaptation matches or exceeds combined adaptation on six of seven paired evaluation estimates and is significantly better on LPIPS, PSNR, and DreamSim. Adding exocentric data helps egocentric-only training, while adding egocentric data hurts exocentric-only training.\"}]",1784188280,78,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":28},"retailsmv-exocentric-vs-egocentric-adaptation-of-foundation-video-world-models-in-retail","",{"@graph":36,"@context":77},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/retailsmv-exocentric-vs-egocentric-adaptation-of-foundation-video-world-models-in-retail/83476/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"Which training viewpoint produces the best adapted model for retail scenes?","Question",{"text":75,"@type":76},"Exocentric-only adaptation matches or exceeds combined adaptation on six of seven paired evaluation estimates and is significantly better on LPIPS, PSNR, and DreamSim. Adding exocentric data helps egocentric-only training, while adding egocentric data hurts exocentric-only training.","Answer","https://schema.org",{"og:url":52,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,120,123,127],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":46,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":128,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":98,"slug":130},19,"General","general"]