[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84520-en":3,"doc-seo-84520-105":29,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84520,962075006959,"Anda","https://ap-avatar.wpscdn.com/avatar/e0002397efbe92a78e?_k=1776741047341049297",8,"Research & Report","MemoBench Benchmarking World Modeling in Dynamically Changing Environments","Video generation models aim to represent dynamic environments, and recent benchmarks measure memory consistency across frames. However, most tests check consistency only when the target stays visible, while occlusion-focused benchmarks often use static scenes. MemoBench introduces a disappear-and-reappear diagnostic setup where a target object continues a physical process while leaving and returning to view. The benchmark provides 360 ground-truth clips (synthetic and real) and an evaluation suite with automated metrics plus VQA scored by LLMs.","arXiv :2606 .27537v5 [ cs .CV] 13 Jul 2026  \nMemoBench: Benchmarking World Modeling in Dynamically Changing Environments  \nHaoyu Chen 1 Kaichen Zhou 1,2 Hang Hua3 Kaile Zhang4 Jingwen Qian5 Wufei Ma6 Haonan Chen 1 Chunjiang Liu7 Yizhou Zhao7 Xiaoyuan Wang7 Weiyue Li 1 Alan Yuille6 Paul Pu Liang2 Yilun Du 1,8  \n1 Harvard University 2 MIT 3 MIT-IBM Watson AI Lab 4 Boston University  \n5 Google 6 JHU 7 CMU 8 Kempner Institute  \nAbstract. Video generation models aspire to simulate dynamic environments, and several benchmarks now evaluate memory consistency across frames. However, most assess consistency only while the target remains in view, and the few that force objects out of view evaluate static scenes where nothing changes during occlusion. To bridge this gap, we introduce MemoBench, a diagnostic benchmark built around the disappear-andreappear paradigm in dynamically changing environments: a target object undergoes a physical process, disappears from view, and must be correctly recovered in its updated state upon reappearance. We curate 360 ground-truth clips spanning synthetic and real-world scenes, and design an evaluation suite combining automated metrics with VQA-based assessment across four diagnostic pillars. Evaluation of ten state-of-the-art models reveals key insights and open challenges regarding memory consistency under the disappear-and-reappear paradigm. Our dataset, code, and leaderboard are available at [https://github.com/MemoBench-Team](https://github.com/MemoBench-Team).  \nKeywords: World Generation, Video Generation, Memory Consistency  \n1 Introduction  \nThe real world is inherently dynamic, continuously evolving regardless of whether anyone is watching: ice melts, flames flicker, pedestrians walk, and traffic flows. Faithfully modeling such dynamically changing environments is crucial to applications ranging from autonomous driving and robotic manipulation to embodied tasks, where an agent must reason about how the world has changed beyond its field of view. Recent progress in video generation [50, 83] has shown that generative models can serve as world generators, capturing environment dynamicsand enabling prediction under actions or interventions [12, 14 , 56] .  \nDespite this ambition, a fundamental challenge remains under-explored: visual memory under partial observability. In cognitive science, object permanence, the understanding that objects continue to exist when out of sight, is among the earliest cognitive milestones. An analogous capability is crucial for  \n2 H. Chen et al.  \n 196 Synthetic Clips: Barnyard 001,…, Barnyard 008,…  \nVisible Disappear Reappear  \n164 Real Clips: 001_Tablet dissolves,…, 008_Powder pouring,…  \nFig. 1: Overview of MemoBench. Rows 1–2 show a synthetic Visible–Disappear– Reappear sequence and its camera trajectory; Rows 3–4 show a real-world state-change sequence (powder pouring) . MemoBench contains 196 synthetic and 164 real-world clips, evaluated with automated metrics and LLM-judged VQA.  \nvideo generation: as the virtual camera moves, objects inevitably leave and reenter the field of view, and the generative model must faithfully reproduce their appearance, position, and any ongoing state changes upon return [34] . This disappear-and-reappear pattern is ubiquitous in everyday experience. Yet current video generation benchmarks seldom treat this as an explicit evaluation target, leaving it unclear whether generative models truly remember or merely regenerate scene content.  \nExisting benchmarks have advanced the evaluation of world generation along multiple axes, including visual quality, temporal coherence, physical adherence, and scene consistency [1, 8 , 25 , 30], but they predominantly evaluate what is continuously visible across frames. To our knowledge, none directly tests whether a generative model can maintain and update the state of objects that have temporarily left the field of view, under simultaneous camera and scene dynamics, leaving it unclear whether m","cbCaic3jUr6wlFEN","https://ap.wps.com/l/cbCaic3jUr6wlFEN","pdf",10802300,1,41,"English","en",105,"# Introduction\n# Related Work","[{\"question\":\"What does MemoBench provide for evaluation and benchmarking models?\",\"answer\":\"MemoBench includes 360 ground-truth videos at 1920×1080 across synthetic and real-world scenes, along with camera trajectories and depth maps. It evaluates models using automated metrics and LLM-judged VQA, and benchmarks ten state-of-the-art world generation models.\"}]",1784196293,103,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":27},"memobench-benchmarking-world-modeling-in-dynamically-changing-environments","",{"@graph":35,"@context":77},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/memobench-benchmarking-world-modeling-in-dynamically-changing-environments/84520/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"What does MemoBench provide for evaluation and benchmarking models?","Question",{"text":75,"@type":76},"MemoBench includes 360 ground-truth videos at 1920×1080 across synthetic and real-world scenes, along with camera trajectories and depth maps. It evaluates models using automated metrics and LLM-judged VQA, and benchmarks ten state-of-the-art world generation models.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,120,123,127],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":45,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":45,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":45,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":45,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":128,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":98,"slug":130},19,"General","general"]