[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85184-en":3,"doc-seo-85184-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},85184,962075114765,"Quinn","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","PhysMRV Physical Memory Retrieval and Verification for Physics Plausibility Reasoning","Video-language models achieve strong performance in video understanding but remain unreliable for reasoning about physical plausibility, where object interactions, causal dynamics, and fundamental physics must be judged from short evidence. PHYSMRV addresses this via a training-free hierarchical physical memory bank: scene descriptions, event graphs encoding causal structure, and reusable physics-rule summaries. At inference, physically relevant memories guide a frozen VLM to verify plausibility through structured evidence, improving results across multiple VLMs and benchmarks without fine-tuning.","arXiv :2607 . 10 190v 1 [ cs .LG] 11 Jul 2026  \nPhysMRV: Physical Memory Retrieval and Verification for Physics Plausibility Reasoning  \nWenyuan Wang 1 2 Lianyu Hu 1 Hao Wang 2 3 Yang Liu 1  \nAbstract  \nVideo-language models (VLMs) have achieved remarkable performance on video understanding and visual question answering, yet they remain unreliable in reasoning about physical plausibility, where understanding object interactions, causal dynamics, and fundamental physical principles is essential. This limitation is particularly evident on challenging physical reasoning benchmarks, revealing a persistent gap in physical commonsense reasoning. To address this challenge, we propose PHYSMRV, a training-free physical memory and verification framework for physical plausibility reasoning. Unlike retrieval-augmented VLMs that retrieve semantically similar videos as additional context, PHYSMRV transforms training videos into a Hierarchical Memory Bank of structured physical knowledge comprising three complementary levels: scene descriptions capturing visual context, physicalevent graphs modeling object interactions and causal structure, and physics-rule summaries distilling reusable physical principles and cues. During inference, PHYSMRV retrieves physically relevant memories and leverages their structured physical evidence to guide a frozen VLM in verifying physical plausibility, requiring neither fine-tuning nor parameter updates. We evaluate PHYSMRV on three challenging physical reasoning benchmarks, ImplausiBench, IntPhys2, and GRASP Level 2, across multiple state-of-the-art VLMs. Experimental results demonstrate consistent improvements over direct prompting across diverse VLMs and evaluation benchmarks, showing that structured physical memories provide an effective and scalable means of enhancing physical plausibility reasoning without additional training.  \n1. Introduction  \nRecent VLMs have achieved impressive progress in video understanding, demonstrating strong performance in perceptual tasks and reasoning abilities, including spatial reasoning and temporal reasoning.(Bai et al., 2025 ; NVIDIA, 2026 ; Wanget al., 2025 ; An et al., 2026) . However, successful deployment in real-world environments requires capabilities beyond recognizing objects, actions, and scenes. A model also needs to reason about the physical world and determine whether an observed event is physically plausible. For example, it should recognize whether an unsupported object can remain suspended in midair, whether an object should continue to exist after becoming occluded, and whether a collision produces a physically consistent outcome. Such judgments require physical commonsense reasoning about object permanence, causal interactions, and the fundamental principles governing real-world dynamics. This makes physical plausibility reasoning fundamentally different from conventional video understanding: a model may correctly identify visible entities and actions while still failing to detect violations of basic physical laws, because such judgments cannot be directly captured by object-level semantics and instead require an understanding of real-world physical constraints.  \nPrior benchmarks have begun to systematically examine this capability. ImplausiBench focuses on detecting physical violations from visual evidence (Motamed et al., 2025), IntPhys2 evaluates intuitive physics through violation-of-expectation scenarios involving object permanence, continuity, and solidity (Bordes et al., 2025), and GRASP studies physical reasoning in grounded and interactive settings (Jassim et al., 2023) . Figure 1 shows representative examples from these benchmarks. Results on these benchmarks show that current multimodal models still fall substantially below human-level physical understanding, suggesting that VLMs lack a reliable mechanism for verifying whether observed evidence satisfies physical  \n1Nanyang Technological University 2Rutgers University 3University of Illinoi","cbCaimtRsoZbZP6S","https://ap.wps.com/l/cbCaimtRsoZbZP6S","pdf",6073669,1,12,"English","en",105,"# Introduction\n# Background and Benchmarks\n## ImplausiBench\n## IntPhys2\n## GRASP Level 2","[{\"question\":\"Why do video-language models struggle with physical plausibility reasoning?\",\"answer\":\"They can recognize visible entities and actions yet fail to verify violations of basic physical laws, because such judgments require explicit understanding of physical constraints beyond object-level semantics.\"},{\"question\":\"What is PHYSMRV and how does it work?\",\"answer\":\"PHYSMRV is a training-free framework that converts training videos into a hierarchical physical memory bank (scene descriptions, event graphs, and physics-rule summaries). During inference, it retrieves physically relevant memories and uses them to guide a frozen VLM to verify plausibility.\"},{\"question\":\"Does PHYSMRV require fine-tuning the VLM?\",\"answer\":\"No. PHYSMRV improves physical plausibility reasoning without fine-tuning or updating model parameters, relying instead on retrieval and structured verification signals.\"}]",1784201608,30,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"physmrv-physical-memory-retrieval-and-verification-for-physics-plausibility-reasoning","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/physmrv-physical-memory-retrieval-and-verification-for-physics-plausibility-reasoning/85184/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do video-language models struggle with physical plausibility reasoning?","Question",{"text":75,"@type":76},"They can recognize visible entities and actions yet fail to verify violations of basic physical laws, because such judgments require explicit understanding of physical constraints beyond object-level semantics.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is PHYSMRV and how does it work?",{"text":80,"@type":76},"PHYSMRV is a training-free framework that converts training videos into a hierarchical physical memory bank (scene descriptions, event graphs, and physics-rule summaries). During inference, it retrieves physically relevant memories and uses them to guide a frozen VLM to verify plausibility.",{"name":82,"@type":73,"acceptedAnswer":83},"Does PHYSMRV require fine-tuning the VLM?",{"text":84,"@type":76},"No. PHYSMRV improves physical plausibility reasoning without fine-tuning or updating model parameters, relying instead on retrieval and structured verification signals.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":28,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]