[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84125-en":3,"doc-seo-84125-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84125,687197207057,"Sage","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Bridging Physical Reasoning and Task Generalization via Visual Action Outcome Reasoning Alignment","Vision-language models (VLMs) struggle to generalize in interactive physical reasoning, especially on unseen tasks and environments. Two dominant failure modes arise: hallucinated chain-of-thought (CoT) that contradicts physical reality, and misalignment between reasoning and the agent’s actions. VAORA (Visual Action Outcome Reasoning Alignment) introduces a reward design with Visual Alignment Reward and Visual-Action Alignment Reward to jointly suppress hallucinations and tighten the reasoning–behavior gap. Smooth dense success-probability rewards improve training stability. Experiments on PHYRE and Virtual Tool confirm grounded, generalizable physical intelligence under novel-task and unseen-environment settings.","arXiv :2607 .06522v 1 [ cs .AI ] 7 Jul 2026  \nBridging Physical Reasoning and Task Generalization via Visual Action Outcome Reasoning Alignment  \nHan-Jun Ko∗1 Jr-Jen Chen∗1 Haobo Yuan2 Hsin-Ying Lee2 Tiancheng Shen2  \nMing-Hsuan Yang2 Yu-Chiang Frank Wang 1  \n1National Taiwan University 2The University of California, Merced  \nAbstract  \nVision-language models (VLMs) struggle to generalize in interactive physical reasoning, particularly under unseen tasks and environments. Two key failure modes are prominent: hallucinated chain-of-thought (CoT) reasoning that contradicts physical reality, and misalignment between the model’s reasoning and actions. We present VAORA (Visual Action Outcome Reasoning Alignment), a novel reward design that directly addresses both issues. VAORA introduces two complementary rewards: Visual Alignment Reward, which anchors VLM reasoning to the visual context independent of the agent action itself, and Visual-Action Alignment Reward, which grounds reasoning in the visual outcome induced by the model’s action.  \nTogether, these rewards suppress hallucinated CoT and reduce the gap between reasoning and behavior. To improve training stability, we further employ smooth, dense rewards by estimating success probabilities using a pre-trained in-domain expert agent. Experiments on PHYRE and Virtual Tool support our performances across novel-task and unseen-environment settings, confirming that grounded and generalizable physical intelligence can be induced through VAORA.  \n1 Introduction  \nTrue physical agents must go beyond memorizing scene configurations and instead reason about spatial relationships, dynamics, and causality to act effectively in novel situations. We characterize this capability through two forms of generalization: cross-task transfer within the same environment and cross-environment transfer across distinct physics simulators. Conventional non-vision-languagemodel agents are fundamentally constrained in satisfying these criteria by their architectural design, which typically combines a visual encoder such as Vision Transformer (ViT) [16] or ResNet18 [25] with a Multi-Layer Perceptron (MLP) for direct action prediction [1, 8, 21, 29, 34] . With perception and decision-making collapsed into a single opaque mapping, these architectures lack interpretable intermediate representations, making it difficult to explain or audit their decisions. Moreover, prior work [13, 15, 20, 27] shows that direct visual-to-action learning often exploits spurious correlations rather than transferable physical principles. As a result, agents trained on benchmarks such as PHYRE [8] exhibit brittle behavior under distributional shifts: strong performance on seen tasks comes at the expense of generalization.  \nVision-language models (VLMs) [18, 35] offer a fundamentally different paradigm by replacing reactive mappings with explicit causal reasoning. Through chain-of-thought (CoT) reasoning [26, 40, 41], VLMs can form causal interpretations of physical dynamics, enabling stronger cross-task and cross-environment generalization [47, 49] . However, the two dominant training paradigms, Supervised Fine-Tuning (SFT) [11] and reinforcement learning optimized solely for task success, each introduce structural limitations that hinder this potential [24, 45] .  \n∗ denotes co-first author.  \nPreprint.  \nFigure 1: Two major obstacles in CoT-based physical reasoning. Hallucinated CoT denotes the model producing physically incorrect reasoning that leads to wrong action; Misaligned Action, on the other hand, bypasses physically-aligned reasoning via a visual shortcut and results in wrong action planning. Our VAORA aims to resolve both issues, generating successful action with proper physical reasoning.  \nPrior work shows that SFT primarily teaches models to imitate the linguistic form of expert reasoning without grounding the reasoning process in physical reality [11, 28, 36, 39] . In contrast, successdriven RL often encourages mode","cbCaiiJG7DmJploB","https://ap.wps.com/l/cbCaiiJG7DmJploB","pdf",2102335,5,1,26,"English","en",105,"# Abstract\n# Introduction\n## Failure Modes in CoT-Based Physical Reasoning\n## VAORA Reward Design\n## Training Stability and Reward Estimation\n## Experimental Evaluation and Generalization Results","[{\"question\":\"What problem does VAORA address in physical reasoning with vision-language models?\",\"answer\":\"VAORA addresses two issues: hallucinated chain-of-thought reasoning that conflicts with physical reality, and misalignment between reasoning and the actions that the model plans to take.\"},{\"question\":\"How does VAORA’s Visual Alignment Reward work?\",\"answer\":\"Visual Alignment Reward anchors reasoning to the visual context independent of the model’s action, which suppresses hallucinated reasoning at its source.\"},{\"question\":\"How is training stability improved in VAORA?\",\"answer\":\"VAORA uses smooth, dense reward signals by estimating success probabilities with a pre-trained in-domain expert agent, helping with sparse and noisy rewards in interactive physical reasoning.\"}]",1784193124,66,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"bridging-physical-reasoning-and-task-generalization-via-visual-action-outcome-reasoning-alignment","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/bridging-physical-reasoning-and-task-generalization-via-visual-action-outcome-reasoning-alignment/84125/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does VAORA address in physical reasoning with vision-language models?","Question",{"text":76,"@type":77},"VAORA addresses two issues: hallucinated chain-of-thought reasoning that conflicts with physical reality, and misalignment between reasoning and the actions that the model plans to take.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does VAORA’s Visual Alignment Reward work?",{"text":81,"@type":77},"Visual Alignment Reward anchors reasoning to the visual context independent of the model’s action, which suppresses hallucinated reasoning at its source.",{"name":83,"@type":74,"acceptedAnswer":84},"How is training stability improved in VAORA?",{"text":85,"@type":77},"VAORA uses smooth, dense reward signals by estimating success probabilities with a pre-trained in-domain expert agent, helping with sparse and noisy rewards in interactive physical reasoning.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]