[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85606-en":3,"doc-seo-85606-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},85606,3848291630094,"Emma Wilson","https://eur-avatar.wpscdn.com/davatar_085a072bc5b1113ac321206ff7593b45",8,"Research & Report","MindClaw: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention","Theory of Mind (ToM) supports agents in inferring another actor’s beliefs, goals, and intentions, enabling human-centered embodied assistance. Existing ToM benchmarks largely assess offline mental-state recognition or final action prediction and rarely test whether an embodied agent can remain linked to a changing environment, update actor-specific beliefs, and decide when reasoning is necessary. MindClaw extends robot-centric ToM into a real-time closed-loop framework, integrating multi-source inputs, belief memory, embodied cognitive triggers, mental reasoning, and action generation.","MindClaw: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention  \nRuoxuan Zhang, Qiaoqiao Wan, Zhengguang Wang, Chenghao Yu, Hongxia Xie , Jianlong Fu, Wen-Huang  \nCheng  \narXiv :2606 .0 1063v2 [ cs .AI] 13 Jul 2026  \nAbstract—Theory of Mind (ToM) enables an agent to reason about another actor’s beliefs, goals, and intentions, which is essential for human-centered embodied assistance. Existing ToM benchmarks have advanced text and multimodal mental-state recognition, but they mostly evaluate offline question answering or final action prediction. They do not fully test whether an embodied agent can stay connected to a changing environment, update actor-specific beliefs, decide when reasoning is needed, and intervene only when help is useful. Building on MindPower, we extend robot-centric ToM reasoning to a real-time closedloop setting and introduce MindClaw, a framework for embodied mental-state reasoning with precision intervention. MindClaw connects multi-source inputs, belief memory, an embodied cognitive trigger skill, mental reasoning, and action generation, allowing the agent to output helpful actions at the right time while remaining silent when intervention is unnecessary. Experiments show that direct VLM baselines struggle with task awareness and intervention calibration, while MindClaw achieves the best overall performance, demonstrating the importance of triggerskill optimization for closed-loop embodied ToM assistance.  \nI. INTRODUCTION  \nTheory of Mind (ToM) is the ability to infer other agents’mental states, including what they believe, desire, and intend. It is a core mechanism behind social reasoning and collaborative behavior [1], [2], [3] . In embodied intelligence, ToMis especially important because an assistant should not only perceive the physical world, but also understand how a human may subjectively interpret that world. The Belief–Desire– Intention (BDI) framework provides a classical computational view of this process: an agent forms beliefs from observations, derives desires from goals, and commits to intentions that guide action [4] . This perspective naturally raises a key question for embodied agents: can they reason about a human’s mental state and use that reasoning to decide when and how to help?  \nRecent multimodal ToM benchmarks have made important progress toward this goal. Text-based benchmarks evaluate social reasoning in narratives, while multimodal benchmarks introduce images or videos to test belief, goal, and intention understanding [5], [6], [7] . However, most existing benchmarks still treat ToM as an offline recognition or questionanswering problem. The model is usually asked to infer a character’s belief or goal after observing a fixed clip. Such a formulation is useful for measuring mental-state recognition,  \nCorresponding author: Hongxia Xie.  \nRuoxuan Zhang, Qiaoqiao Wan, Zhengguang Wang, Chenghao Yu, and Hongxia Xie are with Jilin University. Jianlong Fu is with Microsoft Asia. Wen-Huang Cheng is with National Taiwan University.  \nbut it does not require an embodied assistant to maintain state over time, decide whether reasoning should be activated, or act back into an environment. As a result, models can appear competent at answering ToM questions while still failing to behave as useful interactive helpers.  \nMindPower [8] takes an important step beyond role-centric ToM recognition by introducing a robot-centric reasoning hierarchy that connects perception, belief, desire, intention, decision, and action. It explicitly evaluates whether VLMbased embodied agents can generate decisions and action plans under false-belief correction and implicit-goal completion tasks. This is a major advance over benchmarks that only ask for human mental-state labels. Nevertheless, MindPower still represents assistance mainly as a final decision/action output after a complete reasoning hierarchy. The agent observes the scenario, reasons through the hierarchy, and then producesan ","cbCaiv2LRgwm74Z7","https://ap.wps.com/l/cbCaiv2LRgwm74Z7","pdf",1660580,1,9,"English","en",105,"# Introduction\n## Theory of Mind in embodied intelligence\n## Limits of existing ToM benchmarks\n## MindPower and robot-centric reasoning hierarchy\n## Precision-intervention principle and when to act","[{\"question\":\"What problem does MindClaw address in existing ToM benchmarks?\",\"answer\":\"MindClaw addresses the gap where prior benchmarks mainly test offline recognition or final action prediction, without evaluating continuous state maintenance, adaptive belief updating, selective reasoning activation, and appropriate intervention timing in embodied environments.\"},{\"question\":\"How does MindClaw decide whether to intervene?\",\"answer\":\"MindClaw uses an embodied cognitive trigger skill at each step to decide which internal cognitive operation to perform—such as updating belief memory, invoking mental reasoning, generating actions, or emitting noop—so the agent intervenes only when help is warranted.\"},{\"question\":\"What experimental outcome does the abstract report for MindClaw?\",\"answer\":\"Experiments indicate that direct VLM baselines struggle with task awareness and intervention calibration, while MindClaw achieves the best overall performance, highlighting the importance of optimizing the trigger-skill for closed-loop embodied ToM assistance.\"}]",1784204885,23,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"mindclaw-closed-loop-embodied-mental-state-reasoning-for-precision-intervention","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/mindclaw-closed-loop-embodied-mental-state-reasoning-for-precision-intervention/85606/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-22","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does MindClaw address in existing ToM benchmarks?","Question",{"text":75,"@type":76},"MindClaw addresses the gap where prior benchmarks mainly test offline recognition or final action prediction, without evaluating continuous state maintenance, adaptive belief updating, selective reasoning activation, and appropriate intervention timing in embodied environments.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does MindClaw decide whether to intervene?",{"text":80,"@type":76},"MindClaw uses an embodied cognitive trigger skill at each step to decide which internal cognitive operation to perform—such as updating belief memory, invoking mental reasoning, generating actions, or emitting noop—so the agent intervenes only when help is warranted.",{"name":82,"@type":73,"acceptedAnswer":83},"What experimental outcome does the abstract report for MindClaw?",{"text":84,"@type":76},"Experiments indicate that direct VLM baselines struggle with task awareness and intervention calibration, while MindClaw achieves the best overall performance, highlighting the importance of optimizing the trigger-skill for closed-loop embodied ToM assistance.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]