[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86382-en":3,"doc-seo-86382-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86382,34359740700684,"Finn","https://ap-avatar.wpscdn.com/avatar/1f400023980c374ae676?_k=1777273430885731487",8,"Research & Report","Replanning Human-Robot Collaborative Tasks with Vision-Language Models via Semantic and Physical Dual-Correction","Human–robot collaborative assembly demands that robots interpret ambiguous human corrective instructions and turn them into physically executable motions. Vision–language models (VLMs) offer semantic understanding but can choose logically inconsistent targets and misjudge whether actions will succeed. This work presents a replanning framework that converts human instructions into action-target candidates, then applies an Internal Correction Model for pre-execution logical verification and an External Correction Model for post-execution visual verification. Integrated VLM reasoning with 6-DoF grasp generation and collision-free planning, achieving 66.7% real-world success for object fixation and strong tool-selection accuracy.","arXiv :2602 . 14551v2 [ cs .RO] 13 Jul 2026  \nFULL PAPER  \nReplanning Human–Robot Collaborative Tasks with Vision–Language Models via Semantic and Physical Dual–Correction  \nTaichi Katoa , Takuya Kiyokawab* , Namiko Saitob,c , and Kensuke Haradab,da Department of Systems Science, School of Engineering Science, The University of Osaka, 1-3 Machikaneyama, Toyonaka, Osaka, Japan;  \nb Department of Systems Innovation, Graduate School of Engineering Science, The University of Osaka, 1-3 Machikaneyama, Toyonaka, Osaka, Japan;  \nc Microsoft Research Asia, Shinagawa, Tokyo, Japan;  \nd Industrial Cyber-physical Systems Research Center, The National Institute of Advanced Industrial Science and Technology (AIST), 2-3-26 Aomi, Koto-ku, Tokyo, Japan.  \nARTICLE HISTORY  \nCompiled July 14, 2026  \nABSTRACT  \nHuman–robot collaborative assembly requires robots to interpret ambiguous cor  \nrective instructions while producing physically executable motions. Vision–language  \nmodels (VLMs) provide semantic reasoning but may select logically inconsistent  \ntargets or misjudge execution outcomes. We propose a replanning framework that  \nmaps human instructions to Action Target candidates, including grasp poses and  \ntool selections, and combines an Internal Correction Model for pre-execution logical  \nverification with an External Correction Model for post-execution visual verification.  \nThe framework integrates VLM reasoning with 6-DoF grasp generation and collision  \nfree trajectory planning. Simulation ablations show configuration-dependent effects:  \ninternal correction improves candidate validity, whereas external correction enables  \nrecovery for a low-latency VLM but can reduce success when visual verification pro  \nduces false negatives. Experiments with an upper-body humanoid robot achieved  \n66.7% success in real-world object fixation, 100% in initial tool selection, and 75.0%  \nin corrective tool selection. These results demonstrate interactive replanning across  \nspatial and semantic collaborative tasks while identifying visual-state verification as  \na key limitation.  \nKEYWORDS  \nReplanning, Vision–Language Models, Human–Robot Collaboration  \n1. Introduction  \nWhile recent advances in manufacturing industry have advanced rapidly, associated social challenges such as workforce shortages and the demand for flexible production systems have been increasing. To address these issues, a new concept known as Industry 5.0 [1] has been proposed. Unlike conventional manufacturing paradigms that focus on machine-centered automation, Industry 5.0 emphasizes human-centered production activities [2] . Within this paradigm, Human–Robot Collaboration (HRC) is  \n* Corresponding Author: Takuya Kiyokawa. Email: [kiyokawa@sys.es.osaka-u.ac.jp](kiyokawa@sys.es.osaka-u.ac.jp)  \nregarded as a key technology for enabling adaptive and interactive assembly tasks.  \nPrevious studies on HRC have proposed various collaborative approaches from perspectives such as motion planning [3], role allocation [4], and safety assurance [5] . They assume predefined workflows or structured commands. In real-world assembly tasks, however, there are many situations in which humans are responsible for processes that robots cannot perform independently, while robots act in a supportive role. In such scenarios, a cooperative relationship is required in which humans take the initiative and robots provide appropriate assistance to complete the task.  \nOne of the main challenges in human–robot collaborative assembly is enabling robots to understand human intentions and actions. When humans work together, they often use ambiguous verbal instructions, such as “hold it a little more to the left” or “pass me a larger tool,” to adjust each other’s actions and smoothly coordinate the task. Humans naturally resolve such ambiguity through shared context and implicit intention, but enabling robots to do the same remains a fundamental challenge. Many existing systems rely on predefined fixed mo","cbCaijvfCg3Cb72u","https://ap.wps.com/l/cbCaijvfCg3Cb72u","pdf",8051817,5,1,20,"English","en",105,"# Introduction\n## Industry 5.0 and HRC Motivation\n## Core Challenge: Interpreting Ambiguous Human Intent\n## Limitations of Direct VLM Application\n## Proposed Dual-Correction Replanning Framework","[{\"question\":\"What problem does the paper address in human–robot collaborative assembly?\",\"answer\":\"It addresses how robots can interpret ambiguous human corrective verbal instructions and generate actions that are both semantically consistent and physically executable.\"},{\"question\":\"How does the proposed framework use vision–language models (VLMs)?\",\"answer\":\"It maps human instructions to a set of action-target candidates (e.g., grasp poses and tool selections) selected through VLM reasoning.\"},{\"question\":\"What roles do the Internal and External Correction Models play?\",\"answer\":\"The Internal Correction Model verifies logical consistency before execution, while the External Correction Model performs post-execution visual verification to support recovery when outcomes differ from expectations.\"}]",1784211384,50,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"replanning-human-robot-collaborative-tasks-with-vision-language-models-via-semantic-and-physical-dual-correction","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/replanning-human-robot-collaborative-tasks-with-vision-language-models-via-semantic-and-physical-dual-correction/86382/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does the paper address in human–robot collaborative assembly?","Question",{"text":76,"@type":77},"It addresses how robots can interpret ambiguous human corrective verbal instructions and generate actions that are both semantically consistent and physically executable.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does the proposed framework use vision–language models (VLMs)?",{"text":81,"@type":77},"It maps human instructions to a set of action-target candidates (e.g., grasp poses and tool selections) selected through VLM reasoning.",{"name":83,"@type":74,"acceptedAnswer":84},"What roles do the Internal and External Correction Models play?",{"text":85,"@type":77},"The Internal Correction Model verifies logical consistency before execution, while the External Correction Model performs post-execution visual verification to support recovery when outcomes differ from expectations.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":29,"slug":113},6,"Technology","technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":22,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":127,"show_sort_weight":22,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":46,"category_name":135,"show_sort_weight":20,"slug":136},19,"General","general"]