[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86104-en":3,"doc-seo-86104-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86104,1374391974468,"Eden","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Affordance-Based Manipulation Planning with Text Goals and Sim-to-Real Generalisation via Real-to-Sim Image Conversion","Affordance-based manipulation planning is developed by combining affordance recognition with action-effect prediction. The system evaluates candidate action plans by comparing predicted multimodal outcomes with text goals specified at run time through a multimodal goal-matching module. Predicted object positions are maintained across future visual states, enabling planning under occlusion and when goal-referenced descriptors no longer identify objects. Real-to-sim image conversion improves physical-robot applicability, and module-level and integrated evaluations are conducted in simulation and on hardware.","Affordance-Based Manipulation Planning with Text Goals and Sim-to-Real Generalisation via Real-to-Sim Image Conversion  \nSolvi Arnold, Rin Karashima, Tadashi Adachi, Takafumi Mochizuki, Kimitoshi Yamazaki  \nAbstract—We present a manipulation planning system based on affordance recognition and action effect prediction. The system reasons through possible futures in visual form, and evaluates candidate plans by agreement of predicted outcomes with text-based goals set at run-time, using a multimodal goal-matching module. Positions of objects named in the goal text are tracked through predictions even when occluded, making it possible to generate action plans even when objects become occluded, or when their initial descriptors cease to identify them in future states. We further expand the system with an image conversion module for translating realworld state images with objects of varied shapes and visual appearances into a consistent visual appearance, to facilitate manipulation planning in a physical robot setup. We evaluate performance of the system ’s modules in isolation, and demonstrate the integrated system ’s manipulation planning capabilities on a set of challenging tasks in both simulation and on hardware.  \nIndex Terms— Manipulation planning, Multi-modal models, Affordances, Natural language instruction, Neural networks  \nI. INTRODUCTION  \nACTION planning has traditionally often been ad  \ndressed using symbolic methods, and more recently using LLMs. Such methods can be effective if a symbolic or linguistic description can adequately convey all the relevant features of the task environment. However, this is rarely the case in real-world manipulation scenarios, limiting the planning fidelity we can achieve using symbolic methods.  \nFurthermore, when effect rules are defined manually, planning ability is constrained by the imagination of the rule-designer, resulting in a system than cannot reliably avoid or exploit effects and side-effects that the designer overlooked. Symbolic reasoning certainly has a role in human action planning, but our reasoning is built on top of a more granular understanding of the physical world.  \nIn a robotics context, which actions are available in a given situation intricately depends on the robot’s physical characteristics and limitations. Furthermore, for an action to be executable and reliably produce the intended effects, it has to be parametrised in a manner that is translatable to unambiguous mo  \nSA is at NICT, Japan. Work performed while at Shinshu University, Japan.([solvi.arnold@nict.go.jp](solvi.arnold@nict.go.jp)). RK, TA, TM are at EPSON Avasys, Japan. KY is at Tohoku University, Japan. Work performed while at Shinshu University, Japan.  \ntions of the target robot platform. Hence, for a robot to use its action repertoire fully and reliably, it must be able to 1) recognise which actions are available to it in a given scene (i.e. recognise the scene ’s affordances), and 2) accurately anticipate the effects of those actions.  \nThe abovementioned concept of affordances [1] has proven useful for structuring robot behaviour. Assigning regions of robots’ (typically continuous) state spaces to qualitatively distinct affordances facilitates the formation of meaningful relations between perception and action spaces. However, the effects of executing affordances so defined are less straightforward to model. We may assume to know our robot in detail, but the same assumption can generally not be made for the environments in which it will operate. Similar observations apply to biological agents. For humans and other animals, nature ’s approach has been to equip us with the ability to learn the effects ofour actions from experience.  \nBased on the above, [2] introduced a manipulation planning method combining recognition of prespecified affordances with learned prediction of affordance effects, both implemented as neural networks (NNs) . The present work builds on [2], addressing some of its li","cbCaitDatF2qf5rJ","https://ap.wps.com/l/cbCaitDatF2qf5rJ","pdf",2288241,4,1,14,"English","en",105,"# Introduction\n# Related Work","[{\"question\":\"How does the system use affordances for manipulation planning?\",\"answer\":\"It first recognizes which affordances are available in a scene, then uses learned prediction of how actions will affect the physical state. Planning is grounded in these affordance-action effect relationships rather than only symbolic rules.\"},{\"question\":\"How are text goals incorporated during planning?\",\"answer\":\"Text-based goals are set at run time and matched against predicted outcomes using a multimodal goal-matching module. Candidate plans are evaluated by the agreement between predicted effects and the text goal constraints.\"},{\"question\":\"What enables sim-to-real generalisation in this approach?\",\"answer\":\"An image conversion module translates real-world state images with varied object appearances into a consistent visual appearance for planning. This supports reliable manipulation planning in a physical robot setup, validated in both simulation and on hardware.\"}]",1784208539,35,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"affordance-based-manipulation-planning-with-text-goals-and-sim-to-real-generalisation-via-real-to-sim-image-conversion","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/affordance-based-manipulation-planning-with-text-goals-and-sim-to-real-generalisation-via-real-to-sim-image-conversion/86104/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"How does the system use affordances for manipulation planning?","Question",{"text":75,"@type":76},"It first recognizes which affordances are available in a scene, then uses learned prediction of how actions will affect the physical state. Planning is grounded in these affordance-action effect relationships rather than only symbolic rules.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How are text goals incorporated during planning?",{"text":80,"@type":76},"Text-based goals are set at run time and matched against predicted outcomes using a multimodal goal-matching module. Candidate plans are evaluated by the agreement between predicted effects and the text goal constraints.",{"name":82,"@type":73,"acceptedAnswer":83},"What enables sim-to-real generalisation in this approach?",{"text":84,"@type":76},"An image conversion module translates real-world state images with varied object appearances into a consistent visual appearance for planning. This supports reliable manipulation planning in a physical robot setup, validated in both simulation and on hardware.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]