[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81726-en":3,"doc-seo-81726-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81726,4810365810221,"Aurora","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Joint Discovery of Object and Action Symbols through Effect Prediction for Robotic Manipulation Planning","The document presents a model for robotic manipulation planning that abstracts continuous, high-dimensional sensorimotor interactions into discrete object and action symbols. It jointly discovers high-level manipulation primitives and object categories via a binary bottleneck, trained with random interactions to predict multi-modal outcomes such as object motion, contact, and force feedback. Using the learned binary representations, a discrete planner exploits intermediate effect-trajectory steps for partial action execution and precise low-level control, and enables few-shot generalization on novel objects based on behavior rather than visual similarity.","Joint Discovery of Object and Action Symbols through Effect Prediction for Robotic Manipulation Planning  \nBurcu Kilic 1 , Berke Kartal 1 , Fatih Dogangun 1 , Erhan Oztop2,3 , Emre Ugur 1  \narXiv :2607 .00031v1 [ cs .RO] 22 Jun 2026  \nAbstract—To perform complex manipulation planning, autonomous robots are required to abstract continuous, highdimensional sensorimotor interactions into discrete object and action representations. Earlier work either categorized objects based on visual appearances, which fails to distinguish objects that appear similar but behave differently, or based on effects under interaction, but was limited to predefined actions. To address these limitations, we propose a model that jointly discovers highlevel manipulation primitives and object categories through a binary bottleneck layer, trained to predict multi-modal outcomes, including object motion, contact, and force feedback, from random interaction data. Building on these discovered binary representations, we leverage a discrete planning method that uses intermediate steps in the predicted effect trajectory to enable partial action executions for precise low-level control. Additionally, we evaluate our framework’s generalization capabilities on novel objects by assigning object categories through comparing a small number of interaction effects with the predicted effects of learned object symbols, enabling few-shot generalization based on behavior rather than visual similarity. We conduct experiments on tabletop repositioning and stacking tasks, and confirm that our effect-driven planning approach outperforms both a state-of-theart method and a visual-based alternative in planning precision across seen and novel objects.  \nIndex Terms—Robot learning, Representation learning, Autonomous mental development, Few shot learning  \nI. INTRODUCTION  \nAutonomous robot development aims to build intelligent agents capable of learning through own experience, reasoning about their surroundings, and adapting to novel scenarios [1] . However, a fundamental challenge is that reasoning over highdimensional continuous sensorimotor spaces becomes unmanageable in long-horizon manipulation tasks. Operating in such complex sensorimotor spaces requires a layer of abstraction on both percepts and actions [2], [3] . For example, an infant that stacks multiple blocks on top of each other does not plan low-level joint movements directly from environmental pixels; rather, it thinks in terms of high-level intentions like picking a block and placing it on top of another one.  \nPrior studies either manually constructed such abstractions [4] or used reconstruction-based clustering [5], ignoring the role of interaction. However, Sun [6] argued that the concepts of objects and skills emerge directly from the agent’s interaction with the world and are linked to their goals and needs. Following this principle, an autonomous robot must discover its own representations by continuously exploring and manipulating objects, while learning to make sense of  \nThis work was supported by the Scientific and Technological Research Council of Turkey (TUBITAK) ARDEB 1001 Program (124E227) and by the INVERSE Project under Grant 101136067, funded by the European Union.  \n1Bogazici University, 2 Ozyegin University, 3 Osaka University.  \nthe physical dynamics. For this physical exploration to be effective, the robot’s interactions cannot be limited to purely visual observations. Multi-modal feedback such as force and contact can convey complementary information to visual feedback about action outcomes, an observation also supported by developmental studies showing that young infants learn action properties through both visual and tactual interactions with objects [7] . Motivated by this, our approach allows a robot to autonomously categorize objects and learn highlevel manipulation skills directly from multi-modal physical interactions.  \nThrough continuous physical interactions, children learn to categ","cbCaijo3Ubk7yWxu","https://ap.wps.com/l/cbCaijo3Ubk7yWxu","pdf",5906532,4,1,10,"English","en",105,"# Introduction\n## Problem of long-horizon sensorimotor reasoning\n## Limitations of existing object/skill abstractions\n## Motivation for multimodal interaction and effect-based learning\n## Temporal structure and planning precision","[{\"question\":\"What core challenge does the work address in robotic manipulation planning?\",\"answer\":\"It addresses how to make high-dimensional, continuous sensorimotor reasoning tractable for long-horizon manipulation by abstracting interactions into discrete object and action representations.\"},{\"question\":\"How does the proposed model learn object categories and manipulation primitives?\",\"answer\":\"It uses a joint discovery model with a binary bottleneck layer trained to predict multi-modal interaction outcomes (e.g., motion, contact, and force feedback) from random interaction data.\"},{\"question\":\"How does the approach enable planning and generalization on novel objects?\",\"answer\":\"The discrete planning method leverages intermediate steps in the predicted effect trajectory to execute partial actions with precise low-level control, and it assigns object categories by matching a few observed interaction effects to predicted effects for few-shot generalization based on behavior.\"}]",1784175662,25,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"joint-discovery-of-object-and-action-symbols-through-effect-prediction-for-robotic-manipulation-planning","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/joint-discovery-of-object-and-action-symbols-through-effect-prediction-for-robotic-manipulation-planning/81726/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What core challenge does the work address in robotic manipulation planning?","Question",{"text":75,"@type":76},"It addresses how to make high-dimensional, continuous sensorimotor reasoning tractable for long-horizon manipulation by abstracting interactions into discrete object and action representations.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed model learn object categories and manipulation primitives?",{"text":80,"@type":76},"It uses a joint discovery model with a binary bottleneck layer trained to predict multi-modal interaction outcomes (e.g., motion, contact, and force feedback) from random interaction data.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the approach enable planning and generalization on novel objects?",{"text":84,"@type":76},"The discrete planning method leverages intermediate steps in the predicted effect trajectory to execute partial actions with precise low-level control, and it assigns object categories by matching a few observed interaction effects to predicted effects for few-shot generalization based on behavior.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":22,"slug":133},"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]