[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86242-en":3,"doc-seo-86242-105":30,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86242,1374391974585,"Genevieve","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","EDAR: Learning Environment-Dependent Action Representations for Robotic Manipulation","Learning effective action representations is critical for robotic manipulation, where raw control trajectories are often noisy, redundant, and hard to model directly. Existing approaches mostly encode the internal structure of action streams while treating environment roles as implicit. Manipulation instead requires environment-dependent semantics because identical action segments can yield different outcomes across scene contexts. EDAR proposes environment-dependent action tokens grounded in executable control structure and expected visual consequences, improving downstream policy learning in long-horizon manipulation on simulated and real benchmarks.","EDAR: Learning Environment-Dependent Action Representations for Robotic  \nManipulation  \nYuecheng Xu 1†, Tong Yang2 ,3†, Jingkai Jia 1 , Chi Zhang3 ,  \nXuelong Li3‡, Wenqiang Zhang 1 ,2‡  \n1 College of Intelligent Robotics and Advanced Manufacturing, Fudan University  \n2 Shanghai Key Lab of Intelligent Information Processing,  \nCollege of Computer Science and Artificial Intelligence, Fudan University  \n3TeleAI, China Telecom  \n{ycxu25, [tongyang23](tongyang23}@m.fudan.edu.cn)[}](tongyang23}@m.fudan.edu.cn)[@m.fudan.edu.cn](tongyang23}@m.fudan.edu.cn) , [wqzhang@fudan.edu.cn](wqzhang@fudan.edu.cn)  \narXiv :2607 . 11427v1 [ cs .RO] 13 Jul 2026  \nAbstract  \nLearning effective action representations is critical for robotic manipulation, where raw control trajectories are often noisy, redundant, and difficult to model directly. Existing methods mainly encode the structure of the action stream itself, treating the role of actions in the environment as implicit. Yet manipulation is about changing the world: the same action segment can induce different outcomes under different scene contexts, making action semantics inherently environment-dependent. We propose EDAR, an Environment-Dependent Action Representation that grounds action tokens in both executable control structure and expected visual consequences. By coupling motor commands with their environment-conditioned effects, EDAR encourages the learned action space to capture interaction semantics rather than merely command-level patterns. Experiments on simulated and real-robot manipulation benchmarks demonstrate that EDAR improves downstream policy learning, especially in long-horizon manipulation. These results highlight the importance of grounding action representations in executable control structure and environment-conditioned visual change.  \n1. Introduction  \nRaw action trajectories, especially those collected from demonstrations or teleoperation, are often noisy, redundant, and uneven. They contain high-frequency jitter, taskirrelevant micro-corrections, and embodiment-specific control patterns, making them difficult to model directly as a reliable interface for policy learning. This issue becomes more pronounced in manipulation, where policies must gen-  \n†Equal contribution. ‡Corresponding authors.  \nerate temporally coherent and physically feasible behaviors over long horizons. As a result, learning effective action representations has become an important problem in robotic manipulation. A good action representation should provide a compact and learnable interface between perception and control, while preserving the information necessary for accurate and robust task execution. The central question is therefore: what kind of action representation is appropriate for robotic manipulation?  \nMost existing action representations [24, 33] are learned primarily from the action stream itself. They may differ inform, ranging from compact trajectory codes to action tokens or frequency-based representations, but they typically share a trajectory-structure-based view: the goal is to preserve or regularize the organization of the control trajectory, such as its geometry, temporal variation, smoothness, or predictability. This view is useful for reducing redundancy and making actions easier to model, but it remains largely action-only. It describes how a command sequence is structured, while leaving its role in the environment implicit. This is limiting for manipulation, where an action is not merely a motor signal but a means of changing the world. The same gripper-closing and rightward-motion segment may be inconsequential in free space, grasp an object under contact, or close a drawer when aligned with a handle, as shown in Fig. 1. Such cases suggest that action semantics are not determined by the command alone, but by the environment-conditioned transition it induces. Therefore, an appropriate action representation for manipulation should not only encode the execution pattern of ","cbCaiilWzI8yLjpm","https://ap.wps.com/l/cbCaiilWzI8yLjpm","pdf",19259559,4,1,19,"English","en",105,"# Introduction\n## Problem: action-only representations\n## Proposed method: EDAR\n## Experimental evaluation","[{\"question\":\"How does EDAR address environment-dependent action semantics?\",\"answer\":\"EDAR grounds action tokens in both executable control structure and expected visual consequences. It couples motor commands with environment-conditioned effects so that action embeddings reflect interaction semantics rather than command-level similarity alone.\"}]",1784209737,48,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":28},"edar-learning-environment-dependent-action-representations-for-robotic-manipulation","",{"@graph":36,"@context":77},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/edar-learning-environment-dependent-action-representations-for-robotic-manipulation/86242/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"How does EDAR address environment-dependent action semantics?","Question",{"text":75,"@type":76},"EDAR grounds action tokens in both executable control structure and expected visual consequences. It couples motor commands with environment-conditioned effects so that action embeddings reflect interaction semantics rather than command-level similarity alone.","Answer","https://schema.org",{"og:url":52,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,120,123,127],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":46,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":98,"slug":129},"General","general"]