[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83564-en":3,"doc-seo-83564-105":29,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},83564,34359740700684,"Finn","https://ap-avatar.wpscdn.com/avatar/1f400023980c374ae676?_k=1777273430885731487",8,"Research & Report","AutoSpeed Annotation-Free Stage-Adaptive Motion Speed Learning for Robot Manipulation","AutoSpeed addresses robot manipulation efficiency by learning stage-adaptive motion speeds without speed or stage annotations. Existing imitation learning policies typically match expert demonstration speed and use a fixed temporal prediction horizon, restricting flexibility and throughput across task stages. AutoSpeed generates candidate future trajectories under multiple speed ratios, evaluates them via a composite cost balancing prediction error and horizon length, and optimizes the policy toward the minimum-cost candidate. Frequency-domain modulation using DCT preserves motion continuity, reducing execution time while improving success rates.","arXiv :2607 .0 105 1v2 [ cs .RO] 5 Jul 2026  \nAutoSpeed: Annotation-Free Stage-Adaptive Motion Speed Learning for Robot Manipulation  \nQingda Hu 1 * , Ziheng Qiu 1 * , Jieru Zhao2 , Zhongxue Gan 1 , and Wenchao Ding 1†  \n1 College of Intelligent Robotics and Advanced Manufacturing, Fudan University, Shanghai, China  \n2 School of Computing, Shanghai Jiao Tong University, Shanghai, China  \n* Equal contribution. †Corresponding author: [dingwenchao@fudan.edu.cn](dingwenchao@fudan.edu.cn)  \nAbstract. Different stages of manipulation tasks exhibit varying levels of difficulty, suggesting stage-dependent motion speeds and temporal prediction horizons. However, existing IL-based visuomotor policies typically imitate the execution speed of expert demonstrations and operate with a fixed temporal prediction horizon, limiting flexibility and overall task throughput. In this paper, we introduce AutoSpeed, a modelagnostic learning framework that enables existing visuomotor policies to predict trajectories with stage-adaptive motion speeds, without requiring speed or stage annotations. We treat future trajectories at different speeds as candidate optimization targets, evaluate each candidate using a composite cost that trades off prediction error against prediction horizon, and optimize the policy toward the minimum-cost candidate.  \nWith a fixed-length action sequence, speed modulation adjusts the effective temporal prediction horizon: simple stages are executed faster with a longer prediction horizon, whereas complex stages are executed more slowly with a shorter prediction horizon. Specifically, we implement speed modulation in the frequency domain via the discrete cosine transform (DCT), which enables smooth, non-integer speed scaling and thus preserves motion continuity. Extensive evaluations show that AutoSpeed substantially reduces task execution time while also improving success rates. Under the AutoSpeed framework, the inferred motion speeds exhibit a strong correspondence with task stages.  \nKeywords: Imitation Learning · Adaptive Motion Speed · Visuomotor Policy Learning  \n1 Introduction  \nImitation learning (IL) is widely adopted for visuomotor policy learning. In recent years, it has catalyzed a diverse range of visuomotor policies [8, 13, 21, 42] for robotic manipulation, including prominent Vision-Language-Action (VLA) models [3,11,24,39] . Existing IL-based policies typically mimic the motion speed exhibited in expert demonstrations. [6] Regardless of the stage of the task, the  \n2 Q. Hu et al.  \nFig. 1: Stage-aware motion speed adaptation. Motion speeds in expert demonstrations are often suboptimal. AutoSpeed aims to train policies to predict future trajectories with stage-aware motion speed without requiring speed or stage annotations.  \npolicy predicts actions over a fixed future horizon, i.e., it uses fixed-length action chunking [8,40] . However, the motion speeds in expert demonstrations are often suboptimal, thereby limiting the efficiency and performance of the learned policy. In practical deployment scenarios, task completion efficiency is often critical, especially in industrial settings.  \nDifferent stages of manipulation tasks exhibit varying levels of difficulty [7, 33], suggesting that both the motion speed and the temporal prediction horizon should be stage-dependent. Empirically, we find a consistent positive relationship between motion speed and effective prediction horizon, aligning with prior findings in human motor and cognitive control [9, 28] . As shown in Fig. 1, in easy stages (e.g., free-space reaching, or gross repositioning), long-horizon action chunks can be predicted reliably from the current observation [25, 35], allowing faster execution. In contrast, fine-grained stages (e.g., insertion, or sustained interaction) demand higher precision and tighter perception–action coupling, and are therefore better served by shorter effective prediction horizons and slower execution. This is consistent with f","cbCaiv1T3DQzuW77","https://ap.wps.com/l/cbCaiv1T3DQzuW77","pdf",5620592,1,18,"English","en",105,"# Introduction\n## Stage-dependent speed and prediction horizon\n## AutoSpeed framework and candidate supervision targets\n## Frequency-domain speed modulation and Ratio Head","[{\"question\":\"How is motion speed modulation implemented to preserve continuity?\",\"answer\":\"AutoSpeed performs speed modulation in the frequency domain using the discrete cosine transform (DCT). This enables smooth non-integer speed scaling while preserving motion continuity, which supports precise fine-grained manipulation.\"}]",1784188866,45,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":27},"autospeed-annotation-free-stage-adaptive-motion-speed-learning-for-robot-manipulation","",{"@graph":35,"@context":77},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/autospeed-annotation-free-stage-adaptive-motion-speed-learning-for-robot-manipulation/83564/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"How is motion speed modulation implemented to preserve continuity?","Question",{"text":75,"@type":76},"AutoSpeed performs speed modulation in the frequency domain using the discrete cosine transform (DCT). This enables smooth non-integer speed scaling while preserving motion continuity, which supports precise fine-grained manipulation.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,120,123,127],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":45,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":45,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":45,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":45,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":128,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":98,"slug":130},19,"General","general"]