[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86116-en":3,"doc-seo-86116-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86116,1374391974468,"Eden","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","SegDiff Segmented Trajectory Diffusion for Consistent and Adaptive Robot Manipulation","Imitation learning lets robots learn manipulation skills from demonstrations, but existing methods either suffer from compounding errors in short-horizon continuous action prediction or require external planning in discrete keypose approaches. SegDiff introduces a closed-loop visuomotor policy that decomposes demonstrations into motion segments between keyposes and predicts a continuous trajectory from the current state to the next keypose for long-horizon accuracy. Using diffusion and DDIM inversion, SegDiff adds Dynamic Temporal Ensembling to handle multi-modal actions, reduce discontinuities, and maintain real-time refinement in dynamic environments.","arXiv :2607 . 11027v1 [ cs .RO] 13 Jul 2026  \nSegDiff: Segmented Trajectory Diffusion for Consistent and Adaptive Robot Manipulation  \nHaidong Cao1,2, Wenjun Cao1,2, Quanhao Li1,2, Sicheng Xie1,2 Zhiying Du1,2, Jiaqi Leng1,2, Zuxuan Wu1,2,†, Yu-Gang Jiang1,2  \n1Institute of Trustworthy Embodied AI, Fudan University, China  \n2Shanghai Key Laboratory of Multimodal Embodied AI, China  \n†Corresponding author  \nAbstract  \nImitation learning enables robots to acquire manipulation skills from demonstrations by mapping observations to actions. Existing approaches predict either short-horizon continuous action sequences or discrete keyposes. However, continuous prediction methods suffer from compounding errors due to short prediction horizons and struggle with multi-modal action distributions, whereas keypose-based methods necessitate an external planner, constraining real-time applicability. To address these challenges, we introduce SegDiff, a closed-loop visuomotor policy that integrates the strengths of both paradigms. SegDiff decomposes demonstrations into motion segments between keyposes and learns to predict the continuous trajectory from the current state to the next keypose, enabling long-horizon prediction with real-time refinement. Furthermore, we leverage the capability of diffusion models and DDIM inversion to propose a Dynamic Temporal Ensembling mechanism, which allows the policy to efficiently respond to dynamic environments and mitigate discontinuities caused by inconsistent multi-modal sampling. SegDiff demonstrates significant performance gains over existing approaches across various simulated and real-world scenarios, indicating its strong ability to reason over extended temporal dependencies while maintaining real-time adaptability and control stability.  \n1 Introduction  \nRobotic manipulation is a critical skill for intelligent agents, yet developing robust policies remains challenging due to the complexity of the physical world. Imitation learning (IL) [2, 32, 35, 36, 48], which learns directly from expert demonstrations, has thus emerged as a powerful and pragmatic paradigm. Offline imitation learning (OIL) [1, 2, 7, 27] further improves data efficiency and safety by learning from fixed expert datasets without costly real-world interaction. Among OIL approaches, behavior cloning (BC) [2] remains the most fundamental and widely used paradigm. Traditional BC predicts the immediate action from the current observation, but even small deviations between the predicted and expert actions caused by perception noise or imperfect policy approximation can quickly accumulate and drive the robot into out-of-distribution (OOD) states.  \nRecent approaches [4, 5, 9, 24, 42, 50, 51, 54] extend the prediction horizon by predicting short action  \n(b) Keypose Prediction (e.g. PerAct, RVT)  \nPredict Keypose Only  \nExecute via Path Planner  \nStart  Keypose  End  Action to execute  Action for buffer  \nFigure 1 Comparison of conventional continuous prediction methods, keypose prediction methods, and SegDiff. SegDiff combines the advantages of both paradigms by enforcing keypose constraints within continuous trajectory generation via Segmented Trajectory Modeling, while improving action consistency and real-time responsiveness through Dynamic Temporal Ensembling.  \nsequences as shown in Fig. 1(a), with some employing temporal ensembling [5, 54] between consecutive predictions to smooth motions. Although these methods improve short-term consistency, they still suffer from compounding errors over extended durations. Another major challenge arises from the multi-modal nature of the action distribution in robotic manipulation tasks. The same observation can correspond to multiple valid action modes such as grasping an object from different directions. Averaging across such modes can yield ambiguous actions and cause unsafe behaviors.  \nKeypose-based prediction methods [11, 13, 14, 18, 23, 38, 49] mitigate compounding errors by focusing on di","cbCaimOeATsdRvfl","https://ap.wps.com/l/cbCaimOeATsdRvfl","pdf",3347640,5,1,26,"English","en",105,"# Introduction\n## Problem of continuous and keypose prediction\n## SegDiff overview and key ideas\n# Method: Segmented Trajectory Modeling\n## Receding horizon control and action buffer alignment\n## Dynamic Temporal Ensembling via DDIM inversion","[{\"question\":\"What limitations do current continuous action prediction methods face in robot manipulation?\",\"answer\":\"They use short prediction horizons, so small deviations can compound over time. They also struggle with multi-modal action distributions, which can cause unsafe averaging across valid action modes.\"},{\"question\":\"How does SegDiff combine continuous prediction with keypose constraints?\",\"answer\":\"SegDiff decomposes demonstrations into motion segments between keyposes and learns to predict the continuous trajectory from the current state to the next keypose. This uses keyposes as anchors to reduce compounding errors while keeping continuous control.\"},{\"question\":\"How does SegDiff remain responsive and consistent in dynamic environments?\",\"answer\":\"It uses receding horizon control, executing only the initial trajectory portion and refining via Dynamic Temporal Ensembling. DDIM inversion helps validate and update the action buffer, discarding outdated predictions to mitigate discontinuities from inconsistent multi-modal sampling.\"}]",1784208617,66,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"segdiff-segmented-trajectory-diffusion-for-consistent-and-adaptive-robot-manipulation","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/segdiff-segmented-trajectory-diffusion-for-consistent-and-adaptive-robot-manipulation/86116/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What limitations do current continuous action prediction methods face in robot manipulation?","Question",{"text":76,"@type":77},"They use short prediction horizons, so small deviations can compound over time. They also struggle with multi-modal action distributions, which can cause unsafe averaging across valid action modes.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does SegDiff combine continuous prediction with keypose constraints?",{"text":81,"@type":77},"SegDiff decomposes demonstrations into motion segments between keyposes and learns to predict the continuous trajectory from the current state to the next keypose. This uses keyposes as anchors to reduce compounding errors while keeping continuous control.",{"name":83,"@type":74,"acceptedAnswer":84},"How does SegDiff remain responsive and consistent in dynamic environments?",{"text":85,"@type":77},"It uses receding horizon control, executing only the initial trajectory portion and refining via Dynamic Temporal Ensembling. DDIM inversion helps validate and update the action buffer, discarding outdated predictions to mitigate discontinuities from inconsistent multi-modal sampling.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]