[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83837-en":3,"doc-seo-83837-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83837,2336464648322,"Aria","https://ap-avatar.wpscdn.com/avatar/2200025388227c56fec?_k=1778556882303663488",8,"Research & Report","SEAM Smooth Execution of Action-Chunked Motion for Vision-Language-Action Policies","Vision-Language-Action (VLA) policies that generate actions in fixed-length chunks can suffer multimodal bifurcation, where adjacent chunks sampled from independent Gaussian latents converge to incompatible trajectory modes, causing abrupt discontinuities at chunk boundaries. Existing fixes trade reliability and smoothness against extra inference cost via per-step backpropagation, rejection sampling, or retraining. SEAM (Smooth Execution of Action-chunked Motion) is a training-free inference-time method for flow-matching VLAs using Velocity-guided Loss Steering (VLS). On LIBERO-10 with π0.5, SEAM lowers boundary jerk and transition discontinuity while keeping task success and loop cost near the baseline.","SEAM: Smooth Execution of Action-Chunked Motion for Vision-Language-Action Policies  \nDijia Zhan∗ , Xuemiao Xu†, Jinyi Li∗ , Jie Tang†  \nSouth China University of Technology  \n∗Equal contribution †Corresponding authors  \narXiv :2607 .04609v 1 [ cs .RO] 6 Jul 2026  \nAbstract  \nVision-Language-Action (VLA) policies that execute fixedlength action chunks can exhibit multimodal bifurcation: across-chunk inconsistency in which adjacent chunks generated from independent Gaussian latents can converge to incompatible trajectory modes, producing abrupt discontinuities at chunk boundaries. Existing remedies either require backpropagation through the policy at each denoising step, rely on rejection sampling, or require retraining, each trading computational cost or task reliability for smoother transitions. We propose SEAM (Smooth Execution of Actionchunked Motion), a training-free inference-time method for flow matching VLAs. SEAM exploits a simple synchronousexecution insight: after the robot consumes the executed prefix, the previous chunk’s unexecuted tail is already available as an analytic consistency reference. Its core mechanism, Velocity-guided Loss Steering (VLS), derives a timedependent target from this tail and applies a closed-form correction after each Euler step without backpropagating through the policy network. On LIBERO-10 with π0.5 , SEAM reduces boundary jerk by 28%, reduces chunk transition discontinuity by 27%, preserves baseline-level task success, and keeps denoising-loop cost near the unguided baseline.  \nIntroduction  \nVision-Language-Action (VLA) policies have become central to language-conditioned robotic manipulation, from RT- 1 and RT-2 to OpenVLA, π0 , and π0.5 (Brohan et al. 2022; Zitkovich et al. 2023; Kim et al. 2024; Black et al. 2024; Physical Intelligence et al. 2025) . Given visual observations and a natural-language instruction, these policies generate robot action trajectories that must remain task-directed and physically smooth. Recent flow matching VLAs such as π0.5 (Physical Intelligence et al. 2025) model this generation process as an ODE that transports Gaussian noise into an action trajectory.  \nA key design choice in these policies is action chunking (Zhao et al. 2023; Chi et al. 2025): the model predicts H future actions at once, executes only the first K ≪ H , and then predicts the next chunk. Chunking improves local temporal coherence because each prediction contains a short motion plan, but because only K of the H predicted actions are executed before re-prediction, consecutive chunks share an overlap region of H−K steps. Ideally, the beginning of the new chunk should agree with the unexecuted tail of the  \nprevious one. In practice, however, the two chunks are generated from independent noise samples zn , zn+1 ∼ N (0, I), so their overlap predictions can disagree even when the observations are nearby in time.  \nThis disagreement can manifest as multimodal bifurcation (Figure 1a), a cross-chunk inconsistency in which adjacent chunks choose incompatible trajectory modes. Contactrich manipulation often admits multiple valid local strategies (Shafiullah et al. 2022), and a flow matching policy represents these alternatives by mapping different noise samples to different trajectory modes. Adjacent chunks can therefore select incompatible modes: one chunk may continue an approach from one side of an object, while the next begins a different approach, creating a sharp reversal exactly at the chunk transition (Wang 2026; Black, Galliker, and Levine 2025) . This artifact is not merely cosmetic; it increases jerk and discontinuity atthe chunk boundary, which can introduce large motion discontinuities and reduce execution reliability.  \nExisting cross-chunk remedies occupy different points on a cost–smoothness spectrum. They share a common objective: reducing chunk-boundary artifacts by encouraging the new chunk to agree with the unexecuted tail of the previous chunk, but they instantiate this obj","cbCairxdISN1ZRBT","https://ap.wps.com/l/cbCairxdISN1ZRBT","pdf",9866769,2,1,9,"English","en",105,"# Introduction\n## Motivation: multimodal bifurcation from action chunking\n## Existing cross-chunk remedies and their limitations\n# Proposed method: SEAM overview\n## Velocity-guided Loss Steering (VLS)\n## Empirical results on LIBERO-10","[{\"question\":\"What problem does SEAM address in vision-language-action policies?\",\"answer\":\"SEAM addresses cross-chunk inconsistency caused by action chunking, where independently sampled chunks can align to incompatible trajectory modes and create sharp discontinuities at chunk boundaries.\"},{\"question\":\"How does SEAM improve chunk-boundary smoothness during inference?\",\"answer\":\"SEAM is a training-free, inference-time method that uses Velocity-guided Loss Steering (VLS) to apply an analytic correction after each Euler step, guiding the overlap toward the unexecuted tail of the previous chunk without backpropagating through the policy network.\"},{\"question\":\"What performance impact does SEAM report on LIBERO-10 with π0.5?\",\"answer\":\"SEAM reduces boundary jerk by 28% and chunk transition discontinuity by 27%, preserves baseline-level task success, and keeps denoising-loop computational cost near the unguided baseline.\"}]",1784190883,23,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"seam-smooth-execution-of-action-chunked-motion-for-vision-language-action-policies","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/seam-smooth-execution-of-action-chunked-motion-for-vision-language-action-policies/83837/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-18","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does SEAM address in vision-language-action policies?","Question",{"text":75,"@type":76},"SEAM addresses cross-chunk inconsistency caused by action chunking, where independently sampled chunks can align to incompatible trajectory modes and create sharp discontinuities at chunk boundaries.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does SEAM improve chunk-boundary smoothness during inference?",{"text":80,"@type":76},"SEAM is a training-free, inference-time method that uses Velocity-guided Loss Steering (VLS) to apply an analytic correction after each Euler step, guiding the overlap toward the unexecuted tail of the previous chunk without backpropagating through the policy network.",{"name":82,"@type":73,"acceptedAnswer":83},"What performance impact does SEAM report on LIBERO-10 with π0.5?",{"text":84,"@type":76},"SEAM reduces boundary jerk by 28% and chunk transition discontinuity by 27%, preserves baseline-level task success, and keeps denoising-loop computational cost near the unguided baseline.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]