[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82423-en":3,"doc-seo-82423-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82423,13056703020460,"Valentina","https://ap-avatar.wpscdn.com/avatar/be000253dac470eee5d?_k=1778207105932848923",8,"Research & Report","PAC-ACT: Post-training Actor-Critic for Action Chunking Transformers","Precision industrial contact manipulation demands robot policies that remain deterministic, safe under contact-force constraints, and robust to pose perturbations, especially in high-frequency real-time control. Vision-language-action models generalize well but often incur excessive inference and GPU-memory overhead, limiting practical deployment. Action-chunking vision-action policies offer low-latency, temporally continuous behavior, yet their behavior-cloning training suffers from error accumulation under distribution shift. PAC-ACT introduces a reinforcement-learning post-training framework to improve task completion, contact stability, and force safety while retaining efficient chunking inference and low resource usage.","PAC-ACT: Post-training Actor-Critic for Action  \nChunking Transformers  \nYujie Pang  \nDept. of Mechanical and Energy Engineering  \nSouthern University of Science and Technology Shenzhen, China 0009-0003-6377-6949  \nZudong Li  \nDept. of Mechanical and Energy Engineering  \nSouthern University of Science and Technology Shenzhen, China 0009-0009-5574-9602  \narXiv :2607 .09590v1 [ cs .RO] 10 Jul 2026  \nAbstract—Precision industrial contact manipulation requires robots to complete tasks reliably under pose perturbations and contact-force constraints, placing strong demands on policy determinism, contact-force safety, and disturbance resistance. Vision-language-action models provide strong task generalization, but they usually introduce higher inference and GPU-memory overhead, which still limits deployment in high-frequency control scenarios. In contrast, vision-action chunking policies provide low-latency inference and temporally continuous actions, making them more suitable for industrial scenarios that require predictable and real-time behavior. However, these policies are mainly trained through behavior cloning, and error accumulation under distribution shift limits further improvements in task success and contact stability. To address this issue, this paper proposes PAC-ACT, a reinforcement-learning post-training framework for pretrained action-chunking policies. The core idea is to improve task completion and contact stability through reinforcement learning while preserving the lightweight and efficient properties of vision-action chunking models. Specifically, PAC-ACT reformulates step-wise policy optimization asa chunk-level decision process, aligning reinforcement-learning updates with action-chunk generation; it further builds policy and value-network architectures adapted to pretrained networksand introduces a hybrid behavior-prior constraint to prevent the policy from drifting away from the action distribution learned by the pretrained policy. Experiments on industrially motivated precision-contact benchmarks show that the method improves task success, contact stability, and force safety while maintaining low latency and low GPU-memory usage. In particular, on the Contour task, the median peak contact force is significantly reduced, and the proportion of force readings above 60N is reduced by 46 times. An additional sparse-reward ablation verifies the necessity of key framework components and shows that PACACT can still explore effectively under randomized initial poses when only the task-success reward is retained together with the behavior-prior constraint.  \nIndex Terms—Visuomotor control, action chunking, reinforcement learning fine-tuning, imitation learning, precision contact manipulation  \nI. INTRODUCTION  \nModel-based or vision-servoed manipulation pipelines often depend on accurate camera calibration, hand-eye calibration, and object or end-effector pose estimation. Errors in these calibration and localization steps can propagate through the visualservoing control law and degrade positioning accuracy, which  \nbecomes particularly problematic in precision contact tasks [1], [2] . In recent years, end-to-end visuomotor policies represented by Action Chunking Transformer (ACT) and Diffusion Policy have directly mapped high-dimensional visual observations to low-dimensional actions, reducing the need for explicit geometric action design in several manipulation settings [3], [12] . In particular, ACT predicts multiple future action steps at once through action chunking, improving temporal continuity and execution stability as reported in fine-grained bimanual manipulation experiments [3] . This paradigm has been extended to bimanual manipulation [3], multi-task robot learning [23], and large-scale vision-language-action (VLA) models [14], [15], demonstrating the broad applicability of action chunking in robot control.  \nHowever, such action chunking policies are still fundamentally trained under the offline behavior ","cbCaifuDuKVyMonx","https://ap.wps.com/l/cbCaifuDuKVyMonx","pdf",7296589,1,12,"English","en",105,"# Abstract\n# Introduction\n## Motivation: calibration/localization errors and contact constraints\n## Action chunking and its limitations under behavior cloning\n## Challenges of chunk-level reinforcement learning and fine-tuning\n# Proposed method: PAC-ACT (post-training framework)","[{\"question\":\"Why are vision-language-action models limited in high-frequency industrial control?\",\"answer\":\"They can introduce higher inference latency and GPU-memory overhead, which restricts deployment in real-time settings requiring frequent control updates.\"},{\"question\":\"What main limitation affects action-chunking policies trained with behavior cloning?\",\"answer\":\"Error accumulation under distribution shift can cause failures during long-horizon execution, which is especially problematic in precision contact tasks with millimeter-level pose sensitivity and force safety requirements.\"},{\"question\":\"How does PAC-ACT address contact stability and force safety during post-training?\",\"answer\":\"PAC-ACT reformulates optimization as a chunk-level decision process aligned with chunk generation, adapts policy/value networks to the pretrained model, and uses a hybrid behavior-prior constraint to prevent drift from the pretrained action distribution.\"}]",1784180288,30,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"pac-act-post-training-actor-critic-for-action-chunking-transformers","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/pac-act-post-training-actor-critic-for-action-chunking-transformers/82423/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why are vision-language-action models limited in high-frequency industrial control?","Question",{"text":75,"@type":76},"They can introduce higher inference latency and GPU-memory overhead, which restricts deployment in real-time settings requiring frequent control updates.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What main limitation affects action-chunking policies trained with behavior cloning?",{"text":80,"@type":76},"Error accumulation under distribution shift can cause failures during long-horizon execution, which is especially problematic in precision contact tasks with millimeter-level pose sensitivity and force safety requirements.",{"name":82,"@type":73,"acceptedAnswer":83},"How does PAC-ACT address contact stability and force safety during post-training?",{"text":84,"@type":76},"PAC-ACT reformulates optimization as a chunk-level decision process aligned with chunk generation, adapts policy/value networks to the pretrained model, and uses a hybrid behavior-prior constraint to prevent drift from the pretrained action distribution.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":28,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]