[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83475-en":3,"doc-seo-83475-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83475,1099513958762,"Logic","https://ap-avatar.wpscdn.com/avatar/1000023916a998db790?x-image-process=image/resize,m_fixed,w_180,h_180&k=1784791008015729253",8,"Research & Report","Unleashing More Actions via Action Compositional Training for VLA Models","Vision-Language-Action (VLA) models enable robots to execute language-instructed manipulation, yet conventional training tends to overfit to demonstration-specific behavioral patterns. Even when out-of-distribution demands only require novel combinations of known sub-skills, policies often fail due to missing composition coverage. Expanding datasets helps but remains costly because robot demonstrations require labor-intensive teleoperation and physical setup. ACT-VLA tackles this by offline latent-task-driven augmentation that synthesizes physically valid demonstrations for policy training. Simulation evaluations show markedly higher success rates and improved compositional generalization.","Unleashing More Actions via Action Compositional  \nTraining for VLA Models  \nKai Peng∗ School of Artificial Intelligence Shenzhen Technology University Shenzhen, China [2410263047@stumail.sztu.edu.cn](2410263047@stumail.sztu.edu.cn)  \nJie Lu∗ School of Artificial Intelligence Shenzhen Technology University Shenzhen, China [2510263005@stumail.sztu.edu.cn](2510263005@stumail.sztu.edu.cn)  \nXiaojiang Peng† School of Artificial Intelligence Shenzhen Technology University Shenzhen, China [pengxiaojiang@sztu.edu.cn](pengxiaojiang@sztu.edu.cn)  \narXiv :2607 .0035 1v 1 [ cs .RO] 1 Jul 2026  \nAbstract—Vision-Language-Action (VLA) models excel at robotic manipulation, driven by the scale and diversity of demonstration data. However, standard training paradigms often cause VLA models to severely overfit to specific behavioral patterns, rendering them unable to generalize to out-of-distribution scenarios even when those scenarios merely require novel combinations of identical sub-skills. While expanding datasets can mitigate this overfitting, acquiring high-quality robot data remains notoriously labor-intensive and cost-prohibitive. To resolve this impasse without expensive human teleoperation and to truly unleash more actions—i.e., enable VLA models to compose known sub-skills into a much broader set of executable behaviors beyond the original demonstrations—we propose ACT-VLA (Action Compositional Training for VLA Models), an offline data augmentation framework that leverages the model’s latent task representations to synthesize novel, physically valid demonstrations directly from existing tasks for policy training. By eliminating additional manual data collection, our method automatically expands the training distribution and mitigates overfitting. We evaluate our approach on challenging manipulation tasks in simulation. Experiments demonstrate that while baseline VLA models generalize poorly due to original distribution overfitting, policies trained with our synthesized data achieve substantially higher success rates, validating that leveraging existing tasks for automated demonstration synthesis provides an effective, scalable, and dataefficient route to broadening VLA generalization.  \nIndex Terms—Vision-Language-Action models, robotic manipulation, compositional generalization  \nI. INTRODUCTION  \nRobotic manipulation has undergone a fundamental transformation with the rise of large-scale imitation learning [1],[2] . By training on extensive collections of human demonstrations, modern robot learning systems have achieved remarkable dexterity across a broad range of manipulation tasks. Among these approaches, Vision-Language-Action (VLA) models [3]–[8] have emerged as a particularly promising paradigm, integrating visual perception, language understanding, and action generation into a unified framework capable of following natural-language instructions and executing complex robotic behaviors.  \nDespite their impressive performance, the generalization capability of VLA models remains fundamentally constrained by  \nthe diversity and coverage of training data. Existing imitation ∗ These authors contributed equally to this work.  \n†Corresponding author: Xiaojiang Peng ([pengxiaojiang@sztu.edu.cn](pengxiaojiang@sztu.edu.cn))  \nlearning pipelines learn manipulation behaviors directly from demonstrations, causing policies to rely heavily on previously observed task distributions. As shown by recent evaluations of VLA robustness and compositional generalization [9]–[11], models often struggle when confronted with novel combinations of familiar skills. Although individual sub-skills may have been successfully learned, the model frequently fails to execute them in unseen sequences because the corresponding demonstrations were absent during training. Since the number of possible task compositions grows combinatorially with the number of available skills, exhaustive demonstration coverage becomes practically impossible.  \nThis challenge is exace","cbCaitpzhXpqUXlN","https://ap.wps.com/l/cbCaitpzhXpqUXlN","pdf",6507050,5,1,7,"English","en",105,"# Introduction\n## Problem: distribution overfitting and limited composition coverage\n## Motivation: scalable data generation beyond teleoperation\n## Proposed method: ACT-VLA for offline demonstration synthesis\n## Evaluation: simulation results and success-rate gains","[{\"question\":\"What problem does ACT-VLA address in VLA model training?\",\"answer\":\"ACT-VLA addresses overfitting to demonstration-specific behavioral patterns, which limits generalization when tasks require novel combinations of familiar skills that were not composed in training.\"},{\"question\":\"How does ACT-VLA generate additional training data without extra teleoperation?\",\"answer\":\"ACT-VLA uses an offline data augmentation framework that leverages the model’s latent task representations to synthesize new, physically valid demonstrations directly from existing tasks.\"},{\"question\":\"How is ACT-VLA evaluated and what do the results show?\",\"answer\":\"ACT-VLA is evaluated on challenging manipulation tasks in simulation. Policies trained with synthesized data achieve substantially higher success rates than baseline models, validating improved compositional generalization.\"}]",1784188252,18,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"unleashing-more-actions-via-action-compositional-training-for-vla-models","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/unleashing-more-actions-via-action-compositional-training-for-vla-models/83475/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does ACT-VLA address in VLA model training?","Question",{"text":76,"@type":77},"ACT-VLA addresses overfitting to demonstration-specific behavioral patterns, which limits generalization when tasks require novel combinations of familiar skills that were not composed in training.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does ACT-VLA generate additional training data without extra teleoperation?",{"text":81,"@type":77},"ACT-VLA uses an offline data augmentation framework that leverages the model’s latent task representations to synthesize new, physically valid demonstrations directly from existing tasks.",{"name":83,"@type":74,"acceptedAnswer":84},"How is ACT-VLA evaluated and what do the results show?",{"text":85,"@type":77},"ACT-VLA is evaluated on challenging manipulation tasks in simulation. Policies trained with synthesized data achieve substantially higher success rates than baseline models, validating improved compositional generalization.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":20,"slug":137},19,"General","general"]