[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84754-en":3,"doc-seo-84754-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84754,137441390410,"Hazel","https://ap-avatar.wpscdn.com/avatar/2000252f4ab5702993?_k=1776741390130283984",8,"Research & Report","Simple-to-Complex Structured Demonstrations for Vision-Language-Action Learning","Vision-Language-Action (VLA) models advance robotic manipulation by unifying visual perception, language understanding, and action generation, yet demonstration collection is often treated as an afterthought. The work argues that how demonstrations are organized directly impacts policy learning efficiency, training stability, and generalization. It proposes a simple-to-complex structured demonstration collection strategy using a dual-arm robot, decomposing tasks into sub-skills, standardizing environments, and ordering data by increasing complexity. Evaluations on block grasping/sorting and towel folding show higher success rates and steadier training than end-to-end trajectory collection.","arXiv :2607 .0459 1v 1 [ cs .RO] 6 Jul 2026  \nSimple-to-Complex Structured Demonstrations for Vision-Language-Action Learning  \nXINCHUAN QIU∗ , Graduate School of Advanced Science and Engineering, Hiroshima University, Japan YI YU∗†, Graduate School of Advanced Science and Engineering, Hiroshima University, Japan  \nVision-Language-Action (VLA) models have demonstrated strong capabilities in robotic manipulation by integrating visual perception, language understanding, and robot action generation. Existing research has primarily focused on improving model architectures, training strategies, and dataset scale, while little attention has been paid to how demonstrations are collected and organized. We identify demonstration organization as a fundamental yet overlooked aspect of imitation learning, as it directly affects policy learning efficiency, training stability, and policy generalization. To address this gap, we propose a simple-to-complex structured demonstration collection strategy for VLA learning using a dual-arm robotic platform. Instead of treating demonstrations as independent task trajectories, our approach systematically organizes data through three general principles: (i) decomposing complex manipulation tasks into progressively learnable sub-skills,(ii) standardizing the interaction environment to reduce unnecessary variability, and (iii) organizing demonstrations according to progressively increasing task complexity. This structured design enables VLA models to first acquire fundamental manipulation skills before learning increasingly complex task compositions, facilitating more effective learning of long-horizon manipulation tasks. We evaluate the proposed strategy on two representative robotic manipulation tasks: block grasping and sorting, and towel folding. Experimental results show consistent improvements in task success rate and training stability compared with the baseline method of directly collecting end-to-end complete task trajectories. These findings highlight demonstration organization as a previously underexplored but important factor in VLA learning and provide practical insights into efficient skill acquisition, scalable dataset construction, and long-horizon robotic manipulation.  \nCCS Concepts: • Computing methodologies → Artificial intelligence; • Computing methodologies → Cognitive robotics; • Computing methodologies → Knowledge representation and reasoning;  \nAdditional Key Words and Phrases: VLA Learning, Imitation Learning, Robot Manipulation, Demonstration Collection, Dual-Arm Robots, Long-Horizon Manipulation  \nACM Reference Format:  \nXinchuan Qiu and Yi Yu. 2026. Simple-to-Complex Structured Demonstrations for Vision-Language-Action Learning. In Proceedings of Make sure to enter the correct conference title from your rights confirmation email (Conference acronym ’XX). ACM, New York, NY, USA, 20 pages. [https://doi.org/XXXXXXX.XXXXXXX](https://doi.org/XXXXXXX.XXXXXXX)  \n1 Introduction  \nVision-Language-Action (VLA) models have recently emerged as a promising paradigm for robotic manipulation by jointly modeling visual observations, language instructions, and robot actions. Benefiting from large-scale robotic  \n∗ These authors contributed equally to this work.  \n†Corresponding author.  \nAuthors’ Contact Information: Xinchuan Qiu, Graduate School of Advanced Science and Engineering, Hiroshima University, Higashi-Hiroshima, Japan, [qiuxinchuan2025@163.com](qiuxinchuan2025@163.com); Yi Yu, Graduate School of Advanced Science and Engineering, Hiroshima University, Higashi-Hiroshima, Japan.  \nPermission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. ","cbCaijn88LeUNjVG","https://ap.wps.com/l/cbCaijn88LeUNjVG","pdf",3467894,3,1,20,"English","en",105,"# Introduction\n## Motivation: Why demonstration organization matters\n## Proposed approach: Simple-to-complex structured collection\n## Evaluation: Dual-arm tasks and results","[{\"question\":\"What problem does the paper identify in existing VLA imitation learning research?\",\"answer\":\"The paper highlights that research mostly focuses on architectures, training strategies, and dataset size, while paying little attention to how robotic demonstrations are collected and organized.\"},{\"question\":\"How does the proposed demonstration collection strategy work?\",\"answer\":\"It organizes demonstrations using three principles: progressively decomposing complex tasks into sub-skills, standardizing the interaction environment, and ordering demonstrations from simpler to more complex task compositions.\"},{\"question\":\"What evidence supports the effectiveness of structured demonstrations?\",\"answer\":\"On a dual-arm platform for block grasping/sorting and towel folding, the structured approach improves task success rate and training stability compared with collecting end-to-end complete task trajectories.\"}]",1784198046,50,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"simple-to-complex-structured-demonstrations-for-vision-language-action-learning","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/simple-to-complex-structured-demonstrations-for-vision-language-action-learning/84754/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-22","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper identify in existing VLA imitation learning research?","Question",{"text":75,"@type":76},"The paper highlights that research mostly focuses on architectures, training strategies, and dataset size, while paying little attention to how robotic demonstrations are collected and organized.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed demonstration collection strategy work?",{"text":80,"@type":76},"It organizes demonstrations using three principles: progressively decomposing complex tasks into sub-skills, standardizing the interaction environment, and ordering demonstrations from simpler to more complex task compositions.",{"name":82,"@type":73,"acceptedAnswer":83},"What evidence supports the effectiveness of structured demonstrations?",{"text":84,"@type":76},"On a dual-arm platform for block grasping/sorting and towel folding, the structured approach improves task success rate and training stability compared with collecting end-to-end complete task trajectories.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":29,"slug":113},6,"Technology","technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":22,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":127,"show_sort_weight":22,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":46,"category_name":135,"show_sort_weight":106,"slug":136},19,"General","general"]