[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85139-en":3,"doc-seo-85139-105":30,"detail-sidebar-cat-0-en-105":95},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85139,1374391974468,"Eden","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Maximizing Human Efficiency in Large-Scale Robot Post-Training via VLAC-CUT Guided Pipeline","Adapting Vision Language Action (VLA) models to downstream tasks requires multiple rounds of post-training because early data cannot cover all edge cases. The report defines human efficiency as policy improvement and task throughput gained per unit of human labor and time, and presents a human-efficient pipeline where a Teleoperator performs high-value remote interventions and recovery demonstrations while a Floor Operator monitors multiple robots, triggers takeovers, and resets scenes. VLAC-CUT automatically curates rollout trajectories by segmenting progress, idle, failure-inducing, and recovery portions. Validated on four real-world manipulation tasks, iterative post-training reaches 80%–95% success and 1.7×–4.2× higher throughput; under the same budget, VLAC-CUT guided reuse outperforms HITL-only training.","arXiv :2607 .09776v 1 [ cs .RO] 8 Jul 2026  \nMAXIMIZING HUMAN EFFICIENCY IN LARGE-SCALE ROBOT POST-TRAINING VIA VLAC-CUT GUIDED PIPELINE  \nShaopeng Zhai*†, Qi Zhang* , Tianyi Zhang* , Haoran Zhang* , Fuxian Huang*  \nZijun Xu, Zhanhui Lin  \nContinuous Learning Team, Shanghai AI Lab  \n*Equal contribution. †Project Lead.  \nABSTRACT  \nWhen adapting Vision Language Action (VLA) models to downstream tasks, multiple rounds of post training are required because a single round of data cannot resolve all issues, making continuous iterations necessary to progressively address the weaknesses exposed in previous rounds. In this report, we aim to maximize human efficiency during post-training, defined as the policy improvement and task throughput achieved per unit of human labor and time.  \nWe propose a human-efficient post-training pipeline that enables a small number of human operators to supervise multiple robots. The pipeline is built around a specialized division of labor: a trained Teleoperator focuses on high-value remote interventions and recovery demonstrations, while a Floor Operator monitors multiple robots, triggers takeovers, and performs physical resets. This role specialization reduces task switching, lowers operator training costs, and allows limited human labor to supervise more robot interaction across a larger fleet. To improve data utilization efficiency, we introduce VLAC-CUT as an automatic rollout curation tool. It segments autonomous robot trajectories into progress-making, idle, failure-inducing, and recovery portions, preserving useful segments while filtering harmful or uninformative ones. The curated rollout data are combined with Human-in-the-Loop data for the next post-training round. We validate the proposed pipeline on four real-world manipulation tasks. Across iterative posttraining rounds, the final policies achieve 80%–95% success rates and improve task throughput by 1.7×–4.2× over the base model. Under the same humanintervention budget, VLAC-CUT guided rollout reuse outperforms HITL-only training in both success rate and throughput.  \n1 INTRODUCTION  \nAfter pre-training, generalist policies like Vision Language Action (VLA) models require post training to achieve reliable performance on downstream tasks. This phase typically begins with collecting task specific data to fine tune the VLA. However, because human operators cannot foresee all potential edge cases during the initial data collection, the models fine tuned on this preliminary data inevitably exhibit flaws. Consequently, post training inherently involves multiple iterations. In each round, the policy is evaluated, and its specific failure modes are explicitly recorded to guide targeted data collection in the subsequent phase. This iterative process is crucial for progressively mitigating weaknesses and improving the actual task success rate. However, each iteration also requires humans to evaluate failures, intervene in difficult states, reset physical scenes, and collect or curate new training data. As these costs accumulate across repeated post-training rounds, human labor becomes a central bottleneck for scaling real-world VLA adaptation. Therefore, the key question is not only how to improve the policy, but how much policy improvement and task throughput can be obtained per unit of human labor and time.  \nCurrently, real world post training paradigms generally fall into three categories. The first is real world Reinforcement Learning (RL), for example,[] . This encompasses both fully online methods,  \nwhere the boundaries between training iterations are entirely blurred, and near online methods with clearly defined iterations, such as the approach used in pi0.6 . These approaches require the model to autonomously collect data in the real world, generating a mixture of successful and failed rollouts. Unlike traditional supervised learning that relies purely on positive samples, RL methods can effectively extract useful training signals from th","cbCaieaj2WgUqGdr","https://ap.wps.com/l/cbCaieaj2WgUqGdr","pdf",8696106,2,1,29,"English","en",105,"# Abstract\n# Introduction\n## Problem of multi-round post-training costs\n## Existing paradigms for real-world adaptation\n## Goal: maximizing human efficiency\n# Proposed human-efficient pipeline\n## Role specialization: Teleoperator and Floor Operator\n## VLAC-CUT guided rollout curation\n# Experimental validation and results","[{\"question\":\"Why is multi-round post-training necessary for VLA models in real-world tasks?\",\"answer\":\"Because initial task-specific data collected by humans cannot foresee all edge cases, the fine-tuned model retains flaws. Post-training therefore iterates by evaluating failures, recording failure modes, and collecting targeted data in subsequent rounds.\"},{\"question\":\"What does the report mean by human efficiency in robot post-training?\",\"answer\":\"Human efficiency is defined as the policy improvement and task throughput achieved per unit of human labor and time, measured mainly via success rate and system throughput.\"},{\"question\":\"How do Teleoperator and Floor Operator roles improve supervision efficiency across many robots?\",\"answer\":\"Teleoperators handle high-value remote interventions and recovery demonstrations, while Floor Operators monitor multiple robots, trigger takeovers, and perform physical resets. This specialization reduces context switching, lowers operator training costs, and increases the number of robot interactions a limited team can oversee.\"},{\"question\":\"What is VLAC-CUT and how does it affect data reuse during post-training?\",\"answer\":\"VLAC-CUT automatically curates rollout trajectories by segmenting them into progress-making, idle, failure-inducing, and recovery portions, keeping useful segments while filtering harmful or uninformative ones. Combined with Human-in-the-Loop data, it improves success rate and throughput compared with HITL-only training under the same budget.\"}]",1784201333,73,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":90,"head_meta":92,"extra_data":94,"updated_unix":28},"maximizing-human-efficiency-in-large-scale-robot-post-training-via-vlac-cut-guided-pipeline","",{"@graph":36,"@context":89},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/maximizing-human-efficiency-in-large-scale-robot-post-training-via-vlac-cut-guided-pipeline/85139/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81,85],{"name":72,"@type":73,"acceptedAnswer":74},"Why is multi-round post-training necessary for VLA models in real-world tasks?","Question",{"text":75,"@type":76},"Because initial task-specific data collected by humans cannot foresee all edge cases, the fine-tuned model retains flaws. Post-training therefore iterates by evaluating failures, recording failure modes, and collecting targeted data in subsequent rounds.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What does the report mean by human efficiency in robot post-training?",{"text":80,"@type":76},"Human efficiency is defined as the policy improvement and task throughput achieved per unit of human labor and time, measured mainly via success rate and system throughput.",{"name":82,"@type":73,"acceptedAnswer":83},"How do Teleoperator and Floor Operator roles improve supervision efficiency across many robots?",{"text":84,"@type":76},"Teleoperators handle high-value remote interventions and recovery demonstrations, while Floor Operators monitor multiple robots, trigger takeovers, and perform physical resets. This specialization reduces context switching, lowers operator training costs, and increases the number of robot interactions a limited team can oversee.",{"name":86,"@type":73,"acceptedAnswer":87},"What is VLAC-CUT and how does it affect data reuse during post-training?",{"text":88,"@type":76},"VLAC-CUT automatically curates rollout trajectories by segmenting them into progress-making, idle, failure-inducing, and recovery portions, keeping useful segments while filtering harmful or uninformative ones. Combined with Human-in-the-Loop data, it improves success rate and throughput compared with HITL-only training under the same budget.","https://schema.org",{"og:url":51,"og:type":91,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":93,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":96},[97,101,105,109,114,119,124,127,132,135,139],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":106,"show_sort_weight":107,"slug":108},"Exam",70,"exam",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},5,"Comic",60,"comic",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},6,"Technology",50,"technology",{"id":120,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":122,"slug":123},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":125,"slug":126},30,"research-report",{"id":128,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":130,"slug":131},9,"Religion & Spirituality",20,"religion-spirituality",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":130,"slug":134},"World Cup","world-cup",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":136,"slug":138},10,"Lifestyle","lifestyle",{"id":140,"doc_module":4,"doc_module_name":46,"category_name":141,"show_sort_weight":110,"slug":142},19,"General","general"]