[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85613-en":3,"doc-seo-85613-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85613,3848291630094,"Emma Wilson","https://eur-avatar.wpscdn.com/davatar_085a072bc5b1113ac321206ff7593b45",8,"Research & Report","Let It Be Simple One-Step Action Generation for Vision-Language-Action Models","Generating diverse images from sparse text is difficult, while generating compact actions from rich observations is comparatively easier. Under a condition–target perspective, Vision-Language-Action (VLA) aligns more naturally with image-to-text rather than text-to-image. The paper formalizes this via the irreducible velocity loss Rv(t,c) in standard flow matching and validates it with controlled toy experiments and an image-to-text MNIST task. High-noise training improves one-step VLA decoding on LIBERO, reaching 95.6% on LIBERO-Long and staying competitive across related benchmarks and real robot tasks.","Let It Be Simple: One-Step Action Generation for Vision-Language-Action Models  \nYitong Chen 1,2 Shiduo Zhang2,3 Jingjing Gong2,3 Xipeng Qiu2,3  \n1University of Science and Technology of China 2 Shanghai Innovation Institute 3Fudan University  \n[cyt050719@mail.ustc.edu.cn](cyt050719@mail.ustc.edu.cn), [xpqiu@fudan.edu.cn](xpqiu@fudan.edu.cn)  \narXiv :2606 .05737v2 [ cs .CV] 13 Jul 2026  \nAbstract—Generating diverse images from sparse text is hard; generating compact actions from rich observations is easier. From the condition-target view, Vision-Language-Action (VLA) thus aligns with image-to-text, not text-to-image. We formalize this view through the irreducible velocity loss Rv (t, c) of standard flow matching and validate it with a controlled 8-mode toy experiment and image-to-text MNIST task. We then show that high-noise training boosts one-step VLA decoding on standard LIBERO, achieving 95.6% on LIBERO-Long, and remains competitive across LIBERO-Plus, LIBERO-Pro, and real-world robot tasks, while ablations that weaken the condition or expand the horizon predictably erase the one-step gain. These results suggest that whether one-step action generation works in VLA depends not on specialized training, but on the condition-target structure.  \nI. INTRODUCTION  \nDiffusion and flow models usually trade inference time for sample quality. In image generation, one-step or fewstep sampling is hard because a class label or text prompt can still leave a broad, high-dimensional, and multimodal image distribution [35] . This is why strong image generators often rely on extra objectives, teachers, or distillation machinery [41, 45, 15, 16, 5] .  \nCondition-target structure describes the residual generative problem after conditioning: how complex is the target distribution that remains after observing c? For VLA action generation, at each decision point, the policy receives images, language, and proprioceptive state [4, 8, 32], then predicts a short action chunk. If the condition encoder captures the scene and task prompt well, the remaining conditional action distribution can be simpler than that of text-conditional image generation.  \nStandard diffusion and flow matching models [23, 1, 25] use the conditional flow matching loss. At each time t, the velocity prediction loss has an irreducible lower bound; we denote this irreducible velocity loss as Rv (t, c), which over the entire time interval—not the training-time empirical loss—characterizes the intrinsic difficulty of one-step generation.  \nWe make three contributions. First, we formulate the difficulty of one-step flow generation through the lens of condition-target structure and the irreducible velocity-loss profile. Second, we use a simple 8-mode ring toy experiment and an image-to-text MNIST task to explore how different condition strengths shape Rv (t, c) and thereby affect the noiseendpoint prediction difficulty. Third, we test the resulting simple recipe in VLA policies: high-noise training improves one-step decoding on standard LIBERO [24], while weakening the condition or increasing target complexity via longer horizons reduces performance as expected. We further validate this view on LIBERO-Plus, LIBERO-Pro, and real-world robot tasks.  \nWe show that VLA flow matching occupies a special regime: when the condition is informative and the action target is compact, standard flow matching, without distillation or auxiliary objectives, already can yield strong one-step performance.  \nII. RELATED WORK  \nVLA action generation. Robot policies have increasingly adopted vision-language-action models, from autoregressive systems such as RT-1, RT-2, OpenVLA, FAST, and FASTer [6, 52, 18, 34, 26] to continuous diffusion or flow policies such as Diffusion Policy, Octo, π0 , and SimVLA [8, 32, 4, 29] . Autoregressive policies benefit from language-model infrastructure, but action tokenization and decoding order become design choices. Continuous flow policies avoid an action codebook, while","cbCaivJSGJS71VOV","https://ap.wps.com/l/cbCaivJSGJS71VOV","pdf",2687055,3,1,13,"English","en",105,"# Introduction\n# Related Work","[{\"question\":\"What viewpoint does the paper use to explain one-step action generation in VLA models?\",\"answer\":\"It uses a condition–target structure perspective, describing the remaining target complexity after conditioning on observations and language so that the conditional action distribution can be simpler than in text-conditional image generation.\"},{\"question\":\"How is the intrinsic difficulty of one-step generation characterized?\",\"answer\":\"The paper defines an irreducible velocity loss Rv(t,c) for standard flow matching, and uses its profile over the full time interval (not only empirical training loss) to reflect intrinsic difficulty.\"},{\"question\":\"Which training strategy improves one-step decoding, and where is it tested?\",\"answer\":\"High-noise training boosts one-step VLA decoding, validated on standard LIBERO (95.6% on LIBERO-Long) and further checked on LIBERO-Plus, LIBERO-Pro, and real-world robot tasks, with ablations showing predictable loss of the one-step gain when conditions weaken or horizons expand.\"}]",1784204927,33,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"let-it-be-simple-one-step-action-generation-for-vision-language-action-models","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/let-it-be-simple-one-step-action-generation-for-vision-language-action-models/85613/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What viewpoint does the paper use to explain one-step action generation in VLA models?","Question",{"text":75,"@type":76},"It uses a condition–target structure perspective, describing the remaining target complexity after conditioning on observations and language so that the conditional action distribution can be simpler than in text-conditional image generation.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is the intrinsic difficulty of one-step generation characterized?",{"text":80,"@type":76},"The paper defines an irreducible velocity loss Rv(t,c) for standard flow matching, and uses its profile over the full time interval (not only empirical training loss) to reflect intrinsic difficulty.",{"name":82,"@type":73,"acceptedAnswer":83},"Which training strategy improves one-step decoding, and where is it tested?",{"text":84,"@type":76},"High-noise training boosts one-step VLA decoding, validated on standard LIBERO (95.6% on LIBERO-Long) and further checked on LIBERO-Plus, LIBERO-Pro, and real-world robot tasks, with ablations showing predictable loss of the one-step gain when conditions weaken or horizons expand.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]