[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82310-en":3,"doc-seo-82310-105":30,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82310,1374391974564,"Clementine","https://ap-avatar.wpscdn.com/avatar/14000253aa45c000a9e?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779874745381141002",8,"Research & Report","Shortcut Trajectory Planning for Efficient Offline Reinforcement Learning","Diffusion-based trajectory planners perform well in offline reinforcement learning, yet their iterative denoising incurs high inference cost. Consistency-based planners lower sampling steps but often require a two-stage teacher–student distillation pipeline, raising training expense and risking instability. Shortcut Trajectory Planning (STP) introduces an offline model-based RL framework that uses shortcut models as efficient trajectory generators. STP trains a conditional shortcut trajectory model in a single stage, enables adjustable one-step or few-step inference via step-size conditioning, and ranks candidates using a feasibility-aware critic correction, achieving strong results on D4RL locomotion, navigation, manipulation, and dexterous control benchmarks.","arXiv :2607 .09336v 1 [ cs .LG] 10 Jul 2026  \nShortcut Trajectory Planning for Efficient Offline Reinforcement Learning  \nGuanquan Wang [guanquan-wang@g. ecc. u-tokyo. ac.jp](guanquan-wang@g. ecc. u-tokyo. ac.jp)  \nDepartment of Information and Communication Engineering The University of Tokyo  \nYoshimasa Tsuruoka [yoshimasa-tsuruoka@g. ecc. u-tokyo. ac.jp](yoshimasa-tsuruoka@g. ecc. u-tokyo. ac.jp)  \nDepartment of Information and Communication Engineering The University of Tokyo  \nAbstract  \nDiffusion-based trajectory planners have shown strong performance in offline reinforcement learning, but their iterative denoising process often incurs high inference cost. Consistencybased planners reduce the number of sampling steps, yet they typically rely on a two-stage teacher–student distillation pipeline that increases training cost and may introduce instability. We propose Shortcut Trajectory Planning (STP), an offline model-based reinforcement learning framework that incorporates shortcut models as efficient trajectory generators. STP trains a conditional shortcut trajectory model in a single stage, supports adjustable one-step and few-step inference through step-size conditioning, and selects candidate plans using a critic augmented with feasibility-aware correction. Across standard D4RL benchmarks, including locomotion, navigation, manipulation, and dexterous control tasks, STP achieves strong performance while simplifying the training pipeline for fast generative planning.  \n1 Introduction  \nDiffusion-based trajectory planners have shown strong performance in offline reinforcement learning by modeling behavior trajectories as generative distributions and using planning-time sampling for decision making (Janner et al., 2022; Ajay et al., 2023) . However, their iterative denoising process often requires many sampling steps, which makes them computationally expensive for real-time or high-frequency control. Recent consistency-based planners, including Consistency Planning (CP) (Wang et al., 2024a) and Consistency Trajectory Planning (CTP) (Wang et al., 2026), address this limitation by replacing multi-step diffusion sampling with consistency-based generation (Song et al., 2023; Ding & Jin, 2024; Kim et al., 2024b) . These methods substantially reduce inference cost while maintaining competitive planning performance.  \nDespite their fast inference, CP and CTP still rely on a two-stage training pipeline. They first train an Elucidated Diffusion Model (EDM) (Karras et al., 2022) as a teacher and then distill it into a consistency model (Song et al., 2023; Luo et al., 2023) . This teacher-student procedure increases the total training cost and introduces an additional source of instability: the final planner depends on both a successfully trained teacher and a successfully distilled student. Such instability is particularly undesirable in offline RL, where learning is already challenging due to multimodal trajectory distributions and distributional shift (Fujimoto et al., 2019; Kumar et al., 2020; Yu et al., 2020; Kidambi et al., 2020) .  \nIn this paper, we propose Shortcut Trajectory Planning (STP), a shortcut-model-based trajectory planner for offline model-based RL. Shortcut models (Frans et al., 2025) learn step-size-conditioned updates and can therefore support one-step, few-step, and multi-step generation within a single network. Unlike consistencybased planners that require teacher-student distillation, shortcut models can be trained directly in a single stage, simplifying the learning pipeline while preserving the ability to perform fast sampling. This makes shortcut models a natural fit for trajectory planning, where different tasks may require different trade-offs between planning speed and trajectory quality.  \nWe evaluate STP on standard offline RL benchmarks (Fu et al., 2020) and compare it with diffusionbased, consistency-based, and CTM-based planners. Our experiments show that STP achieves competitive planning performa","cbCaiqrPdnrymGCI","https://ap.wps.com/l/cbCaiqrPdnrymGCI","pdf",844779,2,1,16,"English","en",105,"# Abstract\n# Introduction\n# Related Work","[{\"question\":\"What is the core design of Shortcut Trajectory Planning (STP)?\",\"answer\":\"STP uses shortcut models as efficient trajectory generators and trains a conditional shortcut trajectory model in a single stage, supporting adjustable one-step or few-step inference through step-size conditioning and candidate selection via a feasibility-aware critic correction.\"}]",1784179530,40,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":28},"shortcut-trajectory-planning-for-efficient-offline-reinforcement-learning","",{"@graph":36,"@context":77},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/shortcut-trajectory-planning-for-efficient-offline-reinforcement-learning/82310/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-22","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"What is the core design of Shortcut Trajectory Planning (STP)?","Question",{"text":75,"@type":76},"STP uses shortcut models as efficient trajectory generators and trains a conditional shortcut trajectory model in a single stage, supporting adjustable one-step or few-step inference through step-size conditioning and candidate selection via a feasibility-aware critic correction.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,111,114,119,122,126],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":46,"category_name":109,"show_sort_weight":29,"slug":110},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":112,"slug":113},30,"research-report",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},9,"Religion & Spirituality",20,"religion-spirituality",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":120,"show_sort_weight":117,"slug":121},"World Cup","world-cup",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":123,"slug":125},10,"Lifestyle","lifestyle",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":98,"slug":129},19,"General","general"]