[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81692-en":3,"doc-seo-81692-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81692,1649267921044,"Ava Thompson","https://us-avatar.wpscdn.com/avatar/1800007509477c92dfb?_k=1782875107921204101",8,"Research & Report","WARP-RM A Warp-Augmented Relative Progress Reward Model for Data Curation","Scaling imitation learning depends on large datasets, but human teleoperation produces mixed-quality demonstrations with hesitations and recoveries. WARP-RM addresses progress-reward supervision noise and costly subtask labeling by learning dense, signed relative progress fully self-supervised from successful demonstrations. Time-warp augmentations (variable playback speeds and reversals) generate per-frame progress targets, while WARP-BC uses predicted scalar reward estimates to upweight high-advantage action chunks during behavior cloning. Experiments on robotic T-shirt folding and bottle-in-bin show robustness to suboptimal data, higher success and throughput, and released evaluation artifacts.","arXiv :2606 .28320v3 [ cs .RO] 10 Jul 2026  \nWARP-RM: A Warp-Augmented Relative Progress Reward Model for Data Curation  \nJustin Yu1 3 * † Andrew Goldberg1 * Kavish Kondap1 * Karim El-Refai1 * Ethan Ransing1 Qianzhong Chen2 Mac Schwager2 Fred Shentu3 Philipp Wu3 Ken Goldberg1  \n1University of California, Berkeley 2 Stanford University 3XDOF  \n* Equal contribution. †Corresponding author: [yujustin@berkeley.edu](yujustin@berkeley.edu)  \nFigure 1: WARP-RM signed progress measure vˆt on an unseen mixed-quality teleoperated T-shirt-folding demonstration. Large negative magnitudes occur when the right gripper drops the shirt in (b), and near-zero magnitude during stagnation between (g) and (h) . These values are used to filter and weight downstream policy training. Predictions on more examples in Appendix G.  \nAbstract: Scaling imitation learning requires large datasets, yet human teleoperation inevitably produces mixed-quality demonstrations containing hesitationsand recoveries. Prior frame-level progress reward models supervise on absolute temporal progress proxies that suffer from label noise, or require costly human annotations to define subtask boundaries. We present WARP (Warp-Augmented Relative Progress), a novel fully self-supervised algorithm for learning dense, signed relative progress magnitudes directly from successful demonstrations. WARP generates per-frame progress targets via time-warp augmentations of demonstrations (variable playback speeds and reversals) and we train WARP-RM to predict the normalized elapsed time between input frames. Aggregating these predictions across overlapping windows yields a dense frame-level progress signal. We then introduce WARP-BC, which leverages these scalar reward estimates to upweight high-advantage action chunks during behavior cloning, where chunk-level advantage is obtained by aggregating per-frame rewards. We evaluate our approach on a physical bimanual robot system performing a long-horizon deformable object manipulation task: folding T-shirts from a random crumpled start. To evaluate policy robustness against suboptimal data, we construct training datasets of varying quality using episode length as a proxy for teleoperation sub-optimality. As the dataset is widened to admit more inefficiencies, WARP-BC maintains a 19/20 success rate compared to vanilla BC’s collapse to 2/20, improving throughput by up to ∼ 18× . Furthermore, we evaluate a bottle-in-bin placement task in the real-world, as well as in a reproducible simulation of the task, where gains in success, speed, and throughput hold under paired significance tests, and we release all simulation code and evaluation artifacts.  \nProject page: [https://uynitsuj.github.io/warp-rm/](https://uynitsuj.github.io/warp-rm/) .  \nKeywords: Robot Imitation Learning, Data Curation, Reward Models  \n1 Introduction  \nImitation learning has emerged as a powerful framework for long-horizon robot manipulation, enabling complex visuomotor behaviors from human teleoperated demonstrations. Recent advancesin policy modeling and large-scale pretraining [1, 2, 3, 4, 5, 6, 7, 8, 9] have significantly improved expressivity and generalization. However, these methods remain highly sensitive to demonstration quality, particularly in long-horizon manipulation settings, as policies learn to mimic suboptimal pauses and fumbles present in human teleoperation [10, 11, 12] . If trained on, these behaviors can derail policy performance. However, suboptimal sequences often contain valuable recovery behaviors, akin to DAgger data [13, 14, 15, 16, 17] . Some approaches to data curation operate at the trajectory level, discarding entire episodes that fall below a quality threshold [18, 19, 20] . However, this coarse filtering may suffer from two limitations: it discards valuable, high-advantage segments embedded within otherwise suboptimal executions, while simultaneously failing to prune localized hesitations or fumbles within retained demonstrations.  \nTo addr","cbCaimW6GYbZh1RH","https://ap.wps.com/l/cbCaimW6GYbZh1RH","pdf",17069002,2,1,24,"English","en",105,"# Introduction\n# Method\n## WARP-RM\n## WARP-BC\n# Experiments\n## Data curation via episode-quality proxies\n## Real-world and simulation evaluation\n# Contributions","[{\"question\":\"What problem does WARP-RM solve in imitation learning data curation?\",\"answer\":\"WARP-RM learns progress rewards despite mixed-quality teleoperated demonstrations, avoiding label noise from absolute progress proxies and reducing reliance on expensive human annotations for subtask boundaries.\"},{\"question\":\"How does WARP generate training targets for relative progress?\",\"answer\":\"WARP applies time-warp augmentations to successful demonstrations by replaying them with non-uniform velocities, including reversals, then trains a model to predict the normalized elapsed time between input frames.\"},{\"question\":\"How does WARP-BC use WARP-RM outputs during behavior cloning?\",\"answer\":\"WARP-BC leverages scalar reward estimates to upweight high-advantage action chunks, where chunk-level advantage is computed by aggregating per-frame rewards during behavior cloning.\"}]",1784175447,60,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"warp-rm-a-warp-augmented-relative-progress-reward-model-for-data-curation","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/warp-rm-a-warp-augmented-relative-progress-reward-model-for-data-curation/81692/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does WARP-RM solve in imitation learning data curation?","Question",{"text":75,"@type":76},"WARP-RM learns progress rewards despite mixed-quality teleoperated demonstrations, avoiding label noise from absolute progress proxies and reducing reliance on expensive human annotations for subtask boundaries.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does WARP generate training targets for relative progress?",{"text":80,"@type":76},"WARP applies time-warp augmentations to successful demonstrations by replaying them with non-uniform velocities, including reversals, then trains a model to predict the normalized elapsed time between input frames.",{"name":82,"@type":73,"acceptedAnswer":83},"How does WARP-BC use WARP-RM outputs during behavior cloning?",{"text":84,"@type":76},"WARP-BC leverages scalar reward estimates to upweight high-advantage action chunks, where chunk-level advantage is computed by aggregating per-frame rewards during behavior cloning.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,109,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":29,"slug":108},5,"Comic","comic",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]