[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-117403-en":3,"doc-seo-117403-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},117403,7971461740909,"Levi","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Video-Language Critic - Transferable Reward Functions for Language-Conditioned Robotics","Natural language offers an intuitive way for people to specify robot tasks, but grounding language to behavior typically demands large volumes of diverse, language-annotated demonstrations for each target robot. The work disentangles what to accomplish from how to accomplish it, leveraging external observation-only data for the former and robot embodiment–specific information for the latter. It introduces Video-Language Critic, a transferable reward model trained with contrastive learning and temporal ranking to score behavior traces from a separate actor, improving sample efficiency across embodiments on Meta-World.","This is an electronic reprint of the original article.  \nThis reprint may differ from the original in pagination and typographic detail.  \nAlakuijala, Minttu; McLean, Reginald; Woungang, Isaac; Farsad, Nariman; Kaski, Samuel; Marttinen, Pekka; Yuan, Kai  \nVideo-Language Critic : Transferable Reward Functions for Language-Conditioned Robotics  \nPublished in:  \nTransactions on Machine Learning Research  \nPublished: 01/01/2025  \nDocument Version  \nPublisher's PDF, also known as Version of record  \nPublished under the following license:  \nCC BY  \nPlease cite the original version:  \nAlakuijala, M. , McLean, R. , Woungang, I. , Farsad, N. , Kaski, S. , Marttinen, P. , & Yuan, K. (2025) . VideoLanguage Critic : Transferable Reward Functions for Language-Conditioned Robotics. Transactions on Machine  \nLearning Research, 2025, 1-22 . [https://openreview.net/forum?id=jJOVpnNrEp](https://openreview.net/forum?id=jJOVpnNrEp)  \nThis material is protected by copyright and other intellectual property rights, and duplication or sale of all or part of any of the repository collections is not permitted, except that material may be duplicated by you foryour research use or educational purposes in electronic or print form. You must obtain permission for anyother use. Electronic or print copies may not be offered, whether for sale or otherwise to anyone who is not an authorised user.  \nVideo-Language Critic: Transferable Reward Functions for Language-Conditioned Robotics  \nMinttu Alakuijala 1 Reginald McLean2 Isaac Woungang2 Nariman Farsad2 Samuel Kaski 1,3 Pekka Marttinen 1 Kai Yuan4  \n1 Department of Computer Science, Aalto University  \n2 Department of Computer Science, Toronto Metropolitan University  \n3 Department of Computer Science, University of Manchester  \n4 Intel Corporation  \n[mint tu. alakuijala@aalto.fi](mint tu. alakuijala@aalto.fi)  \nReviewed on OpenReview: [https: // openreview. net/ forum? id= jJOVpnNrEp](https: // openreview. net/ forum? id= jJOVpnNrEp)  \nAbstract  \nNatural language is often the easiest and most convenient modality for humans to specify tasks for robots. However, learning to ground language to behavior typically requires impractical amounts of diverse, language-annotated demonstrations collected on each target robot. In this work, we aim to separate the problem of what to accomplish from how to accomplish it, asthe former can benefit from substantial amounts of external observation-only data, and only the latter depends on a specific robot embodiment. To this end, we propose Video-Language Critic, a reward model that can be trained on readily available cross-embodiment data using contrastive learning and a temporal ranking objective, and use it to score behavior traces from a separate actor. When trained on Open X-Embodiment data, our reward model enables 2x more sample-efficient policy training on Meta-World tasks than a sparse reward only, despite a significant domain gap. Using in-domain data but in a challenging task generalization setting on Meta-World, we further demonstrate more sample-efficient training than is possible with prior language-conditioned reward models that are either trained with binary classification, use static images, or do not leverage the temporal information present in video data.1  \n1 Introduction  \nAdvances in natural language processing and vision-language representations have enabled a significant increase in the scalability and generalization abilities of learned control policies for robotics. Methods involving large architectures, such as Transformers (Vaswani et al., 2017), and internet-scale pretraining have transferred well to both high-level (Liang et al., 2022; Vemprala et al., 2023) and low-level (Brohan et al., 2022; Lynch et al., 2022; Shridhar et al., 2022) robotic control. Natural language has many desirable features as a modality for specifying tasks. Unlike structured, hand-designed task sets, natural language is unrestricted and open-domain. Moreover, prompts can be s","cbCailW76TilkBEi","https://ap.wps.com/l/cbCailW76TilkBEi","pdf",5074286,1,23,"English","en",105,"# Introduction\n## Problem Motivation and Background\n## Proposed Approach Overview\n## Reward Model Training and Objectives","[{\"question\":\"Why does grounding natural language to robot behavior require so much data in existing approaches?\",\"answer\":\"Learning usually depends on impractical amounts of diverse, language-annotated demonstrations collected for each target robot, making training data requirements high.\"},{\"question\":\"What does the Video-Language Critic method aim to separate?\",\"answer\":\"It separates “what to accomplish” from “how to accomplish,” so external observation-only data can inform task goals, while embodiment-specific details affect execution.\"},{\"question\":\"How is Video-Language Critic trained and used?\",\"answer\":\"The reward model is trained using readily available cross-embodiment data with contrastive learning and a temporal ranking objective, then scores behavior traces produced by a separate actor.\"}]","Video-Language Critic - Transferable Reward Functions for Language-Conditioned Robotics | PDF",1785675681,58,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"video-language-critic-transferable-reward-functions-for-language-conditioned-robotics","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/video-language-critic-transferable-reward-functions-for-language-conditioned-robotics/117403/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why does grounding natural language to robot behavior require so much data in existing approaches?","Question",{"text":75,"@type":76},"Learning usually depends on impractical amounts of diverse, language-annotated demonstrations collected for each target robot, making training data requirements high.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What does the Video-Language Critic method aim to separate?",{"text":80,"@type":76},"It separates “what to accomplish” from “how to accomplish,” so external observation-only data can inform task goals, while embodiment-specific details affect execution.",{"name":82,"@type":73,"acceptedAnswer":83},"How is Video-Language Critic trained and used?",{"text":84,"@type":76},"The reward model is trained using readily available cross-embodiment data with contrastive learning and a temporal ranking objective, then scores behavior traces produced by a separate actor.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]