[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83494-en":3,"doc-seo-83494-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},83494,687197100911,"Himbo","https://ap-avatar.wpscdn.com/avatar/a000239b6f1da00475?x-image-process=image/resize,m_fixed,w_180,h_180&k=1782698725881665579",8,"Research & Report","VLM-AR3L：用于强化学习绝对与相对奖励的视觉-语言模型","Designing effective reward functions remains a major challenge in reinforcement learning, especially in open-ended and visually complex settings where goals are abstract and hard to quantify. VLM-AR3L introduces a vision-language framework that interprets visual observations in the context of natural-language goals and learns both absolute and relative rewards from VLM-generated preference labels. The absolute model evaluates individual states, while the relative model infers progress or regression by comparing consecutive observations, combining stability with comparative robustness.","VLM-AR3L: Vision-Language Models for Absolute and Relative Rewards in  \nReinforcement Learning  \nKuan-Chen Chen 1 , Winston Chen 1 , Wei-Fang Sun2 and Min-Chun Hu 1  \n1Department of Computer Science, National Tsing Hua University  \n2NVIDIA AI Technology Center (NVAITC)  \narXiv :2607 .00483v2 [ cs .RO] 2 Jul 2026  \nAbstract  \nDesigning effective reward functions remains a major challenge in reinforcement learning (RL), particularly in open-ended environments where task goals are abstract and difficult to quantify. In this work, we present VLM-AR3L, a framework that leverages Vision-Language Models (VLMs) to provide both absolute and relative rewards for RL.  \nVLM-AR3L interprets an agent’s visual observations in the context of a natural language task goal, and learns both absolute and relative rewards from VLM-generated preference labels. The absolute reward model predicts scalar evaluations for individual states, while the relative reward model compares consecutive observations to infer progress or regression toward the task goal. Their integration combines the stability of state-based evaluation with the robustness of comparative supervision. We evaluate VLM-AR3L across benchmarks spanning classic control, manipulation, and open-world embodied tasks, with a particular focus on Minecraft given its visual complexity and long-horizon decision-making requirements. Experimental results show that VLM-AR3L consistently outperforms prior VLM-based reward learning methods. Videos and code are available on the project website: [https://vlm-ar3l.github.io/](https://vlm-ar3l.github.io/) .  \n1 Introduction  \nDesigning effective reward functions remains a major challenge in reinforcement learning (RL) [Laud, 2004; Leike et al., 2018; OpenAI, 2019; Gupta et al., 2022] . This challenge is especially pronounced in open-ended or visually complex environments, where task goals are often abstract and difficult to specify precisely. For example, in open-world environments such as Minecraft, agents are often assigned high-level goals (e.g.,“build a shelter”) that require long-horizon planning and semantic understanding, while intermediate states provide little explicit reward signal. As a result, hand-crafting reward functions in such settings is typically impractical, due to ambiguous objectives, sparse feedback, and the absence of privileged state information. To address this bottleneck, recent approaches have explored leveraging large pre-trained  \nmodels as sources of reward, enabling agents to learn from high-level task descriptions or multimodal alignment signals.  \nLarge language models (LLMs) have been shown to provide structured code or textual feedback for RL agents in text-based or programmatic settings. However, these methods often rely on privileged state access or handcrafted programmatic interfaces, limiting their applicability in real-world visual domains. When working with visual observations, contrastive vision-language models (cVLMs) such as CLIP are widely used to compute similarity scores between observations and goals for reward shaping. While simple and scalable, these methods typically require task-specific fine-tuning or retraining to be effective across different domains.  \nTo move beyond direct similarity-based alignment, preference-based reward modeling has emerged as a promising alternative. These methods learn reward functions by comparing observation pairs and inferring preferences that are typically obtained from human feedback or large pretrained models. This formulation enables finer-grained feedback and greater robustness in visually diverse or ambiguous environments. For instance, RL-VLM-F [Wang et al., 2024] leverages generative vision-language models (VLMs) to generate pairwise preferences, which are then used to train a reward model that assigns scalar values to individual states. We call this approach absolute reward. However, we observe that absolute reward signals can be inconsistent across training steps ","cbCaikzlMI4AQuwY","https://ap.wps.com/l/cbCaikzlMI4AQuwY","pdf",4816590,1,16,"English","en",105,"# Abstract\n# Introduction\n## Problem: Reward function design in RL\n## Prior approaches and limitations\n## Proposed framework: VLM-AR3L\n## Contributions: dual-reward design and relative reward benefits","[{\"question\":\"VLM-AR3L解决强化学习中什么核心难题？\",\"answer\":\"它解决奖励函数难以设计的问题，尤其是在开放式、视觉复杂环境中，任务目标抽象且难以量化时更为突出。\"},{\"question\":\"VLM-AR3L如何得到绝对奖励与相对奖励？\",\"answer\":\"VLM-AR3L在自然语言任务目标语境下解析智能体的视觉观测，并利用VLM生成的偏好标签学习奖励。绝对奖励评估单个状态的标量值，相对奖励通过比较连续观测来推断朝目标的进展或退化。\"},{\"question\":\"为什么相对奖励能提升训练稳定性与学习鲁棒性？\",\"answer\":\"文中指出，绝对奖励在训练过程中可能因不断暴露新状态而出现不一致，导致长时域任务难优化；相对奖励通过比较监督提供更好的时间一致性，并能缓解全局状态顺序不明确（如循环结构）带来的问题。\"}]",1784188425,40,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"vlm-ar3l-vision-language-models-for-absolute-and-relative-rewards-in-reinforcement-learning","",{"@graph":35,"@context":84},[36,53,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/vlm-ar3l-vision-language-models-for-absolute-and-relative-rewards-in-reinforcement-learning/83494/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":4},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"VLM-AR3L解决强化学习中什么核心难题？","Question",{"text":74,"@type":75},"它解决奖励函数难以设计的问题，尤其是在开放式、视觉复杂环境中，任务目标抽象且难以量化时更为突出。","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"VLM-AR3L如何得到绝对奖励与相对奖励？",{"text":79,"@type":75},"VLM-AR3L在自然语言任务目标语境下解析智能体的视觉观测，并利用VLM生成的偏好标签学习奖励。绝对奖励评估单个状态的标量值，相对奖励通过比较连续观测来推断朝目标的进展或退化。",{"name":81,"@type":72,"acceptedAnswer":82},"为什么相对奖励能提升训练稳定性与学习鲁棒性？",{"text":83,"@type":75},"文中指出，绝对奖励在训练过程中可能因不断暴露新状态而出现不一致，导致长时域任务难优化；相对奖励通过比较监督提供更好的时间一致性，并能缓解全局状态顺序不明确（如循环结构）带来的问题。","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,118,121,126,129,133],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":28,"slug":117},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":119,"slug":120},30,"research-report",{"id":122,"doc_module":4,"doc_module_name":45,"category_name":123,"show_sort_weight":124,"slug":125},9,"Religion & Spirituality",20,"religion-spirituality",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":124,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]