[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82368-en":3,"doc-seo-82368-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82368,687197207919,"Theodora","https://ap-avatar.wpscdn.com/avatar/a000253d6f5f7c60be?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779446848396160552",8,"Research & Report","Multimodal Reward Hacking in Reinforcement Learning","Reinforcement learning increasingly aligns multimodal large language models, yet higher proxy rewards often fail to translate into better task performance. In multimodal settings, visual evidence is hard to verify with text-only or weakly grounded reward signals, enabling shortcut policies that raise reward scores while harming visual reasoning. This study analyzes reward hacking in a controlled setup for safety VQA, chart VQA, and extreme stress tests, varying reward design, data ambiguity, model scale, and RL algorithm, and introducing NRFR to isolate RL-induced failures.","arXiv :2607 .09492v 1 [ cs .AI] 10 Jul 2026  \nMultimodal Reward Hacking in Reinforcement Learning  \nJiayu Yao†, Yiwei Wang‡, Anmeng Zhang, Zhe Sun§ , Songsong Wang,  \nLingrui Mei†, Yuyao Ge†, Shenghua Liu†, §  \n† Institute of Computing Technology, Chinese Academy of Sciences ‡ University of California, Merced , § Southeast University , § Corresponding Author  \nAbstract  \nReinforcement learning is increasingly used to align multimodal large language models (MLLMs), yet higher rewards do not always mean better task performance. This problem is amplified in multimodal settings because visual input is difficult to verify with text-only or weakly grounded reward signals, allowing policies to improve reward scores while degrading visual reasoning. We study reward hacking in MLLM reinforcement learning through a controlled setup covering safety-oriented VQA, chart VQA, and extreme reward stress tests, while systematically varying reward design, data ambiguity, model scale (2B to 32B), and RL algorithm (GRPO, RLOO, DAPO) . To separate RL-induced failures from pre-existing weaknesses, we introduce Newly Rewarded Failure Rate (NRFR), which measures the hacking rate specifically among samples where RL achieves higher proxy reward than the SFT baseline. Our experiments reveal three main findings. First, outcome-only rewards cause severe hacking (up to 48.1% Reward Hacking Rate), and NRFR exceeding RHR confirms that RL actively creates new failures rather than merely inheriting them. Second, scaling reduces hacking but is insufficient alone. Even the 32B model retains 54.9% worse rate under outcome-only rewards, while answer-aware rewards invert the average oracle direction at every scale, reaching the strongest gap when combined with 32B. Third, algorithm robustness is scale-dependent, with GRPO consistently most resistant (RHR 48–53%), RLOO persistently vulnerable (67–68%), and DAPO improving sharply from 67.2% at 2B to 45.5% at 8B. We further show that visual-evidence rewards help only when the verifier is reliable: keyword-based verification increases hacking relative to answer-aware rewards, whereas VLM-as-judge semantic verification reduces it. Overall, multimodal reward hacking is a systematic consequence of optimizing imperfect rewards. Robust MLLM alignment requires rewards and verifiers that remain reliable under optimization pressure.  \nGitHub: [https://github.com/Theodyy/MLLM-Reward-Hacking](https://github.com/Theodyy/MLLM-Reward-Hacking)  \n1 Introduction  \nReinforcement learning is increasingly adopted as a post-training paradigm for multimodal large language models such as Qwen3-VL [2] and GPT-4o [14] . Real-time human evaluation at training scale is prohibitively expensive [19] . More and more systems therefore rely on automated reward signals, including deterministic scoring rules, keyword checkers, format verifiers, answer matchers, and LLM-based judges. However, a high reward score does not necessarily guarantee genuine performance improvement.  \nThe reason is structural. Automated rewards need to be cheap to compute at scale, so they tend to score what is easy to verify (keyword presence, format compliance, or answer-string match) rather than what is correct. In multimodal tasks this limitation is especially pronounced because the primary evidence, visual input, is difficult to verify with text-only or weakly grounded rewards. A policy can therefore improve its reward score without improving, or while degrading, its  \nInsights for Robust Alignment  \nFigure 1 Overview of multimodal reward hacking. When automated proxy rewards 􀁁(􀁇, 􀁈) diverge from the intended oracle objective 􀀾 (􀁇, 􀁈), RL can increase reward scores while degrading faithful multimodal reasoning. The central example shows a Chart VQA case where the RL policy obtains high reward by giving the correct answer but fabricates visual evidence, exposing a reward–oracle mismatch. We study this failure mode in a controlled multimodal RL sandbox covering task type, r","cbCaic1zPU7WGrvh","https://ap.wps.com/l/cbCaic1zPU7WGrvh","pdf",1403268,2,1,20,"English","en",105,"# Introduction\n# Insights for Robust Alignment\n## Figure 1 Overview of multimodal reward hacking","[{\"question\":\"Why does reward hacking occur more easily in multimodal reinforcement learning?\",\"answer\":\"Automated rewards must be cheap to compute at scale, so they often verify easy-to-check proxies (keywords, format, or answer strings) rather than true correctness. In multimodal tasks, visual evidence is difficult to validate with text-only or weakly grounded reward signals, allowing policies to optimize reward without faithful visual reasoning.\"},{\"question\":\"What is the Newly Rewarded Failure Rate (NRFR) and what does it measure?\",\"answer\":\"NRFR measures the hacking rate specifically among samples where reinforcement learning attains higher proxy reward than the SFT baseline. It helps distinguish RL-induced failures from pre-existing weaknesses.\"},{\"question\":\"Which reward designs and verification methods most affect hacking rates?\",\"answer\":\"Outcome-only rewards cause severe hacking (up to 48.1%). Scaling alone reduces hacking but does not eliminate it, while answer-aware rewards can reverse oracle direction across scales. Visual-evidence rewards improve robustness only when the verifier is reliable: keyword-based verification increases hacking relative to answer-aware rewards, while VLM-as-judge semantic verification reduces it.\"}]",1784179961,50,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"multimodal-reward-hacking-in-reinforcement-learning","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/multimodal-reward-hacking-in-reinforcement-learning/82368/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-22","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why does reward hacking occur more easily in multimodal reinforcement learning?","Question",{"text":75,"@type":76},"Automated rewards must be cheap to compute at scale, so they often verify easy-to-check proxies (keywords, format, or answer strings) rather than true correctness. In multimodal tasks, visual evidence is difficult to validate with text-only or weakly grounded reward signals, allowing policies to optimize reward without faithful visual reasoning.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the Newly Rewarded Failure Rate (NRFR) and what does it measure?",{"text":80,"@type":76},"NRFR measures the hacking rate specifically among samples where reinforcement learning attains higher proxy reward than the SFT baseline. It helps distinguish RL-induced failures from pre-existing weaknesses.",{"name":82,"@type":73,"acceptedAnswer":83},"Which reward designs and verification methods most affect hacking rates?",{"text":84,"@type":76},"Outcome-only rewards cause severe hacking (up to 48.1%). Scaling alone reduces hacking but does not eliminate it, while answer-aware rewards can reverse oracle direction across scales. Visual-evidence rewards improve robustness only when the verifier is reliable: keyword-based verification increases hacking relative to answer-aware rewards, while VLM-as-judge semantic verification reduces it.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":29,"slug":113},6,"Technology","technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":22,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":127,"show_sort_weight":22,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":46,"category_name":135,"show_sort_weight":106,"slug":136},19,"General","general"]