[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-133247-en":3,"doc-seo-133247-105":31,"detail-sidebar-cat-0-en-105":93},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},133247,962084925290,"Caleb Sterling","https://ap-avatar.wpscdn.com/davatar_085a072bc5b1113ac321206ff7593b45",8,"Research & Report","Seeing What Matters - Visual Preference Policy Optimization for Visual Generation","Reinforcement learning is used for post-training visual generative models, with Group Relative Policy Optimization (GRPO) aligning generators to human preferences. Existing GRPO pipelines rely on a single scalar reward per sample, overlooking spatial and temporal structure and thereby failing to correct localized artifacts or capture fine-grained perceptual cues. Visual Preference Policy Optimization (ViPO) upgrades GRPO by converting scalar feedback into structured, pixel-level advantages. Using a Perceptual Structuring Module, ViPO builds spatially and temporally aware advantage maps via pretrained vision backbones. Experiments on image and video benchmarks show consistent improvements over vanilla GRPO in human-preference alignment and out-of-domain generalization, with architecture-agnostic, lightweight integration into standard GRPO training.","Seeing What Matters: Visual Preference Policy Optimization for Visual  \nGeneration  \nZiqi Ni 1 *, Yuanzhi Liang2∗, Rui Li2 ,3 , Yi Zhou 1†, Haibin Huang2 , Chi Zhang2 , Xuelong Li2†  \n1 Southeast University, 2Institute of Artificial Intelligence (TeleAI), China Telecom  \n3University of Science and Technology of China  \n[zqni@seu.edu.cn](zqni@seu.edu.cn) , [liangyzh18@outlook.com](liangyzh18@outlook.com) , [yizhou.szcn@gmail.com](yizhou.szcn@gmail.com) , xuelong   [li@ieee.org](li@ieee.org)  \narXiv :2511 . 18719v4 [ cs .CV] 15 May 2026  \nAbstract  \nReinforcement learning (RL) has become a powerful tool for post-training visual generative models, with Group Relative Policy Optimization (GRPO) increasingly used to align generators with human preferences. However, existing GRPO pipelines rely on a single scalar reward per sample, treating each image or video as a holistic entity and ignoring the rich spatial and temporal structure of visual content. This coarse supervision hinders the correction of localized artifacts and the modeling of fine-grained perceptual cues. We introduce Visual Preference Policy Optimization (ViPO), a GRPO variant that lifts scalar feedback into structured, pixel-level advantages. ViPO employs a Perceptual Structuring Module that uses pretrained vision backbones to construct spatially and temporally aware advantage maps, redistributing optimization pressure toward perceptually important regions while preserving the stability of standard GRPO. Across both image and video benchmarks, ViPO consistently outperforms vanilla GRPO, improving in-domain alignment with human-preference rewards and enhancing generalization on out-of-domain evaluations. The method is architecture-agnostic, lightweight, and fully compatible with existing GRPO training pipelines, providing a more expressive and informative learning signal for visual generation.  \n1. Introduction  \nReinforcement learning (RL) has recently emerged as an effective framework for aligning visual generative models [1, 13, 22–24, 26, 29, 40] with human preferences [4, 5], enabling scalable supervision beyond paired data. Among RL-based approaches, Group Relative Policy Optimization (GRPO) [6] has attracted attention for its group-wise  \n*Equal contribution. Work done when Ziqi interned at Institute of Artificial Intelligence (TeleAI), China Telecom.  \n†Corresponding authors.  \nFigure 1 . Brief illustration of our work. Existing GRPO for visual generation assigns a single scalar advantage to the entire content, producing coarse feedback that often leads to sub-optimal results. In contrast, our ViPO converts this coarse signal into preference-aware feedback, enabling fine-grained alignment. This allows, for instance, differentiated optimization of the dancing doll and its background, yielding outputs that are more coherent, harmonious, and perceptually pleasing.  \ncomparison-based advantage formulation, which improves optimization stability and sample quality. Recent studies [38, 41] have successfully extended GRPO to diffusion and flow-based generators, confirming its potential for reinforcement-driven alignment in visual generation.  \nHowever, GRPO was originally designed for token-level or sequence-level outputs, such as in language or reasoning tasks. When directly applied to visual data, this formulation assumes that each visual instance, whether a static image ora video, can be represented by a single scalar advantage, ignoring the rich spatial and temporal structure inherent in visual generation. Such simplification makes GRPO less sensitive to regional or semantic variations within visual content, limiting its ability to assign differentiated credit across spatial locations. Consequently, although the framework remains effective in principle, it provides insufficiently structured feedback for complex visual synthesis tasks. Specifically, this coarse feedback directly affects the visual quality and perceptual alignment of generated results. In convention","cbCaipSJiQtF2YLh","https://ap.wps.com/l/cbCaipSJiQtF2YLh","pdf",14767042,5,1,16,"English","en",105,"# Abstract\n# Introduction\n## Motivation: limits of scalar GRPO for visuals\n## Proposed approach: ViPO with perceptual structuring\n## Expected outcome: fine-grained spatial credit assignment","[{\"question\":\"Why does existing GRPO struggle with visual generation?\",\"answer\":\"GRPO typically assigns a single scalar advantage to an entire image or video, ignoring spatial and temporal structure. This produces coarse feedback and weak credit assignment across regions and semantics, limiting perceptual fidelity.\"},{\"question\":\"What is the core idea behind ViPO?\",\"answer\":\"ViPO reformulates GRPO’s advantage representation by lifting scalar rewards into structured, pixel-level advantages. It redistributes supervision according to perceptual relevance across regions rather than treating all pixels equally.\"},{\"question\":\"How does ViPO generate pixel-level advantage maps?\",\"answer\":\"ViPO uses a Perceptual Structuring Module built on pretrained vision backbones to extract spatial and semantic relevance cues. These cues guide advantage assignment without requiring dense annotations.\"}]","Seeing What Matters - Visual Preference Policy Optimization for Visual Generation | PDF",1787216314,40,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":88,"head_meta":90,"extra_data":92,"updated_unix":29},"seeing-what-matters-visual-preference-policy-optimization-for-visual-generation","",{"@graph":37,"@context":87},[38,55,70],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":54},"https://docshare.wps.com/document/seeing-what-matters-visual-preference-policy-optimization-for-visual-generation/133247/",4,{"url":53,"name":13,"@type":56,"author":57,"headline":13,"publisher":59,"fileFormat":62,"inLanguage":24,"description":14,"dateModified":63,"datePublished":64,"encodingFormat":62,"isAccessibleForFree":65,"interactionStatistic":66},"DigitalDocument",{"name":9,"@type":58},"Person",{"url":42,"name":60,"@type":61},"DocShare","Organization","application/pdf","2026-09-01","2026-08-20",true,{"@type":67,"interactionType":68,"userInteractionCount":20},"InteractionCounter",{"@type":69},"ViewAction",{"@type":71,"mainEntity":72},"FAQPage",[73,79,83],{"name":74,"@type":75,"acceptedAnswer":76},"Why does existing GRPO struggle with visual generation?","Question",{"text":77,"@type":78},"GRPO typically assigns a single scalar advantage to an entire image or video, ignoring spatial and temporal structure. This produces coarse feedback and weak credit assignment across regions and semantics, limiting perceptual fidelity.","Answer",{"name":80,"@type":75,"acceptedAnswer":81},"What is the core idea behind ViPO?",{"text":82,"@type":78},"ViPO reformulates GRPO’s advantage representation by lifting scalar rewards into structured, pixel-level advantages. It redistributes supervision according to perceptual relevance across regions rather than treating all pixels equally.",{"name":84,"@type":75,"acceptedAnswer":85},"How does ViPO generate pixel-level advantage maps?",{"text":86,"@type":78},"ViPO uses a Perceptual Structuring Module built on pretrained vision backbones to extract spatial and semantic relevance cues. These cues guide advantage assignment without requiring dense annotations.","https://schema.org",{"og:url":53,"og:type":89,"og:title":13,"og:site_name":60,"og:description":14},"article",{"robots":91,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":94},[95,99,103,107,111,116,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":96,"show_sort_weight":97,"slug":98},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":100,"show_sort_weight":101,"slug":102},"Literature",80,"literature",{"id":54,"doc_module":4,"doc_module_name":47,"category_name":104,"show_sort_weight":105,"slug":106},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":30,"slug":119},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":47,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":47,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":47,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":47,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]