[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86342-en":3,"doc-seo-86342-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86342,7971461741311,"Ophelia","https://ap-avatar.wpscdn.com/avatar/74000253aff267980c6?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779345379180704826",8,"Research & Report","Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation","This paper introduces SpectraReward, a training-free reward function that repurposes pretrained multimodal large language models (MLLMs) as ready-to-use reward models for image-generation reinforcement learning. Instead of scoring images directly, it recovers how well the original prompt can be read from a generated image via a single image-conditioned, teacher-forced forward pass. The reward is the average image-conditioned prompt token log-likelihood. It also proposes Self-SpectraReward, enabling closed-loop self-improvement without external reward models or knowledge, validated across diffusion models, RL algorithms, reward backbones, and out-of-distribution benchmarks.","arXiv :2607 . 11886v1 [ cs .CV] 13 Jul 2026  \nRead It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation  \nRunhui Huang 1 Qihui Zhang3 Zhe Liu 1 Yu Gao2 Jie Wu2 Hengshuang Zhao 1  \n1 The University of Hong Kong, 2ByteDance Seed, 3Peking University  \nAbstract  \nIn this paper, we propose SpectraReward, a training-free reward function that turns pretrained MLLMs into off-the-shelf reward models for image-generation reinforcement learning. Instead of asking the MLLM to judge a generated image or answer decomposed verification questions, SpectraReward measures how well the original prompt can be recovered from the generated image through a single imageconditioned, teacher-forced forward pass. We use the average image-conditioned prompt log-likelihood as the reward, directly reusing the MLLM’s pretrained image-text alignment ability without preference labels, reward-model fine-tuning.  \nWe further introduce Self-SpectraReward, a special case for unified multimodal models where the policy’s own understanding branch serves as the reward model for its generation branch, forming a closed-loop self-improving framework without external reward models or external knowledge. Extensive experiments validate SpectraReward through a broad image-generation RL study covering two diffusion models, three RL algorithms, nine reward MLLM backbones from four MLLM families spanning 4B to 235B parameters, and five out-of-distribution text-toimage benchmarks. Results show that both SpectraReward and Self-SpectraReward significantly and consistently improve generation performance and outperform prior MLLM-derived reward training methods. Further analysis reveals that larger reward MLLMs are not always better, while Self-SpectraReward can match or surpass much larger external reward models, suggesting that reward-policy alignment is a key factor for effective image-generation RL. Project Page: [https://huangrh99](https://huangrh99) .  \n[github.io/SpectraReward/](github.io/SpectraReward/)  \n1 Introduction  \nImage generation has advanced rapidly in recent years, evolving from specialized text-to-image models [10, 24] to unified multimodal models (UMMs) [9, 41, 7, 46, 8, 16, 34] that integrate visual understanding and generation within a single architecture. Reinforcement learning has emerged in parallel as an effective post-training stage [28, 62, 61, 65], consistently lifting compositional fidelity and instruction-following. The success of an RL recipe, however, rests on two complementary pillars. The optimization algorithm governs the stability of long-horizon training, while the reward model determines the ceiling that the trained policy can ultimately approach. However, designing a practical reward model that remains both efficient and reliable is still challenging.  \nRecent studies have made substantial progress in reward modeling for image generation. One line of work builds reward models from large-scale human preference annotations, using these data to align visual-text representations or vision-language models with human judgments [22, 59, 54, 51, 47] . While effective, these methods depend on expensive annotation, difficult data collection, and costly iteration. Another line of work bypasses preference training by reusing pretrained MLLMs as zeroshot reward sources [23, 26, 17] . Direct scalar or logit-based feedback is training-free, but can be sensitive to judge calibration and scoring noise [23, 26] . Question-decomposition pipelines improve  \nPreprint.  \n| (c) Reward MLLM Comparison\u003Cbr>\u003Cbr>\u003Cbr>(d) Quantitative Comparison TIIFBench\u003Cbr>\u003Cbr> |\n| --- |\n| GenEval |\n\nFigure 1: Overview of SpectraReward. (a) Pretrained MLLMs naturally induce a semantic spectrum that measures how well a generated image aligns with the prompt. SpectraReward aggregates this into a reward for T2I RL. (b) During RL training, SpectraReward steadily increases together with visible improvements in complex scene generation. (c) We study nine reward ","cbCailZ2nhEzPDha","https://ap.wps.com/l/cbCailZ2nhEzPDha","pdf",10654901,7,1,22,"English","en",105,"# Introduction\n## Image generation and reinforcement learning\n## Reward model design challenges\n# SpectraReward\n## Training-free prompt recovery from generated images\n## Semantic spectrum to scalar reward\n# Self-SpectraReward\n## Closed-loop multimodal self-improvement","[{\"question\":\"What is SpectraReward and what problem does it solve?\",\"answer\":\"SpectraReward is a training-free reward function that turns pretrained MLLMs into reward models for text-to-image reinforcement learning, avoiding preference-label training and additional reward-model tuning.\"},{\"question\":\"How does SpectraReward compute the reward from a generated image?\",\"answer\":\"It freezes an MLLM, conditions it on the generated image, and runs a single teacher-forced forward pass on the prompt. The reward is the average image-conditioned prompt token log-likelihood, reflecting a semantic spectrum of how well prompt requirements are recovered from the image.\"},{\"question\":\"What is Self-SpectraReward and how does it differ from SpectraReward?\",\"answer\":\"Self-SpectraReward is a special case for unified multimodal models where the policy’s own understanding branch provides the reward signal to its generation branch, forming a closed-loop framework without external reward models or external knowledge.\"}]",1784210627,55,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"read-it-back-pretrained-mllms-are-zero-shot-reward-models-for-text-to-image-generation","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/read-it-back-pretrained-mllms-are-zero-shot-reward-models-for-text-to-image-generation/86342/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What is SpectraReward and what problem does it solve?","Question",{"text":76,"@type":77},"SpectraReward is a training-free reward function that turns pretrained MLLMs into reward models for text-to-image reinforcement learning, avoiding preference-label training and additional reward-model tuning.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does SpectraReward compute the reward from a generated image?",{"text":81,"@type":77},"It freezes an MLLM, conditions it on the generated image, and runs a single teacher-forced forward pass on the prompt. The reward is the average image-conditioned prompt token log-likelihood, reflecting a semantic spectrum of how well prompt requirements are recovered from the image.",{"name":83,"@type":74,"acceptedAnswer":84},"What is Self-SpectraReward and how does it differ from SpectraReward?",{"text":85,"@type":77},"Self-SpectraReward is a special case for unified multimodal models where the policy’s own understanding branch provides the reward signal to its generation branch, forming a closed-loop framework without external reward models or external knowledge.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]