[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83499-en":3,"doc-seo-83499-105":30,"detail-sidebar-cat-0-en-105":87},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83499,687197100911,"Himbo","https://ap-avatar.wpscdn.com/avatar/a000239b6f1da00475?x-image-process=image/resize,m_fixed,w_180,h_180&k=1782698725881665579",8,"Research & Report","Personalized Active Preference Alignment (PAPA) for Diffusion Models","Diffusion models excel at representing complex distributions such as images and text, yet personalized recommender scenarios require steering generation toward regions that maximize evolving user preferences. This objective can be cast as reinforcement learning, where a diffusion model is fine-tuned to maximize a preference-based reward. The barrier is the need to learn a parameterized reward model from large-scale preference data, which is often impractical. PAPA optimizes directly with real-time user feedback, improving feedback efficiency. Extensive experiments, ablations, and the improved EPAPA strategy validate effectiveness and deployment readiness.","PAPA: Online Personalized Active Preference Alignment  \nAnindya Sarkar ∗ Washington University in St. Louis  \n[anindya@wustl.edu](anindya@wustl.edu)  \nNasik Muhammad Nafi * Oak Ridge National Laboratory  \n[nafinm@ornl.gov](nafinm@ornl.gov)  \narXiv :2607 .00486v 1 [ cs .LG] 1 Jul 2026  \nIsaac Lyngaas Oak Ridge National Laboratory  \n[lyngaasir@ornl.gov](lyngaasir@ornl.gov)  \nMuralikrishnan Gopalakrishnan Meena Oak Ridge National Laboratory  \n[gopalakrishm@ornl.gov](gopalakrishm@ornl.gov)  \nYevgeniy Vorobeychik Washington University in St. Louis  \n[yvorobeychik@wustl.edu](yvorobeychik@wustl.edu)  \nAbstract  \nDiffusion models are highly effective at modeling complex data distributions, including images and text. However, in applications like personalized recommender systems, the objective often shifts to modeling specific regions of the distribution that maximize user preferences—initially unknown but gradually uncovered through interactive feedback. This can naturally be framed as a reinforcement learning problem, where the goal is to fine-tune a diffusion model to maximize a reward function based on preferences. However, the main challenge lies in learning a parameterized reward model, which typically requires large-scale preference data—something that is often not feasible in practice. In this work, we introduce Personalized Active Preference Alignment (PAPA), a novel method that bypasses the requirement for a parametrized reward model by directly optimizing the diffusion model using real-time user feedback. PAPA enables feedback-efficient preference alignment, drawing inspiration from the variational inference framework. We demonstrate PAPA ’s effectiveness through extensive experiments and ablation studies across diverse class-conditioned and fine-grained alignment tasks. Additionally, based on theoretical insights, we propose an enhanced fine-tuning strategy, referred to as EPAPA, that requires less computational budget and accelerates the finetuning process, further boosting PAPA’s suitability for realworld deployment. Our code is made publicly available at [https://github.com/NasikNafi/papa](https://github.com/NasikNafi/papa).  \n*Equal Contribution. Accepted at ECML PKDD 2026  \n1. Introduction  \nDiffusion models are deep generative models that generate data by reversing a diffusion process, excelling at capturing complex spaces like natural image manifolds. However, in applications like personalized product recommendations, the goal is to steer generation toward items that align with individual user preferences, which are revealed over time through user activity. A similar challenge arisesin other domains as well. For instance, in image generation, diffusion models are trained on vast datasets scraped from the internet, but practical applications often require images with high aesthetic quality. In fact, many other scenarios share this general structure, such as drug discovery, where the goal is to guide generation toward compounds with high bioactivity. This can be framed as a reinforcement learning (RL) problem, where the objective is to fine-tune the diffusion model to maximize a reward that reflects the desired properties of the user’s preferences. However, these methods rely on extensive preference data to learn the reward model, making them ineffective in scenarios like personalized recommendation platforms, where large-scale preference data for each user is unavailable but can be gathered through costly interactive feedback.  \nThe challenge is twofold: Firstly, achieving this objective requires efficient exploration. However, in highdimensional spaces, such as those of natural images, this goes beyond simply discovering new regions. It also necessitates respecting the structural constraints of the problem. For instance, in areas like product recommendation, valid solutions—such as realistic-looking products—are typically confined to a lower-dimensional manifold within a much larger design space. Therefore, an effect","cbCaidls2lmITFzR","https://ap.wps.com/l/cbCaidls2lmITFzR","pdf",10853033,3,1,33,"English","en",105,"# Introduction\n## Problem Setting: Personalized Active Preference Alignment\n## Challenges: Efficient Exploration and Feedback Cost\n## Reinforcement Learning View and Reward-Model Bottleneck\n## Contributions: PAPA and EPAPA","[{\"question\":\"What is the core idea behind Personalized Active Preference Alignment (PAPA)?\",\"answer\":\"PAPA bypasses learning a parameterized reward model, avoiding the need for large-scale preference datasets that are often infeasible. It instead leverages feedback collected online through user interaction.\"},{\"question\":\"Why is efficient exploration difficult in this online fine-tuning setting?\",\"answer\":\"In many applications, ground-truth rewards depend on costly subjective judgments (e.g., human preference in product recommendation). The model must balance exploration and exploitation while minimizing costly reward queries to avoid disengaging users.\"}]",1784188454,83,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":82,"head_meta":84,"extra_data":86,"updated_unix":28},"personalized-active-preference-alignment-papa-for-diffusion-models","",{"@graph":36,"@context":81},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/personalized-active-preference-alignment-papa-for-diffusion-models/83499/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77],{"name":72,"@type":73,"acceptedAnswer":74},"What is the core idea behind Personalized Active Preference Alignment (PAPA)?","Question",{"text":75,"@type":76},"PAPA bypasses learning a parameterized reward model, avoiding the need for large-scale preference datasets that are often infeasible. It instead leverages feedback collected online through user interaction.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why is efficient exploration difficult in this online fine-tuning setting?",{"text":80,"@type":76},"In many applications, ground-truth rewards depend on costly subjective judgments (e.g., human preference in product recommendation). The model must balance exploration and exploitation while minimizing costly reward queries to avoid disengaging users.","https://schema.org",{"og:url":51,"og:type":83,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":85,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":88},[89,93,97,101,106,111,116,119,124,127,131],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":90,"show_sort_weight":91,"slug":92},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Exam",70,"exam",{"id":102,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},5,"Comic",60,"comic",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},6,"Technology",50,"technology",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":117,"slug":118},30,"research-report",{"id":120,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":122,"slug":123},9,"Religion & Spirituality",20,"religion-spirituality",{"id":122,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":122,"slug":126},"World Cup","world-cup",{"id":128,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":128,"slug":130},10,"Lifestyle","lifestyle",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":102,"slug":134},19,"General","general"]