[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83688-en":3,"doc-seo-83688-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83688,4810365810221,"Aurora","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Long-Term Optimization for Large-Scale Generative Retrieval with Off-Policy REINFORCE","Generative retrieval supports large-scale recommendation, but commonly relies on supervised next-item prediction that does not directly optimize long-term user satisfaction. The work models recommendation as session-level sequential decision-making and trains autoregressive generative retrievers with off-policy REINFORCE using pre-collected data. Multi-step importance-weight approximation replaces prior one-step correction. It enables offline evaluation via a feedback model for user responses, adapting doubly robust OPE and a test-time scaling method selecting recommendations with maximum predicted long-term return.","Long-Term Optimization for Large-Scale Generative Retrieval  \nwith Off-Policy REINFORCE  \nArtem Matveev  \n[matfu21@yandex.ru](matfu21@yandex.ru)[ ](matfu21@yandex.ru)AI VK Moscow, Russia  \nSergei Makeev  \n[neuralsrg@gmail.com](neuralsrg@gmail.com)[ ](neuralsrg@gmail.com)AI VK Moscow, Russia  \nAleksei Krasilnikov  \n[TheKabeton@yandex.ru](TheKabeton@yandex.ru)[ ](TheKabeton@yandex.ru)AI VK Moscow, Russia  \nVladimir Baikalov  \n[deadinside@itmo.ru](deadinside@itmo.ru)[ ](deadinside@itmo.ru)AI VK Moscow, Russia  \nSergei Liamaev  \n[liamaev.sergei@gmail.com](liamaev.sergei@gmail.com)[ ](liamaev.sergei@gmail.com)AI VK Moscow, Russia  \nKirill Khrylchenko  \n[elightelol@gmail.com](elightelol@gmail.com)[ ](elightelol@gmail.com)HSE University Moscow, Russia  \narXiv :2607 .028 18v 1 [ cs .IR] 2 Jul 2026  \nAbstract  \nGenerative retrieval has become a popular paradigm for large-scale recommendation. However, it is typically trained with supervised next-item prediction objectives that do not directly optimize longterm user satisfaction.  \nIn this work, we formulate recommendation as a session-level sequential decision-making problem and introduce an autoregressive approach for training generative retrievers with off-policy REINFORCE on pre-collected data. Unlike the one-step off-policy correction used in prior work, we propose a multi-step approximation of importance weights enabled by the autoregressive formulation. To support offline evaluation, we train a user feedback model that simulates user responses to generated recommendations. This lets us adapt doubly robust off-policy evaluation for sequential decision-making to recommendation, a setting that has received limited attention. We further introduce a feedback-modelbased test-time scaling procedure that simulates future responsesand selects recommendations with the highest predicted long-term returns.  \nExperiments on the public large-scale Yambda-5B dataset show that our RL agent improves offline estimates of cumulative session reward over next-item and next-positive prediction baselines, while largely preserving retrieval quality. Moreover, allocating more inference-time compute to simulating future responses improves model-based long-term return estimates without updating the policy.  \nCCS Concepts  \n• Information systems → Information retrieval; Recommender systems; • Computing methodologies → Reinforcement learning.  \nKeywords  \nOff-Policy Learning, Off-Policy Evaluation, REINFORCE, Recommender Systems, Reinforcement Learning, Importance Sampling, Candidate Generation  \nPermission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for third-party components of this work must be honored. For all other uses, contact the owner/author(s) .  \nKDD’26, Jeju Island, Republic of Korea  \n© 2026 Copyright held by the owner/author(s) .  \nACM Reference Format:  \nArtem Matveev, Sergei Makeev, Aleksei Krasilnikov, Vladimir Baikalov, Sergei Liamaev, and Kirill Khrylchenko. 2026. Long-Term Optimization for Large-Scale Generative Retrieval with Off-Policy REINFORCE. In Proceedings of 5th Workshop on End-End Customer Journey Optimization at KDD 2026. (KDD’26) . ACM, New York, NY, USA, 10 pages.  \n1 Introduction  \nRecommender systems are widely deployed across industry to help users navigate vast collections of content. Their mission is to retain millions of users by presenting the most relevant items. To achieve this goal, recommender systems are typically optimized with respect to immediate feedback signals, such as clicks [28] or repins [31], which can be intuitively viewed as searching for a local optimum in the space of user interests.  \nReinforcement learning. Reinforcement learning (RL) offers a way to extend this optimization objective to account for long-term rewards","cbCaivbQCL2no5Iz","https://ap.wps.com/l/cbCaivbQCL2no5Iz","pdf",2995926,3,1,10,"English","en",105,"# Introduction\n## Long-term optimization in recommender systems\n## Reinforcement learning for long-term rewards\n## On-policy vs. off-policy methods\n## Challenges of offline evaluation and OPE","[{\"question\":\"Why does standard generative retrieval training underperform for long-term satisfaction?\",\"answer\":\"It is typically trained with supervised next-item objectives, which optimize immediate prediction rather than directly maximizing long-term user satisfaction.\"},{\"question\":\"How does the paper formulate and train generative retrievers for long-term optimization?\",\"answer\":\"It models recommendation as session-level sequential decision-making and trains autoregressive generative retrievers using off-policy REINFORCE on pre-collected data, with multi-step importance-weight approximation.\"},{\"question\":\"How is offline evaluation made possible in this setting?\",\"answer\":\"A user feedback model simulates user responses to generated recommendations, enabling doubly robust off-policy evaluation for sequential decision-making and supporting a test-time scaling procedure for long-term returns.\"}]",1784189744,25,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"long-term-optimization-for-large-scale-generative-retrieval-with-off-policy-reinforce","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/long-term-optimization-for-large-scale-generative-retrieval-with-off-policy-reinforce/83688/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why does standard generative retrieval training underperform for long-term satisfaction?","Question",{"text":75,"@type":76},"It is typically trained with supervised next-item objectives, which optimize immediate prediction rather than directly maximizing long-term user satisfaction.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the paper formulate and train generative retrievers for long-term optimization?",{"text":80,"@type":76},"It models recommendation as session-level sequential decision-making and trains autoregressive generative retrievers using off-policy REINFORCE on pre-collected data, with multi-step importance-weight approximation.",{"name":82,"@type":73,"acceptedAnswer":83},"How is offline evaluation made possible in this setting?",{"text":84,"@type":76},"A user feedback model simulates user responses to generated recommendations, enabling doubly robust off-policy evaluation for sequential decision-making and supporting a test-time scaling procedure for long-term returns.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":22,"slug":133},"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]