[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86213-en":3,"doc-seo-86213-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86213,1374391974564,"Clementine","https://ap-avatar.wpscdn.com/avatar/14000253aa45c000a9e?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779874745381141002",8,"Research & Report","User Preference Induction with LLMs for Offline Top-N Recommendation Evaluation","Offline evaluation remains the standard approach for comparing top-N recommender systems, but it suffers from incomplete relevance information in typical benchmark datasets. Only a small portion of user–item preferences is observed, and unjudged items are often treated as non-relevant under a missing-as-negative assumption. This can bias metrics, unfairly penalize plausible recommendations, and favor popularity or highly exposed items. A proposed LLM-based framework expands relevance judgments by inducing textual user profiles and using an LLM relevance judge for candidate items lacking labels, applied to pooled outputs for improved, more robust ranking assessment.","User Preference Induction with LLMs for Offline Top-N Recommendation Evaluation  \nDavid Otero  \n[david.otero.freijeiro@udc.es](david.otero.freijeiro@udc.es)[ ](david.otero.freijeiro@udc.es)Information Retrieval Lab, CITIC Universidade da Coruña A Coruña, Spain  \nJavier Parapar  \n[javier.parapar@udc.es](javier.parapar@udc.es)[ ](javier.parapar@udc.es)Information Retrieval Lab, CITIC Universidade da Coruña A Coruña, Spain  \narXiv :2607 . 1 1354v 1 [ cs .IR] 13 Jul 2026  \nAbstract  \nOffline evaluation is the standard methodology for comparing topN recommender systems, yet it relies on incomplete relevance information. In most benchmark datasets, only a small subset of user–item preferences is observed, and unjudged items are commonly treated as non-relevant. This missing-as-negative assumption can bias evaluation, penalize plausible recommendations with no recorded feedback, and favour algorithms that concentrate on popular or highly exposed items. We propose an LLM-based framework to expand relevance judgements for offline recommender evaluation. Our approach uses large language models in two complementary roles. First, a preference induction stage summarizes each user’s historical interactions into a textual profile that captures their tastes and interests. Second, conditioned on this profile, an LLM acts as a relevance judge for candidate recommended items that lack observed labels in the original test data. To make this process tractable and evaluation-focused, we apply judgement expansion to a pooled candidate set built from the top-ranked outputs of multiple recommenders. The resulting enriched judgements provide additional relevance evidence for previously unobserved user–item pairs, enabling ranking metrics to be computed on amore complete basis. Experimental results show that this approach is a promising strategy for improving the robustness of offline topN evaluation and mitigating the popularity-sensitive distortions caused by sparse feedback.  \nKeywords  \nRecommender Systems Evaluation, LLM-as-a-judge, Pooling  \nACM Reference Format:  \nDavid Otero and Javier Parapar. 2026. User Preference Induction with LLMs for Offline Top-N Recommendation Evaluation. In . ACM, New York, NY, USA, 10 pages. [https://doi.org/XXXXXXX.XXXXXXX](https://doi.org/XXXXXXX.XXXXXXX)  \n1 Introduction  \nRecommender Systems (RS) are a central component of modern information access, helping users navigate large item spaces by selecting a small set of potentially relevant options [35] . Because  \nPermission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission [and/or a fee. Request permissions from permissions@acm.org](and/or a fee. Request permissions from permissions@acm.org).  \nConference’17, Washington, DC, USA  \n© 2026 Copyright held by the owner/author(s) . Publication rights licensed to ACM. ACM ISBN 978-1-4503-XXXX-X/18/06  \n[https://doi.org/XXXXXXX.XXXXXXX](https://doi.org/XXXXXXX.XXXXXXX)  \ndeploying and testing recommendation algorithms with real users is costly and operationally complex, offline evaluation remains the standard first step for model development and comparison. Offline protocols are attractive because they are efficient, reproducible, and allow controlled experimentation without requiring a live system [13] .  \nThe evaluation methodology used in recommender systems has evolved substantially over time. Early work framed recommendation as a rating prediction task, where the objective was to estimate the exact score that a user would assign to an item [13, 14] . Under this ","cbCaiv8XoncnrOsF","https://ap.wps.com/l/cbCaiv8XoncnrOsF","pdf",536864,4,1,10,"English","en",105,"# Introduction\n## Offline evaluation for top-N recommendation\n## Limitations of missing-as-negative relevance assumptions\n## Motivation for LLM-based judgement expansion","[{\"question\":\"Why does offline top-N evaluation for recommender systems become biased?\",\"answer\":\"Because benchmark test sets contain only a small fraction of labeled user–item interactions, unjudged items are commonly treated as non-relevant. The missing-as-negative assumption conflates “no observed feedback” with “no preference,” biasing metrics and comparisons.\"},{\"question\":\"How does the proposed method use LLMs to expand relevance judgments?\",\"answer\":\"It uses LLMs in two roles: a preference induction stage converts each user’s interaction history into a textual profile, and then an LLM relevance judge assigns relevance to candidate items that lack observed labels in the test data.\"},{\"question\":\"What is the purpose of applying judgement expansion to a pooled candidate set?\",\"answer\":\"To keep the process tractable and evaluation-focused, the framework performs judgement expansion on a pooled set formed from top-ranked outputs across multiple recommenders. This yields enriched relevance evidence for previously unobserved user–item pairs.\"}]",1784209502,25,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"user-preference-induction-with-llms-for-offline-top-n-recommendation-evaluation","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/user-preference-induction-with-llms-for-offline-top-n-recommendation-evaluation/86213/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why does offline top-N evaluation for recommender systems become biased?","Question",{"text":75,"@type":76},"Because benchmark test sets contain only a small fraction of labeled user–item interactions, unjudged items are commonly treated as non-relevant. The missing-as-negative assumption conflates “no observed feedback” with “no preference,” biasing metrics and comparisons.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed method use LLMs to expand relevance judgments?",{"text":80,"@type":76},"It uses LLMs in two roles: a preference induction stage converts each user’s interaction history into a textual profile, and then an LLM relevance judge assigns relevance to candidate items that lack observed labels in the test data.",{"name":82,"@type":73,"acceptedAnswer":83},"What is the purpose of applying judgement expansion to a pooled candidate set?",{"text":84,"@type":76},"To keep the process tractable and evaluation-focused, the framework performs judgement expansion on a pooled set formed from top-ranked outputs across multiple recommenders. This yields enriched relevance evidence for previously unobserved user–item pairs.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":22,"slug":133},"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]