[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83560-en":3,"doc-seo-83560-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83560,34359740700684,"Finn","https://ap-avatar.wpscdn.com/avatar/1f400023980c374ae676?_k=1777273430885731487",8,"Research & Report","As It Was Aligning LLM Search Evaluation with Historical User Preferences","Large-scale search systems evolve faster than human quality assurance, especially for long-tail intents and multilingual queries. LLM-as-a-judge offers scalable SERP evaluation, but judgments driven only by semantic similarity or world knowledge can deviate from true user preferences on ambiguous queries. The work proposes a behavior-grounded LLM judge using a Query–Relevance–Impressions (QRI) card, summarizing historical interactions to resolve ambiguity, improve relevance consistency, and better match user and human outcomes.","As It Was: Aligning LLM Search Evaluation with Historical User Preferences  \nAli Vardasbi  \nSpotify Netherlands [aliv@spotify.com](aliv@spotify.com)  \nClaudia Hauff  \nSpotify Netherlands [claudiah@spotify.com](claudiah@spotify.com)  \nGustavo Penha  \nSpotify United States [gustavop@spotify.com](gustavop@spotify.com)  \nHugues Bouchard Spotify  \nSpain [hb@spotify.com](hb@spotify.com)  \nEnrico Palumbo Spotify  \nItaly [enricop@spotify.com](enricop@spotify.com)  \nMounia Lalmas  \nSpotify United Kingdom [mounia@acm.org](mounia@acm.org)  \narXiv :2607 .0 1040v 1 [ cs .IR] 1 Jul 2026  \nAbstract  \nLarge-scale search systems evolve faster than human quality assurance scales, especially for long-tail intents and multilingual queries. LLM-as-a-judge approaches are a scalable alternative for evaluating the relevance of search engine result pages (SERPs), but judgments based solely on semantic similarity or world knowledge can drift from actual user preferences, particularly for ambiguous queries.  \nWe introduce a behavior-grounded LLM judge that augments each SERP item with a lightweight, auditable behavioral prior in the form of a Query–Relevance–Impressions (QRI) card. Each card summarizes how users have historically interacted with similar queries and results, providing compact empirical evidence that the judge can cite to resolve ambiguity and make more consistent relevance judgments, while still relying on semantic reasoning.  \nIn a large-scale music search evaluation at Spotify, using relevance estimates derived from historical user interactions across 6,000 recomposed SERPs, the behavior-grounded judge achieves stronger alignment with user preferences, improving Spearman rank correlation by approximately +5% overall and yielding a +91% relative improvement on disagreement cases. On a multilingual human-judged dataset spanning five languages, grounding further increases correlation with human relevance judgments by +15% . Importantly, when evaluated against outcomes from a live A/B test, the grounded judge shows consistently higher alignment with the observed winning model. While absolute alignment remains moderate, these findings demonstrate that lightweight behavioral grounding can improve the reliability and practical usefulness of LLM-based evaluation in real-world search systems.  \nCCS Concepts  \n• Information systems → Information retrieval.  \nKeywords  \nLLM-as-a-judge, behavioral grounding, user preference alignment  \nThis work is licensed under a Creative Commons Attribution 4 .0 International License.  \nSIGIR ’26, Melbourne, VIC, Australia  \n© 2026 Copyright held by the owner/author(s) .  \nACM ISBN 979-8-4007-2599-9/2026/07  \n[https://doi.org/10.1145/3805712.3808488](https://doi.org/10.1145/3805712.3808488)  \nACM Reference Format:  \nAli Vardasbi, Gustavo Penha, Enrico Palumbo, Claudia Hauff, Hugues Bouchard, and Mounia Lalmas. 2026. As It Was: Aligning LLM Search Evaluation with Historical User Preferences. In Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR’26), July 20–24, 2026, Melbourne, VIC, Australia. ACM, New York, NY, USA, 6 pages. [https://doi.org/10.1145/3805712.3808488](https://doi.org/10.1145/3805712.3808488)  \n1 Introduction  \nSearch systems in production evolve continuously, with frequent updates to ranking models, retrieval stacks, and user experiences. Ensuring that evaluation keeps pace with these changes requires scalable and reliable assessment methods. As a result, LLM-as-ajudge approaches have become attractive as a scalable evaluation layer, enabling relevance assessment and rationales directly from a query, context, and the SERP. This work is grounded in production music search at Spotify, where search surfaces, ranking models, and user expectations evolve rapidly across languages and regions.  \nPlain LLM judges that rely solely on semantic similarity and catalog-centric reasoning may not fully capture how users interpret and engage wit","cbCainmb8aC2v2Xn","https://ap.wps.com/l/cbCainmb8aC2v2Xn","pdf",1053628,3,1,6,"English","en",105,"# Abstract\n# Introduction\n# Related Work","[{\"question\":\"What problem does the paper address in LLM-based search evaluation?\",\"answer\":\"LLM judges can drift from actual user preferences when judgments rely only on semantic similarity or background knowledge, especially for ambiguous and underspecified queries in multilingual settings.\"},{\"question\":\"How does the proposed behavior-grounded judge work?\",\"answer\":\"It attaches a Query–Relevance–Impressions (QRI) card to each SERP item, summarizing historical query-result interaction statistics to provide lightweight empirical grounding alongside semantic reasoning.\"},{\"question\":\"What evidence is reported to show improved evaluation quality?\",\"answer\":\"Experiments in Spotify music search show better alignment with user preferences (including gains in Spearman rank correlation) and stronger agreement with human judgments across a multilingual dataset, with consistent alignment versus live A/B test outcomes.\"}]",1784188832,15,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"as-it-was-aligning-llm-search-evaluation-with-historical-user-preferences","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/as-it-was-aligning-llm-search-evaluation-with-historical-user-preferences/83560/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address in LLM-based search evaluation?","Question",{"text":75,"@type":76},"LLM judges can drift from actual user preferences when judgments rely only on semantic similarity or background knowledge, especially for ambiguous and underspecified queries in multilingual settings.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed behavior-grounded judge work?",{"text":80,"@type":76},"It attaches a Query–Relevance–Impressions (QRI) card to each SERP item, summarizing historical query-result interaction statistics to provide lightweight empirical grounding alongside semantic reasoning.",{"name":82,"@type":73,"acceptedAnswer":83},"What evidence is reported to show improved evaluation quality?",{"text":84,"@type":76},"Experiments in Spotify music search show better alignment with user preferences (including gains in Spearman rank correlation) and stronger agreement with human judgments across a multilingual dataset, with consistent alignment versus live A/B test outcomes.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]