[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83709-en":3,"doc-seo-83709-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83709,4398048949847,"Eliana","https://ap-avatar.wpscdn.com/avatar/400002536579ef2da7f?_k=1778318612642679267",8,"Research & Report","APeB Benchmarking Personalization Ability of Large Language Model Agents","LLM-powered agents face personalization failures when users provide raw, underspecified queries. Personalization requires inferring latent intent, extracting preferences from noisy interaction histories, and choosing among highly competing alternatives. Existing benchmarks seldom measure this end-to-end capability because they typically assume user-refined queries or simplified histories. APeB introduces personalized product search (PPS) with heterogeneous action logs, pairing ambiguous intents with rich histories and user-viewed candidates, then evaluates multi-step agent workflows.","APeB: Benchmarking Personalization Ability of Large Language Model Agents  \nGarry Yang1,* , Zizhe Chen1,* , Xinru Chen2 , Yongqiang Chen1 Jianxiang Wang2 , Deyu Zou1 , Linyi Ding2 , Jialiang Wu2 Yunzhong He2 , Yu Gong2 , James Cheng1,†, Huaixiao Tou2,†  \n1The Chinese University of Hong Kong 2ByteDance  \narXiv :2607 .03 162v 1 [ cs .AI] 3 Jul 2026  \nAbstract  \nLLM-powered agents struggle with personalization when users issue raw, underspecified queries. In this setting, agents must infer latent intent, extract preferences from noisy interaction histories, and select among competing alternatives. Existing benchmarks rarely test this capability, as they often rely on userrefined queries or simplified histories. We introduce personalized product search (PPS), atestbed for agentic personalization under raw queries and diverse histories. We construct Agent Personalized Benchmark (APeB) from action logs, pairing underspecified intents with rich histories and user-viewed candidate items. Evaluating state-of-the-art LLMs with multistep agent workflows, we find that models handle explicit queries well but struggle with early-stage queries requiring intent and preference discovery. Rubric analysis attributes this gap mainly to ineffective history use. A simple history-aware query-refinement pipeline, VQRA, yields consistent gains, highlighting the need for dedicated history-utilization modules in personalized agents.  \n1 Introduction  \nLarge Language Model (LLM)-powered agents extend foundation models with additional modules for memory, planning, and execution, enabling multi-step, goal-directed reasoning over long contexts (Peng et al., 2025 ; Wang et al., 2024 ; Zhang et al., 2025c ; Hu et al., 2025) . As shown in Figure 1, a central challenge for LLM-powered agents is personalization: user objectives are not explicitly specified; preferences are implicit and partially observed through histories; and correctness is usercentric rather than globally optimal (Russell, 2021) . In practice, users implicitly reason over their own past experience, translate vague intent into specific criteria, and discriminate between plausible options  \n*Equal contribution.†Corresponding authors.  \nFigure 1: Personalization challenge for agents: infer intent and preferences from diverse history to make a preference-aligned decision.  \n(Downey et al., 2008 ; Punj and Moore, 2009 ; Liu et al., 2023) . Accordingly, a personalization agent should inherently align its planning and action to one particular user’s intents rather than generate generic responses following a fully specified goal (Zhang et al., 2025c) . Therefore, we ask  \nHow can we benchmark the agentic personalization capability of LLMs?  \nTo our knowledge, existing benchmarks individually address one or two of these axes but never combine noisy heterogeneous histories, vague-torefined search trajectories, and closely competing candidates. Recommendation benchmarks cast personalization as next action prediction from historical signals, emphasizing preference deduction over objective understanding (Zhu et al., 2018 ; Liu et al., 2016) . Rationale prediction benchmarks evaluate localized reasoning, focusing on generating intermediate justifications in isolation rather than reaching a user-aligned outcome (Wang et al., 2025c ; Yang et al., 2025b) . Multi-turn Memory benchmarks model personalization over largely homogeneous conversations, evaluating relevant informa-  \n\n| Category | Persona | User Goal | Candidate Set | Evaluation |\n| --- | --- | --- | --- | --- |\n| Generic Shopping | – | Explicit, multi-constraint | Web-scale / open-ended | Multi-step web shopping |\n| Recommendation | Homogeneous, low-semantic | – | Open-ended | Next-item prediction |\n| Personalized Product Search | Homogeneous, low-semantic | Explicit, user-refined | Randomly sampled | Product finding |\n| Rationale Prediction | Synthetic, surveys | – | Random / open-ended | Localized reasoning |\n| Multi-turn Memory | Homogeneous, synt","cbCaifrIAxD0V3qt","https://ap.wps.com/l/cbCaifrIAxD0V3qt","pdf",35057768,6,1,32,"English","en",105,"# Abstract\n# Introduction\n## Personalization challenge for LLM-powered agents\n## Limitations of existing benchmarks\n## APeB and PPS task design","[{\"question\":\"Why is personalization difficult for LLM-powered agents under raw user queries?\",\"answer\":\"Users often provide vague, underspecified queries, so agents must infer latent intent and recover preferences from noisy interaction histories before making a user-aligned choice.\"},{\"question\":\"What is the core contribution of APeB?\",\"answer\":\"APeB builds a personalized product search benchmark that combines underspecified intents with diverse, heterogeneous histories and closely competing candidate items for agentic evaluation.\"},{\"question\":\"What do the results indicate about current state-of-the-art LLM agents?\",\"answer\":\"Models perform well on explicit queries but struggle with early-stage queries that require intent and preference discovery, with rubric analysis attributing the gap mainly to ineffective history usage.\"}]",1784189905,81,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"apeb-benchmarking-personalization-ability-of-large-language-model-agents","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/apeb-benchmarking-personalization-ability-of-large-language-model-agents/83709/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why is personalization difficult for LLM-powered agents under raw user queries?","Question",{"text":76,"@type":77},"Users often provide vague, underspecified queries, so agents must infer latent intent and recover preferences from noisy interaction histories before making a user-aligned choice.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What is the core contribution of APeB?",{"text":81,"@type":77},"APeB builds a personalized product search benchmark that combines underspecified intents with diverse, heterogeneous histories and closely competing candidate items for agentic evaluation.",{"name":83,"@type":74,"acceptedAnswer":84},"What do the results indicate about current state-of-the-art LLM agents?",{"text":85,"@type":77},"Models perform well on explicit queries but struggle with early-stage queries that require intent and preference discovery, with rubric analysis attributing the gap mainly to ineffective history usage.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]