[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83567-en":3,"doc-seo-83567-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83567,34359740700684,"Finn","https://ap-avatar.wpscdn.com/avatar/1f400023980c374ae676?_k=1777273430885731487",8,"Research & Report","MemSyco-Bench：评测代理记忆中的奉承倾向","Memory is a core capability in LLM-based agents, enabling long-term collaboration through extracted, stored, and retrieved user-specific information. Yet retrieved memories can harm reliability when outdated beliefs, preferences, or earlier decisions are applied in new contexts. This risk manifests as memory-induced sycophancy, where agents over-align with historical user claims instead of current evidence or objective reasoning. MemSyco-Bench provides five targeted tasks to test correct rejection, scope handling, conflict resolution, update tracking, and personalization via valid memory.","arXiv :2607 .0 107 1v2 [ cs .IR] 2 Jul 2026  \nMEMSYCO-BENCH: BENCHMARKING SYCOPHANCY IN AGENT MEMORY  \nZhishang Xiang 1 Zerui Chen 1∗, Yunbo Tang 1 , Zhimin Wei 1 , Ruqin Ning2 , Yujie Lin 1 , Qinggang Zhang2†, Jinsong Su 1†  \n1Xiamen University  \n2Jilin University  \n[xiangzhishang@stu.xmu.edu.cn](xiangzhishang@stu.xmu.edu.cn) ; [chenzerui1@stu.xmu.edu.cn](chenzerui1@stu.xmu.edu.cn) ;  \n[qinggangzhang@jlu.edu.cn](qinggangzhang@jlu.edu.cn) ; [jssu@xmu.edu.cn](jssu@xmu.edu.cn) ;  \nABSTRACT  \nMemory has emerged as a cornerstone of modern LLM-based agents, supporting their evolution from single-turn assistants to long-term collaborators. However, memory is not always beneficial: retrieved memories often induce a critical issue of sycophancy, causing agents to over-align with the user at the cost of factual accuracy or objective reasoning. Despite this emerging risk, existing memory benchmarks primarily evaluate whether memories are correctly stored, retrieved, or updated, while overlooking how retrieved memories influence downstream reasoning and decision-making. To bridge this gap, we propose MemSyco-Bench, a comprehensive benchmark for evaluating memory-induced sycophancy in agent systems. MemSyco-Bench measures when memory should influence a decision and how valid memory should be used. Specifically, it covers five tasks that assess whether agents can reject memory as factual evidence, respect its applicable scope, resolve conflicts between memory and objective evidence, track memory updates, and use valid memory for personalization. All related resources are collected for the community at [https://github.com/XMUDeepLIT/MemSyco-Bench](https://github.com/XMUDeepLIT/MemSyco-Bench).  \n§ MemSyco-Bench 􀂌 Leaderboard  \nFigure 1: We introduce MemSyco-Bench, a comprehensive benchmark for evaluating sycophancy in agent systems, where retrieved historical memories improperly influence agent reasoning. MemSycoBench assesses whether agents can appropriately reject, constrain, update, reconcile, or leverage retrieved memories across diverse reasoning scenarios. Through extensive experiments, we show that existing memory systems often increase sycophancy and struggle with appropriate memory use.  \n∗Equal contribution.†Corresponding author.  \n1 INTRODUCTION  \nLLM-based agents are rapidly evolving from single-turn assistants into long-term collaborators that interact with users across tasks and sessions(Wang et al., 2024a; Zhao et al., 2023) . Unlike conventional LLMs, these agents are expected to accumulate experience, maintain user-specific knowledge, and adapt their behavior over prolonged interactions(Zhang et al., 2025b; Hu et al., 2025) . To support these capabilities, long-term memory has become a fundamental component of modern agent systems (Chhikara et al., 2025; Xu et al., 2026b) . In a typical memory pipeline, agents extract information from past interactions, store it in an external memory bank, retrieve relevant memories for a new request, and inject them into the context for response generation (Zhong et al., 2024) . This process allows agents to preserve user-specific information beyond the current context window, improving personalization, task continuity, and interaction consistency (Westhäußer et al., 2025) .  \nHowever, memory is not always beneficial. Once retrieved, memories become part of the reasoning context and participate in the agent’s decision-making. This becomes risky when historical user beliefs, preferences, or previous decisions are outdated, outside the current scope, or contradicted by objective evidence. We refer to this failure as memory-induced sycophancy: the agent relies on historical user memory when it should instead follow current evidence or task requirements, causing the response to favor prior user views over objective reasoning. As illustrated in Figure 1, a neutral factual question may ask,“Can the Great Wall be seen from space?\" If the retrieved memory contains a familiar but incorrect user beli","cbCaigvDHoi9e168","https://ap.wps.com/l/cbCaigvDHoi9e168","pdf",8504330,2,1,43,"English","en",105,"# Abstract\n# Introduction\n## Memory pipeline and agent long-term collaboration\n## Memory-induced sycophancy as a failure mode\n## Limitations of existing memory benchmarks\n# MemSyco-Bench overview\n## Leaderboard and evaluation scenarios","[{\"question\":\"What problem does MemSyco-Bench address in agent memory systems?\",\"answer\":\"MemSyco-Bench targets memory-induced sycophancy, where retrieved historical user memories improperly influence agent reasoning and decision-making, reducing factual accuracy or objective reasoning.\"},{\"question\":\"How does memory-induced sycophancy differ from conventional sycophancy?\",\"answer\":\"It shifts the influence source from the current prompt to retrieved historical memories, expands beyond merely agreeing to misusing memories as evidence or outside scope, and can persist across sessions to repeatedly affect later responses.\"},{\"question\":\"What aspects of memory use does MemSyco-Bench evaluate?\",\"answer\":\"It evaluates when memory should influence a decision and how valid memory should be used through five tasks, including rejecting memory as evidence, respecting applicable scope, resolving conflicts with objective evidence, tracking memory updates, and leveraging valid memory for personalization.\"}]",1784188886,108,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"memsyco-bench-benchmarking-sycophancy-in-agent-memory","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/memsyco-bench-benchmarking-sycophancy-in-agent-memory/83567/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does MemSyco-Bench address in agent memory systems?","Question",{"text":75,"@type":76},"MemSyco-Bench targets memory-induced sycophancy, where retrieved historical user memories improperly influence agent reasoning and decision-making, reducing factual accuracy or objective reasoning.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does memory-induced sycophancy differ from conventional sycophancy?",{"text":80,"@type":76},"It shifts the influence source from the current prompt to retrieved historical memories, expands beyond merely agreeing to misusing memories as evidence or outside scope, and can persist across sessions to repeatedly affect later responses.",{"name":82,"@type":73,"acceptedAnswer":83},"What aspects of memory use does MemSyco-Bench evaluate?",{"text":84,"@type":76},"It evaluates when memory should influence a decision and how valid memory should be used through five tasks, including rejecting memory as evidence, respecting applicable scope, resolving conflicts with objective evidence, tracking memory updates, and leveraging valid memory for personalization.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]