[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86389-en":3,"doc-seo-86389-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86389,8796095360427,"Lucas Martin","https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d",8,"Research & Report","LMEB: Long-horizon Memory Embedding Benchmark","Memory embeddings underpin memory-augmented systems, yet existing text-embedding benchmarks largely evaluate short, passage-based retrieval and do not measure how models handle long-horizon memory retrieval with fragmented, context-dependent, and temporally distant information. LMEB introduces a standardized framework with 22 datasets and 193 zero-shot retrieval tasks spanning episodic, dialogue, semantic, and procedural memory. Experiments cover 15 embedding models, revealing non-monotonic scaling, and orthogonal skill relationships versus MTEB.","arXiv :2603 . 12572v4 [ cs .CL] 13 Jul 2026  \nLMEB: Long-horizon Memory Embedding  \nBenchmark  \nXinping Zhao 1 ,2 , Xinshuo Hu, Jiaxin Xu 1 , Danyu Tang 1 , Xin Zhang 1 , Mengjia Zhou 1 ,  \nYan Zhong3 , Yao Zhou, Zifei Shan, Meishan Zhang 1 , Baotian Hu 1 ,2 ∗ , Min Zhang 1 ,2  \n1Harbin Institute of Technology (Shenzhen); 2 Shenzhen Loop Area Institute (SLAI); 3Peking University  \n[zhaoxinping@stu.hit.edu.cn](zhaoxinping@stu.hit.edu.cn) , [mason.zms@gmail.com](mason.zms@gmail.com) , {hubaotian, [zhangmin2021}@hit.edu.cn](zhangmin2021}@hit.edu.cn)  \n§ GitHub * HuggingFace 􀂌 Leaderboard  \nAbstract  \nMemory embeddings are crucial for memory-augmented systems, such as Open  \nClaw, but their evaluation is underexplored in current text embedding benchmarks,  \nwhich narrowly focus on traditional passage retrieval and fail to assess models’  \nability to handle long-horizon memory retrieval tasks involving fragmented, context  \ndependent, and temporally distant information. To address this gap, we introduce  \nthe Long-horizon Memory Embedding Benchmark (LMEB), a comprehensive  \nframework for evaluating embedding models on complex, long-horizon memory  \nretrieval. LMEB comprises 22 datasets and 193 zero-shot retrieval tasks spanning  \nfour memory types: episodic, dialogue, semantic, and procedural. These memory  \ntypes differ in terms of level of abstraction and temporal dependency, capturing  \ndistinct aspects of memory retrieval that reflect the diverse challenges of the real  \nworld. We evaluate 15 widely used embedding models, ranging from hundreds of  \nmillions to ten billion parameters. The results reveal that (1) LMEB provides a  \nreasonable level of difficulty; (2) Larger models do not always perform better; (3)  \nLMEB and MTEB measure orthogonal capabilities. This suggests that the field  \nhas yet to converge on a universal model capable of excelling across all memory  \nretrieval tasks, and that strong performance on traditional passage retrieval does  \nnot necessarily transfer to long-horizon memory retrieval. LMEB provides a stan  \ndardized and reproducible framework that fills a key gap in memory embedding  \nevaluation and supports future advances in long-term, context-dependent retrieval.  \n1 Introduction  \nMemory embeddings are foundational to a wide range of advanced applications, including agentic systems [OpenClaw Contributors, 2026, Zheng et al., 2025, Fang et al., 2025, Song et al., 2024] and evolving environments [Cao et al., 2025, Ouyang et al., 2025, Chen et al., 2025b] . These memory-augmented systems [Du et al., 2025] require sophisticated mechanisms to store, retrieve, update, and reason over vast amounts of memories, with retrieval being central to their effectiveness [Roediger III and Abel, 2022] . However, despite their growing importance, the evaluation of memory embeddings remains underexplored, especially for long-horizon, context-rich retrieval tasks. Current text embedding benchmarks mainly focus on traditional passage retrieval [Thakur et al., 2021, Muennighoff et al., 2023, Xiao et al., 2024, Enevoldsen et al., 2025], and therefore do not adequately capture the challenges of memory retrieval. Unlike passage retrieval over well-organized documents, long-horizon memory retrieval often requires recalling fragmented, context-dependent, and temporally distant information [Wu et al., 2025, Huet et al., 2025, Kohar and Krishnan, 2025] . As a result, the lack of a comprehensive evaluation protocol for long-horizon memory retrieval leaves a significant gap in understanding how embedding models perform in memory-intensive scenarios.  \n∗Corresponding Author  \nTo bridge this gap, we introduce the Long-horizon Memory Embedding Benchmark (LMEB), a unified framework for evaluating embedding models on complex, long-horizon memory retrieval tasks. Building on the evaluation standards established for text embeddings, such as MTEB [Muennighoff et al., 2023], LMEB extends this evaluation protocol to memory retrieval tasks","cbCaijp7pDPYIOWH","https://ap.wps.com/l/cbCaijp7pDPYIOWH","pdf",4750330,4,1,35,"English","en",105,"# Abstract\n# Introduction\n## Memory embeddings and the evaluation gap\n## LMEB overview and memory types\n### Episodic Memory\n### Dialogue Memory\n### Semantic Memory\n### Procedural Memory","[{\"question\":\"What problem does LMEB address in current text embedding evaluation?\",\"answer\":\"LMEB targets the lack of evaluation for long-horizon, context-rich memory retrieval. Existing benchmarks mainly cover traditional passage retrieval and miss fragmented, temporally distant, and context-dependent memory needs.\"},{\"question\":\"How does LMEB structure its memory retrieval tasks?\",\"answer\":\"LMEB organizes tasks into four memory types: episodic, dialogue, semantic, and procedural memory. These types differ in abstraction level and temporal dependency to reflect real-world challenges.\"},{\"question\":\"What are the main findings from evaluating 15 embedding models on LMEB?\",\"answer\":\"The results show LMEB offers an appropriate difficulty level, larger models do not always outperform, and LMEB and MTEB measure orthogonal capabilities. Strong performance on passage retrieval does not necessarily transfer to long-horizon memory retrieval.\"}]",1784211447,88,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"lmeb-long-horizon-memory-embedding-benchmark","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/lmeb-long-horizon-memory-embedding-benchmark/86389/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does LMEB address in current text embedding evaluation?","Question",{"text":75,"@type":76},"LMEB targets the lack of evaluation for long-horizon, context-rich memory retrieval. Existing benchmarks mainly cover traditional passage retrieval and miss fragmented, temporally distant, and context-dependent memory needs.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does LMEB structure its memory retrieval tasks?",{"text":80,"@type":76},"LMEB organizes tasks into four memory types: episodic, dialogue, semantic, and procedural memory. These types differ in abstraction level and temporal dependency to reflect real-world challenges.",{"name":82,"@type":73,"acceptedAnswer":83},"What are the main findings from evaluating 15 embedding models on LMEB?",{"text":84,"@type":76},"The results show LMEB offers an appropriate difficulty level, larger models do not always outperform, and LMEB and MTEB measure orthogonal capabilities. Strong performance on passage retrieval does not necessarily transfer to long-horizon memory retrieval.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]