[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86539-en":3,"doc-seo-86539-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86539,962075114765,"Quinn","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","The Hidden Footprint: Making Storage a First-Class Metric for LLM Agent Evaluation","LLM agent benchmarks assess task completion, reliability, and inference cost, but typically omit the persistent disk residue created by agent runs, including logs, context snapshots, checkpoints, and debug traces. These retained bytes affect deployability on desktops, compliance with data-retention rules, and behavior at fleet scale. The paper introduces AgentFootprint, a serialization-aware, cross-framework storage-footprint benchmark covering total retention, channel composition, duplication, growth exponent, compressibility, and reconstructability.","The Hidden Footprint: Making Storage a First-Class Metric for LLM Agent  \nEvaluation  \nChenglin Yu 1 , Hongquan Gui 1 , Ying Yu2 , Hongxia Yang3 , Ming Li 1,4∗  \n1Department of Industrial and Systems Engineering, The Hong Kong Polytechnic University  \n2 College of Economics and Management, Zhejiang Normal University  \n3Department of Computing, The Hong Kong Polytechnic University  \n4Research Institute for Generative AI, The Hong Kong Polytechnic University  \n[chenglin.yu@polyu.edu.hk](chenglin.yu@polyu.edu.hk), [ming.li@polyu.edu.hk](ming.li@polyu.edu.hk)  \narXiv :2607 . 1 1 149v 1 [ cs .AI] 13 Jul 2026  \nAbstract  \nLLM agent benchmarks measure task completion, reliability, and inference cost, but not the persistent data an agent run leaves on disk—logs, context snapshots, checkpoints, debug traces. These bytes are absent from the leaderboards we survey, yet bear on whether an agent system can be deployed on desktops, under data-retention compliance regimes, or at fleet scale. We introduce AgentFootprint, to our knowledge the first systematic cross-framework benchmark of post-run agent storage footprint. Its serialization-aware metric suite—total retention, channel composition, duplication, growth exponent, compressibility, and a conversation-history reconstructability score—addresses a measurement trap: naïve byte-level measurement understates duplication by an order of magnitude, because database paging and JSON escaping obscure repeated content. The footprint has two determinants—the logical volume the agent’s behavior generates, and the amplification the persistence layer adds when storing it—and a fixed-trace control isolates the second: the same trajectory replayed through each persisting framework’s adapter yields retained sizes differing by 6.7× . Under identical models, tools, and tasks, framework configurations at identical 100% accuracy differ by  \n15.7× in retained bytes; the defaults bundle different recovery and audit capabilities, so the suite reports storage jointly with accuracy and reconstructability. Three full-history configurations grow superlinearly on a repeated-observation stress task, and on a deliberately minimal write task framework residue exceeds the delivered output files by orders of magnitude. Exported trajectories from 108 instance-normalized SWE-bench Verified submissions span three orders of magnitude in perinstance volume with no detectable correlation with resolve rate. A content-addressed store reduces retention 4.8–32.7× while preserving all properties checked by our restoration tests, including every reconstructability score: exact history reconstruction does not require megabytes, and these storage metrics can be reported alongside inference cost.  \n1 Introduction  \nLLM agents—systems that pursue multi-step tasks by interleaving model calls with tool use—are moving from demos into production, and the community evaluates them along an expanding set of axes: task success on GAIA (Mialon et al. 2024) and SWE-bench (Jimenez et al. 2024), reliability via pass^k on τ-bench (Yao et al. 2025), and, following Kapoor  \n∗Corresponding author.  \net al. (2025), inference cost. A recent survey of agent evaluation catalogs more than a dozen such metrics (Mohammadi et al. 2025)—none of them includes the storage footprint: the bytes an agent execution leaves behind on disk. This paper measures that axis.  \nUnlike inference cost, which is consumed when execution ends, retention is a stock that persists and accumulates across runs. This persistence is simultaneously useful and costly: it is the substrate for replay debugging, recovery, and compliance audits, and it enlarges long-term storage, governance, and deletion obligations. The magnitudes involved are easy to underestimate. In a production deployment that we measured directly—a supply-chain quoting system (Supp. App. P)—a single task retained 131 MB, including 79 MB of framework residue, compared with 2.6 MB of delivered artifacts; per-task retentio","cbCaihgH3nhJOLYA","https://ap.wps.com/l/cbCaihgH3nhJOLYA","pdf",1009379,3,1,17,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"What problem does the paper address in existing LLM agent benchmarks?\",\"answer\":\"Most benchmarks focus on task success, reliability, and inference cost, but do not measure the persistent storage footprint—what bytes an agent leaves on disk after a run.\"},{\"question\":\"What is AgentFootprint and what does it measure?\",\"answer\":\"AgentFootprint is a cross-framework benchmark and measurement harness that quantifies post-run storage footprint using serialization-aware metrics such as total retention, duplication, growth exponent, compressibility, and reconstructability.\"},{\"question\":\"Why can naïve byte-level measurement be misleading for storage footprint?\",\"answer\":\"Because serialization details like database paging and JSON escaping can mask duplication, causing byte-level approaches to understate repeated content by large factors.\"}]",1784212489,43,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"the-hidden-footprint-making-storage-a-first-class-metric-for-llm-agent-evaluation","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/the-hidden-footprint-making-storage-a-first-class-metric-for-llm-agent-evaluation/86539/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address in existing LLM agent benchmarks?","Question",{"text":75,"@type":76},"Most benchmarks focus on task success, reliability, and inference cost, but do not measure the persistent storage footprint—what bytes an agent leaves on disk after a run.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is AgentFootprint and what does it measure?",{"text":80,"@type":76},"AgentFootprint is a cross-framework benchmark and measurement harness that quantifies post-run storage footprint using serialization-aware metrics such as total retention, duplication, growth exponent, compressibility, and reconstructability.",{"name":82,"@type":73,"acceptedAnswer":83},"Why can naïve byte-level measurement be misleading for storage footprint?",{"text":84,"@type":76},"Because serialization details like database paging and JSON escaping can mask duplication, causing byte-level approaches to understate repeated content by large factors.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]