[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82236-en":3,"doc-seo-82236-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82236,962075114765,"Quinn","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","SherAgent Scaling Attack Investigation in the Wild via LLM-Empowered Iterative Query-Filter Backtracking","Provenance-based attack investigation enables automation by standardizing data and query logic, but real-world deployments face critical obstacles from dependency explosions and fragmented causal chains. To address these issues, the work partners with a large Internet corporation’s SOC handling tens of thousands of alerts daily, analyzing where existing LLM-based workflows fail and why. It introduces SherAgent, an LLM-powered iterative query-filter backtracking system over provenance graphs, dynamically calibrating queries, filtering results, and selecting nodes to mitigate failures.","SherAgent: Scaling Attack Investigation in the Wild via LLM-Empowered Iterative Query-Filter Backtracking  \nZhenyuan Li  \nZhejiang University China  \nXiangmin Shen  \nHofstra University USA  \nZhengkai Wang  \nZhejiang University China  \nRuixiao Lin  \nZhejiang University China  \nLing Jiang  \nTencent Security Keen Lab China  \nSen Nie  \nTencent Security Keen Lab China  \narXiv :2607 .09 176v 1 [ cs .CR] 10 Jul 2026  \nShi Wu  \nTencent Security Keen Lab China  \nAbstract  \nProvenance-based attack investigation enables viable automation by standardizing data and query logic; however, it is critically hindered in practice by dependency explosions and fragmented causal chains in the wild. Towards designing a robust and automated investigation tool, we collaborated with the SOC of a major Internet corporation serving billions of users. By engaging in real-world incident response, we are able to evaluate and refine their existing LLM-based investigation workflows, which processes tens of thousands of raw alerts daily, leaving thousands for manual triage, to find out the root causes of investigation failures and major challenges in their existing tools.  \nMotivated by these findings, we propose SherAgent, an LLMempowered automated investigation system. Operating on an iterative “query-filter” backtracking paradigm over provenance graphs, SherAgent leverages the semantic reasoning capabilities of LLMs to process unstructured data, such as investigation context and threat intelligence. To overcome fragmented causal chains caused by missing events, the system dynamically calibrates query conditions to broaden the search scope. Concurrently, it performs precision result filtering and strategic nodes selection for subsequent exploration, thereby mitigating dependency explosions. Extensive evaluations in the wild demonstrate that SherAgent improves the end-to-end investigation success rate by 31.1% and 63.7% compared to both legacy enterprise baselines and SOTA approaches, respectively. Furthermore, it operates with remarkable efficiency, incurring under $0.10 in API costs and requiring less than 4 minutes per investigation. Finally, our user study confirms that SherAgent provides accurate and clear insights, significantly reducing the analytical overhead for security experts.  \nPermission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission [and/or a fee. Request permissions from permissions@acm.org](and/or a fee. Request permissions from permissions@acm.org).  \nConference acronym ’XX, Woodstock, NY  \n© 2018 Copyright held by the owner/author(s) . Publication rights licensed to ACM. ACM ISBN 978-1-4503-XXXX-X/2018/06  \n[https://doi.org/XXXXXXX.XXXXXXX](https://doi.org/XXXXXXX.XXXXXXX)  \nShouling Ji  \nZhejiang University  \nChina  \nCCS Concepts  \n• Security and privacy → Intrusion detection systems; Operating systems security; • Computing methodologies → Natural language processing; • General and reference → Empirical studies.  \nKeywords  \nAttack Investigation, Provenance Analysis, Dependence Explosion, Broken Chains, Large Language Model, Empirical Study  \nACM Reference Format:  \nZhenyuan Li, Zhengkai Wang, Ling Jiang, Xiangmin Shen, Ruixiao Lin, Sen Nie, Shi Wu, and Shouling Ji. 2018. SherAgent: Scaling Attack Investigation in the Wild via LLM-Empowered Iterative Query-Filter Backtracking. In Proceedings of Make sure to enter the correct conference title from your rights confirmation email (Conference acronym ’XX). ACM, New York, NY, USA, 17 pages. [https://doi.org/XXXXXXX.XXXXXXX](https://doi","cbCaijCIUwL6mNaj","https://ap.wps.com/l/cbCaijCIUwL6mNaj","pdf",1812293,1,17,"English","en",105,"# Abstract\n# Keywords\n# CCS Concepts\n# 1 Introduction\n## Attack investigation challenges and alert fatigue\n## Provenance analysis as a solution\n## Existing approaches: backward tracking vs query-based hunting","[{\"question\":\"What problem does the document identify with provenance-based attack investigation in the wild?\",\"answer\":\"It highlights dependency explosions and fragmented causal chains that hinder automation in real deployments, making investigation workflows fail or require excessive manual triage.\"},{\"question\":\"What is SherAgent, and how does it work conceptually?\",\"answer\":\"SherAgent is an LLM-empowered automated investigation system that runs an iterative “query-filter” backtracking paradigm over provenance graphs to reason over unstructured investigation context and threat intelligence.\"},{\"question\":\"How does SherAgent address missing events and reduce investigation failure modes?\",\"answer\":\"It dynamically calibrates query conditions to broaden search scope when causal chains are broken, while also applying precision result filtering and strategic node selection to mitigate dependency explosions.\"}]",1784179045,43,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"sheragent-scaling-attack-investigation-in-the-wild-via-llm-empowered-iterative-query-filter-backtracking","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/sheragent-scaling-attack-investigation-in-the-wild-via-llm-empowered-iterative-query-filter-backtracking/82236/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the document identify with provenance-based attack investigation in the wild?","Question",{"text":75,"@type":76},"It highlights dependency explosions and fragmented causal chains that hinder automation in real deployments, making investigation workflows fail or require excessive manual triage.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is SherAgent, and how does it work conceptually?",{"text":80,"@type":76},"SherAgent is an LLM-empowered automated investigation system that runs an iterative “query-filter” backtracking paradigm over provenance graphs to reason over unstructured investigation context and threat intelligence.",{"name":82,"@type":73,"acceptedAnswer":83},"How does SherAgent address missing events and reduce investigation failure modes?",{"text":84,"@type":76},"It dynamically calibrates query conditions to broaden search scope when causal chains are broken, while also applying precision result filtering and strategic node selection to mitigate dependency explosions.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]