[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85353-en":3,"doc-seo-85353-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85353,13056703020460,"Valentina","https://ap-avatar.wpscdn.com/avatar/be000253dac470eee5d?_k=1778207105932848923",8,"Research & Report","Agent Hacks Agent: Autoresearch for Production-Agent Red-Teaming","Production LLM agents like Claude Code and Codex can take real actions through deployed interfaces over untrusted content, files, commands, and workspace state, so safety failures manifest as written files, exfiltrated data, or triggered workflows. Red-teaming must track rapid model and tool updates, yet today’s artifacts emphasize attack success while missing the enabling condition behind unsafe trajectories, limiting auditing, patching, and reuse. The work studies autoresearch for production-agent red-teaming and introduces AHA, a falsifiable discovery loop that builds an auditable Vulnerability Concept Graph for reusable, transferable safety knowledge.","AGENT HACKS AGENT:  \nAUTORESEARCH FOR PRODUCTION-AGENT REDTEAMING  \nXutao Mao  \nCity University of Hong Kong  \n[xutao.henry.mao@gmail.com](xutao.henry.mao@gmail.com)  \nXiang Zheng†  \nCity University of Hong Kong  \narXiv :2607 . 1 1698v 1 [ cs .CR] 13 Jul 2026  \nCong Wang†  \nCity University of Hong Kong  \nABSTRACT  \nProduction LLM agents such as Claude Code and Codex act through deployed interfaces over untrusted content, files, commands, and workspace state, so a safety failure here is a real action: a written file, exfiltrated data, or a triggered workflow. Red-teaming these agents must keep pace with every model and tool update, yet today’s tools optimize judged attack success and preserve surface artifacts:  \nbenchmark scores, payloads, archives, strategies, or attack programs. These artifacts record where an attack landed, but not the enabling condition that made the agent trajectory unsafe, so they are hard to audit, patch against, or reuse after the setting changes. We study autoresearch for production-agent red-teaming, using one agentic research environment to automatically discover reusable vulnerability knowledge about another production-style agent. We present AHA, a falsifiable discovery loop: it commits to a vulnerability hypothesis, creates a falsifier, instantiates a scenario-valid attack, executes it in a sandboxed agent harness, reflects on the trajectory, and promotes confirmed findings by an evidence rule into a Vulnerability Concept Graph (VCG) . Each concept is an auditable unit linking an attacker-facing surface to an unsafe trajectory through a claim, enabling condition, falsifier, transfer prediction, and evidence. Across Claude Code and Codex on three scenarios spanning direct and indirect attacks, the discovered concepts share a core that recurs across victim models and agents, the frozen VCG is reusable with no further search, outperforming the strongest frozen discovery baseline by 14.2 percentage points under the same single-shot protocol, and the concepts transfer across scenarios and across direct/indirect attack channels. This makes the artifact directly useful for production triage: a safety team can inspect the enabling condition, patch the agent or workflow, rerun the concept as a check on the fix, and attach new internal concerns through the same build/import scenario workflows. As production agents proliferate, such a VCG turns one-off red-teaming into cumulative, auditable safety knowledge that compounds across models and products. Our code is available at [https:](https:)//[github.com/henrymao2004/Auto-research-red-teaming](github.com/henrymao2004/Auto-research-red-teaming).  \n1 INTRODUCTION  \nProduction LLM agents such as Claude Code and Codex write code, execute tools, and operate as autonomous engineering environments, wired into real file systems, APIs, tool permissions, and team workflows (Anthropic, 2026a ; OpenAI, 2025 ; Guo et al., 2025 ; Meng et al., 2026 ; Pan et al., 2025) . A safety failure here is no longer a model emitting harmful text; it is the agent taking a real action: writing a file, exfiltrating data, modifying code, invoking a tool, or triggering a workflow (Guo et al., 2024 ; Zhang et al., 2024 ; Guo et al., 2025 ; Meng et al., 2026 ; Puppala et al., 2026 ; Wang  \n†Co-corresponding authors.  \nFigure 1: AHA overview. An autoresearch loop turns executed red-team trajectories into a frozen, reusable VCG, the auditable artifact this paper produces and evaluates.  \net al., 2025a ; Pan et al., 2025 ; Zhang & Pei, 2026) . Agent failures are becoming operational failures, and the ones that matter are trajectory-level, where an agent reads attacker-controlled content, chooses a sequence of tool calls, and realizes harm through files, commands, or workspace state (Chen et al., 2026a ; Feng et al., 2026 ; Li et al., 2026b ; Zhang et al., 2026) . Red-teaming such systems must keep pace with every new model, tool integration, permission boundary, and safety patch. Static benchma","cbCaigIJPCMiH8s1","https://ap.wps.com/l/cbCaigIJPCMiH8s1","pdf",2119377,2,1,50,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"Why are safety failures in production LLM agents considered more serious than harmful text outputs?\",\"answer\":\"Because the agents can execute real actions—writing files, exfiltrating data, modifying code, invoking tools, or triggering workflows—so the failure occurs at the trajectory level in the environment.\"},{\"question\":\"What limitation do existing red-teaming artifacts have for auditing and patching?\",\"answer\":\"They mainly record where an attack landed (benchmark scores, payloads, archives, strategies), but not the enabling condition that caused the agent’s unsafe trajectory.\"},{\"question\":\"How does AHA enable reusable and auditable vulnerability knowledge for red-teaming?\",\"answer\":\"AHA runs a falsifiable discovery loop: it commits to a vulnerability hypothesis, creates a falsifier, instantiates a scenario-valid attack, executes in a sandboxed harness, reflects on the trajectory, and promotes confirmed findings into a Vulnerability Concept Graph with evidence and validity conditions.\"}]",1784202743,126,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"agent-hacks-agent-autoresearch-for-production-agent-red-teaming","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/agent-hacks-agent-autoresearch-for-production-agent-red-teaming/85353/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why are safety failures in production LLM agents considered more serious than harmful text outputs?","Question",{"text":75,"@type":76},"Because the agents can execute real actions—writing files, exfiltrating data, modifying code, invoking tools, or triggering workflows—so the failure occurs at the trajectory level in the environment.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What limitation do existing red-teaming artifacts have for auditing and patching?",{"text":80,"@type":76},"They mainly record where an attack landed (benchmark scores, payloads, archives, strategies), but not the enabling condition that caused the agent’s unsafe trajectory.",{"name":82,"@type":73,"acceptedAnswer":83},"How does AHA enable reusable and auditable vulnerability knowledge for red-teaming?",{"text":84,"@type":76},"AHA runs a falsifiable discovery loop: it commits to a vulnerability hypothesis, creates a falsifier, instantiates a scenario-valid attack, executes in a sandboxed harness, reflects on the trajectory, and promotes confirmed findings into a Vulnerability Concept Graph with evidence and validity conditions.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":22,"slug":113},6,"Technology","technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]