[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82720-en":3,"doc-seo-82720-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82720,4398048949847,"Eliana","https://ap-avatar.wpscdn.com/avatar/400002536579ef2da7f?_k=1778318612642679267",8,"Research & Report","CONTRA: 红队测试个性化智能体的配置树搜索方法","Recent LLM-based agent systems extend beyond dialogue into autonomous, customizable workflows through modifiable internal files and installable skills. This added capability increases the risk of unintended harmful actions. The work studies how agent configuration interacts with the likelihood of executing dangerous behaviors without explicit instruction. It proposes CONTRA, an LLM-assisted tree-search method that discovers benign-looking configurations by simulating agent execution, building a skill–malicious action dataset, and evaluating attack success rates.","arXiv :2607 .03220v 1 [ cs .CR] 3 Jul 2026  \nCONTRA: Red-Teaming Configurations of Personalizable Agents  \nJonathan Nöther, Adish Singla, Goran Radanovic  \nMax Planck Institute for Software Systems, Germany  \n{jnoether,adishs,[gradanovic}](gradanovic}@mpi-sws.org)[@mpi-sws.org](gradanovic}@mpi-sws.org)  \nAbstract  \nRecent tools such as OpenClaw have extended the capabilities of LLM-based agents from simple dialog-based systems to fully autonomous agents. These systems allow personalization of the agent through modifiable internal files and the installation of skills. While this enables deployment in a wide range of settings and the automation of diverse tasks, greater capability and autonomy increases the risk of malicious actions being executed unintentionally. In this work, we explore the interplay between agent configuration and the risk of executing dangerous actions without explicit instruction. To this end, we propose CONfiguration Tree-search for Redteaming Agents (CONTRA), an LLM-assisted tree-search algorithm that discovers agent configurations resulting in the execution of malicious actions. CONTRA works by reasoning about benign yet dangerous configurations and evaluating them ina simulated environment. We construct a dataset of the 473 most popular skills from a public repository, along with 2–5 corresponding malicious target actions per skill. In a large-scale analysis, we find that 75. 1% of skills have at least one configuration resulting in the execution of a malicious action, most of which have not been detected as containing malicious content by existing scans. Overall, CONTRA successfully identifies a configuration leading to the execution of the target action in 39 .2% of all tested cases. Our findings demonstrate that current agents provide insufficient safety with respect to personalization.  \n1 Introduction  \nGeneral-purpose personal assistants, such as OpenClaw, possess the versatile reasoning required to manage complex workflows across emails, messages, and private files, as well as interacting with both humans and external agents. While these capabilities allow for novel useful automation, this same integration introduces a significant surface area for malicious actions and unintended consequences.  \nA defining feature of these agents is deep personalization using the agent’s internal configuration files. By maintaining persistent notes of past interactions, the agent evolves a specialized memory and personality that is injected into its context. This allows both the memorization of past experiences, such as enabling learning from past failure, as well as customizing the behavior of the agent towards the user’s needs. Further, the agent’s utility can additionally be extended by installing custom skills, markdown files that instruct the agent on when and how to use external APIs and services. While these features drive utility, they also create a \"black-box\" of behavioral safety, since an agent’s actions are a direct byproduct of its unique configuration. In this paper, we explore the red-teaming of agent configurations with regards to agent skills. We specifically investigate whether, given a specific skill and a malicious target action, there exist a configuration that results in the agent executing that target action. The configuration itself should be benign, i.e. it should not use any adversarial tactics or directly instruct the agent to perform the target action. Systematically answering this allows us to map the latent vulnerabilities introduced by user-driven personalization.  \nPreprint.  \nConfiguration Generation Sandboxed Execution  \nFigure 1: Illustration of CONTRA using a simplified, yet real example. We start with the skill being evaluated, a target action and an archive of previous attempts. We sample one configurations from the archive and instruct the Orchestrator to reason about potential changes that are benign but could lead to the target action. This is then given to the relevant su","cbCaiaki0xNmJ42f","https://ap.wps.com/l/cbCaiaki0xNmJ42f","pdf",1032703,5,1,23,"English","en",105,"# Introduction\n## Configuration Generation and Sandboxed Execution\n## Large-Scale Evaluation","[{\"question\":\"CONTRA研究的核心问题是什么？\",\"answer\":\"研究如何在给定特定技能和恶意目标动作的情况下，找到一种“看似无害”的智能体配置，使其在缺乏明确指令时仍会执行目标恶意动作。\"},{\"question\":\"CONTRA是如何发现能触发恶意行为的配置的？\",\"answer\":\"CONTRA通过对配置文件进行LLM辅助的树搜索，推理哪些配置变更仍“良性”但可能导致目标动作，并在沙箱模拟环境中评估结果，由评审器判断动作与消息是否命中目标。\"},{\"question\":\"实验的主要发现有哪些？\",\"answer\":\"在对ClawHub上473个最受欢迎技能进行大规模测试后，75.1%的技能至少存在一个会导致恶意动作执行的配置；39.2%的测试动作成功被触发，且多数导致触发的配置本身在评估中被认为是良性的。\"}]",1784182490,58,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"contra-red-teaming-configurations-of-personalizable-agents","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/contra-red-teaming-configurations-of-personalizable-agents/82720/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"CONTRA研究的核心问题是什么？","Question",{"text":76,"@type":77},"研究如何在给定特定技能和恶意目标动作的情况下，找到一种“看似无害”的智能体配置，使其在缺乏明确指令时仍会执行目标恶意动作。","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"CONTRA是如何发现能触发恶意行为的配置的？",{"text":81,"@type":77},"CONTRA通过对配置文件进行LLM辅助的树搜索，推理哪些配置变更仍“良性”但可能导致目标动作，并在沙箱模拟环境中评估结果，由评审器判断动作与消息是否命中目标。",{"name":83,"@type":74,"acceptedAnswer":84},"实验的主要发现有哪些？",{"text":85,"@type":77},"在对ClawHub上473个最受欢迎技能进行大规模测试后，75.1%的技能至少存在一个会导致恶意动作执行的配置；39.2%的测试动作成功被触发，且多数导致触发的配置本身在评估中被认为是良性的。","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]