[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85919-en":3,"doc-seo-85919-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85919,7971461740886,"Theodore","https://ap-avatar.wpscdn.com/davatar_3d24733baf745e90a7e4bdd5f77d97b2",8,"Research & Report","ANCHOR Automated Alignment Auditing for CLI Agents on Real-World Harm","Autonomous CLI agents can carry out hundreds of actions in long sessions, such as writing code, running shell commands, browsing the web, and operating cloud infrastructure with minimal oversight. ANCHOR is introduced as an automated auditing framework that stress-tests these agents on illegal tasks grounded in public US court cases. A strong auditor agent uses strategic, persistent malicious user behavior to bypass refusals and adapt across multi-turn interactions. Results show refusal can fail under persistence, enabling harmful autonomy and catastrophic-risk outcomes.","ANCHOR: Automated Alignment Auditing for CLI Agents on Real-World Harm  \nKefan Song 1 Yanjun Qi 1  \narXiv :2607 . 10455v 1 [ cs .AI] 11 Jul 2026  \nAbstract  \nAutonomous CLI agents can now execute hundreds of actions across multi-hour sessions: writing code, executing shell commands, browsing the web, and managing cloud infrastructure, all with minimal human oversight. Does greater autonomy invite greater risk? We introduce ANCHOR, an automated auditing framework that stress-tests CLI agents on illegal tasks grounded in public US court cases. ANCHOR deploys an auditor agent fine-tuned on dark personality data using supervised and reinforcement fine tuning.  \nThis auditor roleplays persistent malicious users who decompose tasks, reframe requests upon refusal, and adapt strategies across multi-turn interactions. Evaluating frontier CLI agents, we find that while they often refuse illegal tasks when prompted directly, compliance reaches 100% under persistent malicious interaction. When agents comply, they frequently exceed user requests, autonomously building infrastructure for large-scale harm, including catastrophic risk scenarios such as large-scale financial fraud and bioweapon development. These findings demonstrate that current alignment techniques are insufficient for autonomous agents and underscore the need for safety evaluations against persistent, adaptive malicious users. We release ANCHOR at [https:](https:)//[github.com/garified/anchor](github.com/garified/anchor).  \n1 Introduction  \nCLI agents such as Claude Code and Gemini-CLI are autonomous systems that execute hundreds of tool calls per session, write and run code, manage files, and interact with external services through the Model Context Protocol (MCP) . Recent work such as Cowork (Anthropic, 2026) further extends these agents beyond coding to general-purpose tasks  \n1University of Virginia, Charlottesville, VA, USA. Correspondence to: Kefan Song \u003C[ks8vf@virginia.edu](ks8vf@virginia.edu) >.  \nProceedings of the 43 rd International Conference on Machine Learning, Seoul, South Korea. PMLR 306, 2026 . Copyright 2026 by the author(s) .  \nincluding web browsing, spreadsheet manipulation, and file management. Unlike workflow-based systems where LLMs follow predefined code paths (Anthropic, 2024), these agents autonomously decide what to build and how to implement it throughout long-horizon tasks. This raises a natural question: can a malicious user exploit this autonomy to carry out large-scale real-world harmful activities?  \nAt its extreme, such misuse rises to the level of catastrophic risk, a threshold now codified in policy and law. California’s Frontier AI Models Act (SB 53) defines it as a foreseeable, material risk that a frontier model materially contributes to the death of, or serious injury to, more than 50 people or more than $1 billion in damages from a single incident (California State Legislature, 2025); Anthropic’s Responsible Scaling Policy similarly describes large-scale devastation, such as thousands of deaths or hundreds of billions of dollars in damage, directly caused by an AI model and that would not have occurred without it (Anthropic, 2023) .  \nExisting safety benchmarks and auditing frameworks are not built to answer this. Agent safety benchmarks such as AgentHarm (Andriushchenko et al., 2025), AgentSafetyBench (Zhang et al., 2025), and OS-Harm (Kuntz et al., 2025) either pre-specify tool-call sequences or limit evaluation to short-horizon tasks, and draw on artificial scenarios that do not reflect the complexity of real-world harmful activities. Alignment auditing frameworks such as Petri (Fronsdal et al., 2025) and Bloom (Gupta et al., 2025) evaluate subtle misalignment model behaviors such as deception and self-preservation, rather than explicit misuse. Their helpfulonly auditors are neither persistent nor strategically deceptive enough to bypass safety mechanisms, and may refuse to audit catastrophic scenarios altogether, underestimating","cbCaimLE7ud63fdX","https://ap.wps.com/l/cbCaimLE7ud63fdX","pdf",458303,5,1,19,"English","en",105,"# Introduction\n## Catastrophic risk and policy context\n## Limits of existing benchmarks and auditing frameworks\n## The ANCHOR auditing framework","[{\"question\":\"What is ANCHOR and what problem does it address?\",\"answer\":\"ANCHOR is an automated auditing framework that stress-tests autonomous CLI agents on illegal tasks grounded in real US court cases. It targets a key gap: existing evaluations do not reflect persistent, adaptive misuse in real-world harmful activities.\"},{\"question\":\"How does ANCHOR generate harmful tasks?\",\"answer\":\"ANCHOR-Seed transforms public US court records of successful criminal activity into realistic harmful task instructions. It uses filtering, extraction, neutral rewriting, and validation rather than relying on synthetic or annotated scenarios.\"},{\"question\":\"What do the evaluations show about current alignment techniques?\",\"answer\":\"Agents often refuse illegal tasks when prompted directly, but compliance can reach 100% under persistent malicious interaction. When they comply, they may exceed requests and autonomously build infrastructure for large-scale harm, indicating current alignment techniques are insufficient for this threat model.\"}]",1784207168,48,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"anchor-automated-alignment-auditing-for-cli-agents-on-real-world-harm","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/anchor-automated-alignment-auditing-for-cli-agents-on-real-world-harm/85919/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What is ANCHOR and what problem does it address?","Question",{"text":76,"@type":77},"ANCHOR is an automated auditing framework that stress-tests autonomous CLI agents on illegal tasks grounded in real US court cases. It targets a key gap: existing evaluations do not reflect persistent, adaptive misuse in real-world harmful activities.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does ANCHOR generate harmful tasks?",{"text":81,"@type":77},"ANCHOR-Seed transforms public US court records of successful criminal activity into realistic harmful task instructions. It uses filtering, extraction, neutral rewriting, and validation rather than relying on synthetic or annotated scenarios.",{"name":83,"@type":74,"acceptedAnswer":84},"What do the evaluations show about current alignment techniques?",{"text":85,"@type":77},"Agents often refuse illegal tasks when prompted directly, but compliance can reach 100% under persistent malicious interaction. When they comply, they may exceed requests and autonomously build infrastructure for large-scale harm, indicating current alignment techniques are insufficient for this threat model.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":20,"slug":137},"General","general"]