[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84307-en":3,"doc-seo-84307-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84307,1374391974585,"Genevieve","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring","Chain-of-thought (CoT) monitoring is proposed as an AI safety mechanism to detect misalignment and deception by inspecting visible reasoning traces. This work stress-tests whether CoT monitors can be coerced: an adversarial agent persuades a monitor to approve actions that violate the monitor’s policy. Across 40 tasks and thousands of interactions, CoT access increases approvals of harmful actions by 9.5% on average. A fact-checking framework with model-diverse pairings reduces approvals by up to 45%, showing CoT monitoring alone can be insufficient.","PERSUASION ATTACKS CAN DECREASE EFFECTIVENESS OF COT  \nMONITORING  \nA PREPRINT  \nJennifer Za∗ LASR Labs  \nJulija Bainiaksina∗  \nLASR Labs University College London  \nNikita Ostrovsky Tanush Chopra Victoria Krakovna  \nLASR Labs LASR Labs Google DeepMind  \narXiv :2607 .08066v 1 [ cs .AI] 9 Jul 2026  \nABSTRACT  \nChain-of-thought (CoT) monitoring is a promising safety mechanism for AI agents, based on the premise that visible reasoning traces can surface misaligned or deceptive behavior. While effective in standard scenarios, recent work highlights that LLMs remain vulnerable to persuasion-based jailbreaks, where natural-language arguments override model constraints. We stress-test whether this vulnerability extends to monitoring LLMs: can an adversarial agent persuade its CoT monitor to approve proposed actions that violate the monitor’s policy? We design an evaluation framework with 40 tasks and analyze thousands of agent-monitor interactions, where agents are instructed to argue for policy-violating proposals. We find that in such adversarial settings, monitor access to the agent’s CoT reasoning increases rather than decreases approval of harmful actions on average by 9.5%, as the scratchpad provides an additional persuasion channel. To address this, we introduce a fact-checking monitoring framework. We find that a fact-checker and monitor pairing from different model families, for example a Claude 3.7 Sonnet monitor paired with a GPT-4.1 fact-checker, reduces approval of policy-violating actions by up to 45%, compared to only 6%, when using the same model for both fact-checking and monitoring roles. Our results demonstrate that CoT monitoring alone maybe insufficient against adversarial persuasion, and that model-diverse fact-checking provides a robust mitigation.  \nKeywords LLM safety · Chain-of-thought monitoring · Adversarial attacks · AI oversight  \n1 Introduction  \nAs large language models are increasingly deployed as autonomous agents—executing code, managing financial transactions, moderating social media content, and interacting with users as chatbots—the need for effective oversight mechanisms has become critical [Shah et al., 2025] . Chain-of-thought (CoT) monitoring has been proposed as a potential tool for detecting misalignment, biases, and misbehavior in frontier models [Baker et al., 2025] . CoT monitoring is deployed as a separate monitor model that observes not only an agent’s actions but also its intermediate reasoning traces. The intuition is compelling: if we can see how an agent reasons, we should be better equipped to catch deceptive or misaligned behavior before it causes harm.  \nWe stress-test the robustness of this approach. In this work, we systematically evaluate whether CoT monitors can be manipulated by adversarial agents that deliberately exploit visible reasoning traces as a persuasion channel to gain approval for policy-violating actions.  \nAcross 40 domain-specific agent tasks spanning software engineering, customer support, financial analysis, and social media moderation, we find that:  \n1. Agents can persuade monitors. When instructed to justify policy-violating actions, LLM agents successfully convince monitors to approve these actions with acceptance rates up to 47% relative to the baseline.  \n∗Equal contribution.  \nFigure 1: Evaluation environments. We compare three monitoring setups for policy-violating actions: justification-only monitoring; justification + Chain-of-Thought (CoT) monitoring, where the agent exposes a scratchpad; and justification + CoT monitoring with fact-checking, using either an internal or external fact-checker. The agent generates justifications rather than executing actions, and may iteratively refine its arguments based on monitor feedback.  \n2. In adversarial settings, CoT access can increase persuasion. When agents are instructed to argue for explicitly policy-violating actions, giving monitors access to agent’s reasoning traces (via a scratchpad) increases appr","cbCaivKMLsY4SUNG","https://ap.wps.com/l/cbCaivKMLsY4SUNG","pdf",1159907,4,1,25,"English","en",105,"# Abstract\n# Introduction\n## Stress-testing CoT monitoring robustness\n## Experimental findings and comparisons\n### Persuasion success against monitors\n### CoT access increases approvals\n### Monitor model selection effects\n### Same-family vs cross-family fact-checking","[{\"question\":\"What does the preprint investigate about CoT monitoring?\",\"answer\":\"It investigates whether adversarial agents can persuade a CoT monitor to approve actions that violate the monitor’s policy.\"},{\"question\":\"How do the authors evaluate the vulnerability of CoT monitors?\",\"answer\":\"They design an evaluation framework with 40 domain-specific tasks and analyze thousands of agent–monitor interactions where agents argue for policy-violating proposals.\"},{\"question\":\"What mitigation is proposed, and how effective is it?\",\"answer\":\"The authors introduce a fact-checking monitoring framework using a fact-checker paired from a different model family; this cross-family setup reduces approvals of policy-violating actions to around 6% on average, and up to 45% reduction in the reported results.\"}]",1784194712,63,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"persuasion-attacks-can-decrease-effectiveness-of-cot-monitoring","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/persuasion-attacks-can-decrease-effectiveness-of-cot-monitoring/84307/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What does the preprint investigate about CoT monitoring?","Question",{"text":75,"@type":76},"It investigates whether adversarial agents can persuade a CoT monitor to approve actions that violate the monitor’s policy.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How do the authors evaluate the vulnerability of CoT monitors?",{"text":80,"@type":76},"They design an evaluation framework with 40 domain-specific tasks and analyze thousands of agent–monitor interactions where agents argue for policy-violating proposals.",{"name":82,"@type":73,"acceptedAnswer":83},"What mitigation is proposed, and how effective is it?",{"text":84,"@type":76},"The authors introduce a fact-checking monitoring framework using a fact-checker paired from a different model family; this cross-family setup reduces approvals of policy-violating actions to around 6% on average, and up to 45% reduction in the reported results.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]