[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83206-en":3,"doc-seo-83206-105":29,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},83206,1374391974468,"Eden","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Reasoning Consistency Scanning: Auditing Chain-of-Thought Validity in AI Safety Evaluations","Chain-of-thought (CoT) explanations in AI safety evaluations can be unfaithful: stated reasoning may not match the process that produced the final output. This work introduces reasoning consistency scanning, a method that checks whether transcript reasoning is logically consistent with the accompanying answer, enabling transcript-only post-hoc auditing. The paper formalizes consistency separately from faithfulness, proposes a six-subtype taxonomy of inconsistency, builds a 60-transcript benchmark, implements an InspectScout scanner, and reports systematic results across models and task types.","arXiv :2607 .07229v 1 [ cs .AI] 8 Jul 2026  \nReasoning Consistency Scanning: A Framework for Auditing Chain-of-Thought Validity in AI Safety Evaluations  \nSilvia Santano  \nAbstract  \nPrior work has shown that chain-of-thought (CoT) reasoning is often unfaithful: a model’s stated reasoning does not reliably reflect the process that produced its output [Turpin et al., 2023, Lanham et al., 2023, Chen et al., 2025, Walden and Wanner, 2026] . Detecting unfaithfulness, though, requires controlled experimental interventions, which cannot be applied to evaluation transcripts after the fact. We turn instead to a more tractable question that has received less attention: whether the stated reasoning is logically consistent with the answer it accompanies. Unlike faithfulness, consistency can be assessed from a transcript alone, with no intervention. We introduce reasoning consistency scanning, a reusable method for detecting this property in AI safety evaluation transcripts. Our contributions are fourfold. First, we formalize reasoning consistency as distinct from faithfulness and define a six-subtype taxonomy of inconsistency. Second, we build a validated benchmark of 60 transcripts, manually adapted from InstrumentalEval outputs. Third, we implement a working scanner for InspectScout, the first to target this property in safety evaluation transcripts. Fourth, we report results across four generator models and three evaluations from inspect evals, showing that reasoning inconsistency is present, detectable, and varies systematically across both models and task types.1  \n1 Introduction  \nChain-of-thought reasoning has become central to how AI safety evaluations are interpreted. When a model is tested for dangerous capability, deception, or alignment-relevant behavior, evaluators increasingly examine not just what the model outputs but how it reasons about the task. Reasoning traces are treated as evidence that a model is genuinely deliberating, recognizing constraints, or weighing consequences, which are all properties that matter for safety assessments.  \nThis practice rests on an assumption that is less examined than it might appear: that the model’s stated reasoning is actually connected to its final answer. However, they often are not. A growing body of research on CoT faithfulness has demonstrated that reasoning traces frequently do not reflect the internal computational process that produced the output [Turpin et al., 2023, Chen et al., 2025] . Models act on information they do not mention in their reasoning, and deny reliance on features that demonstrably influenced their answers [Walden and Wanner, 2026] .  \nHowever, faithfulness and consistency are distinct properties requiring different methods to detect. Faithfulness detection depends on controlled interventions such as injecting hints, perturbing inputs and measuring what changes, and establishes causal claims about internal computation. These methods cannot be applied to existing transcript corpora. Consistency, by contrast, asks a simpler question: ”does the stated reasoning logically connect to and predict the final output, as observable text?” This is detectable from transcripts alone, without any experimental intervention, making it amenable to systematic post-hoc auditing.  \nThe direct consequence is that, if reasoning and answer are decoupled, any eval that uses reasoning traces as evidence of a safety-relevant property is measuring something different from what it claims. This is a direct threat to construct validity, the foundational question of whether an eval is actually measuring the property it was designed to measure. When a model’s reasoning trace does not connect to its output, two problems follow. First, the eval result cannot be trusted. Whether the inconsistency stems from confusion, post-hoc rationalisation, or strategic behavior, theeval is not measuring what it claims. And second, CoT monitoring is undermined. Safety researchers increasingly rely on reaso","cbCaihxfZZ38iylR","https://ap.wps.com/l/cbCaihxfZZ38iylR","pdf",352868,1,13,"English","en",105,"# Introduction\n## Faithfulness versus consistency\n## Construct validity and risks of inconsistency\n## Reasoning consistency scanning methodology\n## Contributions and empirical setup","[{\"question\":\"What are the paper’s main contributions?\",\"answer\":\"It formalizes reasoning consistency with a six-subtype inconsistency taxonomy and decision procedure, creates a validated benchmark of 60 labeled transcripts, implements a working InspectScout scanner using an LLM-as-judge design, and reports results across multiple generator models and evaluation sets.\"}]",1784185936,33,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":27},"reasoning-consistency-scanning-auditing-chain-of-thought-validity-in-ai-safety-evaluations","",{"@graph":35,"@context":77},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/reasoning-consistency-scanning-auditing-chain-of-thought-validity-in-ai-safety-evaluations/83206/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"What are the paper’s main contributions?","Question",{"text":75,"@type":76},"It formalizes reasoning consistency with a six-subtype inconsistency taxonomy and decision procedure, creates a validated benchmark of 60 labeled transcripts, implements a working InspectScout scanner using an LLM-as-judge design, and reports results across multiple generator models and evaluation sets.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,120,123,127],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":45,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":45,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":45,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":45,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":128,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":98,"slug":130},19,"General","general"]