[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83334-en":3,"doc-seo-83334-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83334,687197207919,"Theodora","https://ap-avatar.wpscdn.com/avatar/a000253d6f5f7c60be?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779446848396160552",8,"Research & Report","Overthinking: Amplifying Reasoning Weights to Extract Learned Secrets","Black box auditing of language models can miss subtle misalignment and hidden information. The paper introduces overthinking, amplifying a reasoning model’s tendency to reveal internal information by using reasoning task vectors to amplify the “reasoning direction” in weight space. Given an instruct model M and a distilled reasoning model R, it defines an overthinking model with a scaling factor α>1. Layer-wise attenuation and experiments on 2B–32B models show up to 10× more frequent secret disclosure while preserving coherence within an optimal auditing range.","Overthinking: Amplifying Reasoning Weights to Extract Learned Secrets  \nJack Hopkins * 1 2 Dipika Khullar * 1 2 Fabien Roger 3  \narXiv :2607 .08 173v 1 [ cs .AI] 9 Jul 2026  \nAbstract  \nBlack box auditing of language models is an essential pre-deployment tool, but it may miss subtle forms of misalignment and hidden information. To better elicit hidden information during an auditing process, we introduce overthinking: the process of using reasoning task vectors to amplify the propensity to think out loud of reasoning models. Given the parameters of anon-reasoning instruct model M and reasoningdistilled model R, we define the overthinking model as θ Oα = θM + α (θR − θM ), where α > 1 amplifies reasoning beyond the pure reasoning model R. Additionally, we introduce new layer-wise attenuation strategies that selectively amplify reasoning without losing quality and coherence of model outputs. We demonstrate that overthinking models are more likely to reveal hidden information across four experimental settings, across 2B-32B models. Our findings suggest that reasoning amplification may surface secrets or unintended behaviors acquired during training up to 10 × more frequently than the original reasoning model. How secrets surface depends on the secret type: some require perturbation along the reasoning direction, while others yield to any sufficiently large weight perturbation.  \n1. Introduction  \nEnsuring that AI systems can be audited for unintended behaviors is a central challenge for safe deployment (Christiano et al., 2017 ; Ouyang et al., 2022) . Models are trained on increasingly complex objectives and may acquire unintended goals or behaviors that remain latent under standard evaluation (Denison et al., 2024 ; Hubinger et al., 2024) .  \n1Anthropic Fellows Program 2MATS 3Anthropic. Correspondence to: Jack Hopkins \u003C[jack.hopkins@me.com](jack.hopkins@me.com) >, Dipika Khullar \u003C[dkhullar98@berkeley.edu](dkhullar98@berkeley.edu) >, Fabien Roger \u003Cfa[bien@anthropic.com](bien@anthropic.com) >.  \nProceedings of the 43 rd International Conference on Machine Learning, Seoul, South Korea. PMLR 306, 2026 . Copyright 2026 by the author(s) .  \nBehavioral red-teaming, while valuable, faces fundamental limitations: auditors cannot enumerate all possible failure modes, and models may behave differently under evaluation than deployment (Casper et al., 2024 ; Shevlane et al., 2023) .  \nThis has motivated research into auditing games that formalize the adversarial relationship between models and evaluators (Greenblatt et al., 2024) . As models are increasingly deployed in sensitive contexts, developing systematic techniques to surface what models have learned, beyond what black-box testing can reveal, becomes critical.  \nWe propose using task vector arithmetic (Ilharco et al., 2023) to increase reasoning faithfulness by amplifying the “reasoning direction” in weight space. Given a non-reasoning instruct model M and a reasoning model R, we define the reasoning task vector τreason = θR − θM , capturing the direction in weight space that encodes deliberative reasoning capability. We then apply this task vector to go beyond a regular reasoning model, and get what we call an overthinking model: θ Oα = θM + ατreason = θM + α(θR − θM ), where α > 1 (α = 0 would recover the instruct model M,α = 1 would recover the reasoning model R) . Our hypothesis is that amplification increases the probability that models disclose auditing-relevant information they would otherwise conceal.  \nWe validate this approach using model organisms—models that have secrets analogous to the kind of secret properties that auditing is supposed to catch (Hubinger et al., 2024 ; Cywi´nski et al., 2025) . We study four settings. incontext secrets is a reasoning faithfulness setting inspired by Turpin et al. (2023) where models are asked to answer questions with hidden metadata containing correct answers-which influences their answers without being verbalized in the chain-of-thought","cbCaijeVupft4bMz","https://ap.wps.com/l/cbCaijeVupft4bMz","pdf",1613561,3,1,19,"English","en",105,"# Abstract\n# Introduction\n## Reasoning Vector Task Method\n## Overthinking Models Leak Learned Secrets\n## Auditing Efficacy vs. Coherence","[{\"question\":\"What problem does the paper address in black-box model auditing?\",\"answer\":\"Standard black-box auditing may overlook subtle misalignment and hidden information. The paper targets the gap by improving the elicitation of auditing-relevant disclosures.\"},{\"question\":\"How does the paper define an overthinking model mathematically?\",\"answer\":\"It uses task vector arithmetic: the overthinking parameters are θ_Oα = θ_M + α(θ_R − θ_M), where α\\u003e1 amplifies reasoning beyond the base reasoning model.\"},{\"question\":\"How is auditing effectiveness balanced against output coherence?\",\"answer\":\"The paper reports a window of intermediate α values where models disclose more while remaining coherent, while extreme amplification can destabilize outputs and reduce coherence.\"}]",1784186796,48,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"overthinking-amplifying-reasoning-weights-to-extract-learned-secrets","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/overthinking-amplifying-reasoning-weights-to-extract-learned-secrets/83334/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address in black-box model auditing?","Question",{"text":75,"@type":76},"Standard black-box auditing may overlook subtle misalignment and hidden information. The paper targets the gap by improving the elicitation of auditing-relevant disclosures.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the paper define an overthinking model mathematically?",{"text":80,"@type":76},"It uses task vector arithmetic: the overthinking parameters are θ_Oα = θ_M + α(θ_R − θ_M), where α>1 amplifies reasoning beyond the base reasoning model.",{"name":82,"@type":73,"acceptedAnswer":83},"How is auditing effectiveness balanced against output coherence?",{"text":84,"@type":76},"The paper reports a window of intermediate α values where models disclose more while remaining coherent, while extreme amplification can destabilize outputs and reduce coherence.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},"General","general"]