[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84931-en":3,"doc-seo-84931-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84931,687197207639,"Asher","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Calibration-Family Overfit: Why Trusted Sabotage Monitors Don’t Transfer Across Lineages","Trusted monitoring is a key AI control strategy where a cheaper trusted model flags an untrusted model’s potentially sabotaging actions for auditing or deferral. Existing evaluations treat monitor quality as fixed, but performance may depend on which untrusted policy family it was calibrated on. Using family-conditioned transcripts, a two-way decomposition isolates an interaction term indicating calibration overfit. On code-backdoor transcripts, the cross-family interaction is consistently positive and persists under strict, leak-free controls.","arXiv :2607 .06596v 1 [ cs .CR] 6 Jul 2026  \nCalibration-Family Overfit: Why Trusted Sabotage Monitors Don’t Transfer Across Lineages  \nLucas Pinto  \nIndependent Researcher [lcsdgpinto@gmail.com](lcsdgpinto@gmail.com)  \nAbstract  \nTrusted monitoring is a central defense in AI control: a cheaper trusted model scores an untrusted model’s actions for sabotage, and the most suspicious are audited or deferred. Such monitors are evaluated against one or two untrusted models, and the resulting accuracy is reported asa property of the monitor. We ask whether it is partly a property of the pairing. We make the untrusted policy family the controlled axis: we fit a monitor on family A’s transcripts, apply it to family B, and decompose the cross-family AUROC into how obvious each family’s sabotage is, how capable each monitor is, and the residual own-family advantage after both are removed: the interaction. A positive interaction means a monitor ranks its own lineage’s sabotage above an attack from a different lineage after those row and column main effects are removed; in our matched controls, this remains true when attacks are matched on two obviousness proxies.  \nOn code-backdoor transcripts the interaction is positive and survives the dominant confounds: +0.172 (95% CI [+0.158, +0.185]) on four open-weight families on a strict leak-free basis, holding across an eight-family matrix spanning four frontier vendors, blatancy-matched controls, a trained openweight transcript-reading monitor up to 34B parameters, and a second code benchmark. The gap is geometric: each lineage’s sabotage occupies a different scoring direction, and rotating the monitor’s scoring direction toward the target lineage recovers most of the lost detection while a matched random rotation recovers nothing.  \nTranslated to a defer-to-trusted protocol at a 1% audit budget, an off-lineage monitor catches 19% of attack attempts where an in-lineage monitor catches 41% at the same audit cost, a +21.7-point gap on the strict leak-free basis (42% versus 24% on the transductive basis) . We measure attack-attempt detection as the primary axis; realized sabotage is frontier-only on this substrate, too rare among the open families to power a full matrix, though on the frontier lineages that do realize it the interaction survives and is, by effect size, if anything sharper. We flag the absence of an all-family realized-sabotage matrix as a limitation. The takeaway for practice is that control evaluations should report cross-family transfer matrices, because a single-pairing accuracy overstates the safety a monitor delivers against a lineage it was not calibrated on; the four-step protocol that acts on the gap (report the full matrix, detect off-lineage drift, recalibrate with labeled sabotage, treat pooling as a bounded partial defense) is in Section 5 .  \n1 Introduction  \nAI-control protocols defend against a misaligned untrusted model by routing its actions through a trusted monitor: the monitor assigns a suspicion score, and a fixed fraction of the most-suspicious actions are audited by a more expensive process or deferred to a trusted policy. The protocol’s safety depends on the monitor ranking genuine sabotage above benign behavior at the chosen operating point. Evaluations of these protocols typically  \nmeasure a monitor against one or two untrusted policies and report a single AUROC or a single TPR at a fixed false-positive rate.  \nThat reporting quietly assumes monitor quality is invariant to which untrusted model produced the actions. Deployment violates the assumption in an obvious way: the model you calibrate your monitor on is rarely the exact model you face, and is never the model you will face next quarter. If a monitor’s effectiveness is partly a function of the specific policy lineage it was tuned against, then a published accuracy is an in-lineage number, and the realized accuracy on a different lineage can be lower.  \nWe make the untrusted policy family the cont","cbCailQLvOHPP25d","https://ap.wps.com/l/cbCailQLvOHPP25d","pdf",1441745,1,28,"English","en",105,"# Abstract\n# Introduction\n## Protocol and evaluation setup\n## Three explanations for cross-family drops\n## Headline interaction statistic and findings","[{\"question\":\"What problem does the document address in trusted AI monitoring?\",\"answer\":\"It addresses the assumption that a monitor’s effectiveness is invariant to which untrusted policy produced the actions, noting that deployment shifts the model lineage over time.\"},{\"question\":\"How does the method test whether calibration overfit drives cross-family transfer failure?\",\"answer\":\"It treats the untrusted policy family as the controlled axis, fits monitors on one family’s transcripts, applies them to other families, and decomposes cross-family AUROC into obviousness, capability, and an interaction term.\"},{\"question\":\"What are the main results regarding cross-family sabotage detection?\",\"answer\":\"The interaction term is positive and significant on code-backdoor transcripts, and when converted into a defer-to-trusted protocol under a 1% audit budget, an off-lineage monitor detects substantially fewer attack attempts than an in-lineage monitor at the same cost.\"}]",1784199445,71,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"calibration-family-overfit-why-trusted-sabotage-monitors-dont-transfer-across-lineages","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/calibration-family-overfit-why-trusted-sabotage-monitors-dont-transfer-across-lineages/84931/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the document address in trusted AI monitoring?","Question",{"text":75,"@type":76},"It addresses the assumption that a monitor’s effectiveness is invariant to which untrusted policy produced the actions, noting that deployment shifts the model lineage over time.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the method test whether calibration overfit drives cross-family transfer failure?",{"text":80,"@type":76},"It treats the untrusted policy family as the controlled axis, fits monitors on one family’s transcripts, applies them to other families, and decomposes cross-family AUROC into obviousness, capability, and an interaction term.",{"name":82,"@type":73,"acceptedAnswer":83},"What are the main results regarding cross-family sabotage detection?",{"text":84,"@type":76},"The interaction term is positive and significant on code-backdoor transcripts, and when converted into a defer-to-trusted protocol under a 1% audit budget, an off-lineage monitor detects substantially fewer attack attempts than an in-lineage monitor at the same cost.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]