[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85073-en":3,"doc-seo-85073-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85073,1099514067438,"River Wang","https://ap-avatar.wpscdn.com/avatar/100002539ee87300030?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780474512215547542",8,"Research & Report","Certified Interventional Fidelity Anytime-Valid Adaptive Evaluation of Causal Claims in Mechanistic Interpretability","Mechanistic interpretability typically checks explanation faithfulness by interventions such as swapping hidden states, patching activations, ablating components, or comparing compressed models to originals. These studies often report a single point estimate, even when evaluation is sequential, monitored, or adapted to suspected failures, making stability of “fidelity” or patching claims unclear under finite sampling. Certified Interventional Fidelity (CIF) converts each reported score into a causal estimand and provides confidence intervals and anytime-valid confidence sequences, including adaptive sampling via bounded mixture importance weighting. Hoeffding-style and variance-adaptive betting sequences certify high-fidelity claims while clarifying when differences lack statistical support.","Certified Interventional Fidelity:  \nAnytime-Valid, Adaptive Evaluation of Causal Claims in Mechanistic  \nInterpretability  \nAmir Asiaee 1  \n1Department of Biostatistics, Vanderbilt University Medical Center, Nashville, TN 37232, USA  \narXiv :2607 .08349v 1 [ cs .LG] 9 Jul 2026  \nAbstract  \nMechanistic interpretability often evaluates explanations by intervening on a model: swapping hidden states, patching activations, ablating components, or comparing a compressed model to the original one. These experiments are usually summarized by a point estimate, even though the evaluation may be monitored while it runs or adapted toward suspected failures. This makes it hard to tell whether a reported fidelity or patching effect is a stable causal claim or a consequence of finite sampling and evaluation choices.  \nWe introduce Certified Interventional Fidelity (CIF), a statistical layer for interventional interpretability evaluations. CIF first writes the quantity being reported as a causal estimand: an expectation of a bounded score over a stated input distribution and a stated intervention distribution. It then provides confidence intervals and anytime-valid confidence sequences for this estimand, including under adaptive intervention sampling via bounded mixture importance weighting. We instantiate CIF with Hoeffding-style sequences and variance-adaptive betting sequences, the latter reducing certification cost by 10–30× in our experiments. On MNIST abstractions and GPT-2 Small IOI circuits, CIF certifies high-fidelity claims, shows when apparent method differences are not statistically supported, and makes sensitivity to the intervention distribution explicit.  \n1 INTRODUCTION  \nDeep neural networks can match or exceed human performance, but understanding how they compute remains challenging. Mechanistic interpretability aims to provide explanations that are faithful simplifications of the internal  \ncomputation of a trained model [Olah et al., 2020, Geiger et al., 2025] . A central move in recent interpretability has been the shift from observational probes to interventional evaluations: we modify internal states, paths, or mechanismsand ask whether the resulting behavior supports a proposed mechanistic explanation.  \nThe explanations being evaluated usually fall into two broad classes. The first class consists of abstraction or reduction fidelity claims: a simpler object, such as a high-level causal model, compressed network, pruned network, or circuit, is claimed to preserve the relevant causal behavior of the original model. Causal abstraction and interchange intervention accuracy (IIA) are canonical examples [Geiger et al., 2021, 2025]; recent work on mechanism transformations places structured neural compression in the same interventionalfidelity setting [Asiaee, 2026] . The second class consists of component-effect claims: a set of heads, neurons, paths, or edges is claimed to causally support a behavior. Activation patching, path patching, causal tracing, circuit ablation, and many circuit-discovery evaluations belong to this class [Meng et al., 2022, Goldowsky-Dill et al., 2023, Zhang and Nanda, 2024] . Causal scrubbing and circuit-completeness evaluations sit near the boundary: depending on the reported score, they can be viewed either as testing a reduced explanation’s fidelity or as measuring the effect of selected components [Chan et al., 2022] .  \nThe missing piece: statistical validity. Despite the causal framing, evaluation practice is often informal. A typical workflow is: sample a few thousand input pairs/prompt pairs; sample a few thousand interventions; report a single scalar metric. But interpretability workflows are inherently sequential and adaptive: we monitor results, adjust evaluation sets, and hunt for counterexamples. Without uncertainty quantification and without accounting for adaptivity, it is easy to overstate fidelity, mis-rank methods, or miss rare but important failure modes.  \nWe propose Certified","cbCaio2oHHzciQzj","https://ap.wps.com/l/cbCaio2oHHzciQzj","pdf",453414,2,1,18,"English","en",105,"# Abstract\n# Introduction\n## Background: interventional interpretability\n## Two classes of mechanistic claims\n## Missing piece: statistical validity\n## Certified Interventional Fidelity (CIF) design\n## Positioning relative to prior work\n## Contributions","[{\"question\":\"Why can point-estimate reporting in mechanistic interpretability be misleading?\",\"answer\":\"Because evaluation is often sequential and adaptive, repeated monitoring and finite sampling can turn apparent fidelity or patching effects into unstable claims without uncertainty quantification.\"},{\"question\":\"What does CIF change about how interventional interpretability results are reported?\",\"answer\":\"CIF rewrites the reported score as a causal estimand—an expectation over a specified input distribution and intervention distribution—so the scientific claim is explicit and separated from sampling.\"},{\"question\":\"How does CIF remain valid under adaptive or anytime evaluation?\",\"answer\":\"CIF uses anytime-valid confidence sequences that stay valid under repeated checking or early stopping, and it supports adaptive intervention sampling via bounded mixture importance weighting.\"}]",1784200843,45,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"certified-interventional-fidelity-anytime-valid-adaptive-evaluation-of-causal-claims-in-mechanistic-interpretability","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/certified-interventional-fidelity-anytime-valid-adaptive-evaluation-of-causal-claims-in-mechanistic-interpretability/85073/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why can point-estimate reporting in mechanistic interpretability be misleading?","Question",{"text":75,"@type":76},"Because evaluation is often sequential and adaptive, repeated monitoring and finite sampling can turn apparent fidelity or patching effects into unstable claims without uncertainty quantification.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What does CIF change about how interventional interpretability results are reported?",{"text":80,"@type":76},"CIF rewrites the reported score as a causal estimand—an expectation over a specified input distribution and intervention distribution—so the scientific claim is explicit and separated from sampling.",{"name":82,"@type":73,"acceptedAnswer":83},"How does CIF remain valid under adaptive or anytime evaluation?",{"text":84,"@type":76},"CIF uses anytime-valid confidence sequences that stay valid under repeated checking or early stopping, and it supports adaptive intervention sampling via bounded mixture importance weighting.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]