[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86475-en":3,"doc-seo-86475-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86475,1099514068035,"Ezra","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","When Are Sparse Feature Interventions Actually Localized Matched Evaluation for SAE-Based Safety Control","Sparse autoencoder (SAE) features are used as localized control handles for safety-critical language model behaviors, promising a more efficient behavior-per-perturbation effect than dense activation steering. This advantage depends strongly on baseline matching. A Matched Coherence-Gated (MCG) evaluation protocol compares sparse and dense interventions via complementary matched controls, coherence-gated unsafe harmful compliance, and multi-judge cross-checks. Results across Gemma and Llama variants show efficiency can disappear or reverse under matched-surface conditions, with further artifacts from incoherence and single-judge inflation.","arXiv :2607 . 10226v 1 [ cs .AI] 11 Jul 2026  \nWhen Are Sparse Feature Interventions Actually Localized? Matched Evaluation for SAE-Based Safety Control  \nDaming Luo Christy Liang Junyu Xuan  \nUniversity of Technology Sydney  \n[daming. luo@student. uts. edu. au](daming. luo@student. uts. edu. au)  \n[Jie. Liang@uts. edu. au](Jie. Liang@uts. edu. au) [Junyu. Xuan@uts. edu. au](Junyu. Xuan@uts. edu. au)  \nAbstract  \nSparse autoencoder (SAE) features are increasingly proposed as localized control handles for safety-relevant behavior, on the grounds that a sparse feature intervention changes behavior more efficiently—more behavior per unit of internal perturbation—than dense activation steering. We show that whether this advantage holds depends sharply on how the dense baseline is matched. We introduce Matched Coherence-Gated (MCG) Evaluation, a protocol that brackets a sparseversus-dense comparison with two complementary controls—matched target effect (fix behavior, read off perturbation) and matched perturbation norm (fix perturbation, read off behavior)—counts a jailbreak only when an output is both judge-unsafe and coherent, and cross-checks every headline number with a second behavior-completion judge. Across Gemma-2-2B/9B/27B-it and Llama-3.1-8B-Instruct with two SAE suites (Gemma Scope and Llama Scope), the protocol exposes that matching the total perturbation norm is not sufficient: it leaves the intervention surface unmatched, comparing a single-layer SAE ablation against an all-layer dense direction whose perturbation is diluted across the stack. On Gemma-2-9B, once we also match the surface—same-layer dense steering and dense steering projected onto the SAE decoder span—the apparent SAE efficiency advantage disappears and reverses: at every matched perturbation bin both fair baselines elicit more coherent harmful compliance than SAE ablation (up to −0 .29 true-jailbreak), with SAE additionally paying a capability cost at high perturbation, and the reversal holds under a HarmBench second judge. On Llama-3.1-8B and Gemma-2-27B, SAE retains a large advantage over all-layer dense steering. Independently of the dense baseline, the protocol reveals two further artifacts: high-k SAE ablations trip a safety judge with incoherent text (supported by a human audit), and in a 2B model SAE “jailbreaks” are largely singlejudge inflation that a second judge does not corroborate (κ → 0) while reasoning collapses. We conclude that SAE feature ablation is not a uniformly localized safety handle, and that sparselocalization claims should be evaluated with matched-surface, matched-basis, coherence-gated, multi-judge protocols.  \n1 Introduction  \nSparse autoencoders (SAEs) provide an appealing interface for interpreting and intervening on internal representations of language models (Bricken et al. , 2023; Cunningham et al. , 2023; Lieberum et al. , 2024) . If safety-relevant behavior such as refusal, harmful compliance, or jailbreak susceptibility is mediated by a small number of sparse features, then runtime feature interventions could offer a finer-grained alternative to dense activation steering or weight-level adaptation.  \nThis promise raises a measurement problem. A safety intervention can appear successful for reasons that are not genuine localized control. A method may be weaker than the baseline it is  \ncompared against, preserving utility because it barely changes the model. A method may reach the desired target behavior only at a strength that destroys downstream capability. A high-strength intervention may also produce incoherent or degenerate text that an automated judge labels unsafe, inflating target-behavior metrics without producing meaningful harmful compliance. These failure modes matter because the core claim behind feature-level control is not merely that an intervention changes behavior, but that it changes behavior through a small and specific internal mechanism.  \nWe therefore ask: when are SAE feature interventions a","cbCaijg7FimZzWPw","https://ap.wps.com/l/cbCaijg7FimZzWPw","pdf",511972,3,1,22,"English","en",105,"# Introduction\n## Matched Coherence-Gated (MCG) evaluation protocol\n## Baseline matching and localization analysis\n## Results across model scales and SAE suites","[{\"question\":\"What is the main question the paper investigates about sparse feature interventions?\",\"answer\":\"When SAE feature interventions are truly localized, meaning they improve behavior per perturbation through a small internal mechanism rather than measurement artifacts or capability loss.\"},{\"question\":\"How does the MCG Evaluation protocol make sparse-versus-dense comparisons fair?\",\"answer\":\"It brackets sparse and dense interventions with complementary matched controls: matching target effect and matched perturbation norm, using coherence-gated unsafe harmful compliance and cross-checking headline outcomes with a second judge.\"},{\"question\":\"What major finding challenges the claimed efficiency of SAE ablation for safety control?\",\"answer\":\"On Gemma-2-9B, matching only total perturbation norm is insufficient; once the dense baseline is matched by surface (same-layer or SAE-decoder-span projected), the apparent SAE advantage disappears and can reverse, yielding more coherent harmful compliance for the baselines.\"}]",1784211960,55,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"when-are-sparse-feature-interventions-actually-localized-matched-evaluation-for-sae-based-safety-control","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/when-are-sparse-feature-interventions-actually-localized-matched-evaluation-for-sae-based-safety-control/86475/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the main question the paper investigates about sparse feature interventions?","Question",{"text":75,"@type":76},"When SAE feature interventions are truly localized, meaning they improve behavior per perturbation through a small internal mechanism rather than measurement artifacts or capability loss.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the MCG Evaluation protocol make sparse-versus-dense comparisons fair?",{"text":80,"@type":76},"It brackets sparse and dense interventions with complementary matched controls: matching target effect and matched perturbation norm, using coherence-gated unsafe harmful compliance and cross-checking headline outcomes with a second judge.",{"name":82,"@type":73,"acceptedAnswer":83},"What major finding challenges the claimed efficiency of SAE ablation for safety control?",{"text":84,"@type":76},"On Gemma-2-9B, matching only total perturbation norm is insufficient; once the dense baseline is matched by surface (same-layer or SAE-decoder-span projected), the apparent SAE advantage disappears and can reverse, yielding more coherent harmful compliance for the baselines.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]