[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83676-en":3,"doc-seo-83676-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83676,4810365810221,"Aurora","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Not All Refusals Are Equal How Safety Alignment Fails Cybersecurity at Scale","Safety alignment is central to LLM training, yet it often treats different harm domains as conceptually indistinguishable, creating complications for cybersecurity settings that require legitimate, authorized operations. This work reports large-scale abliteration experiments across 24 open-source LLMs, demonstrating domain-specific removal of refusal behavior using standard methods on a 1T-parameter model. Refusal is distributed broadly across layers, especially in trillion-parameter MoE architectures, enabling capture of cybersecurity-relevant harmful concepts. Results show architecture and safety-training type as the most reliable predictors, with models grouped into three susceptibility tiers and explanations proposed for observed intervention effects.","Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale  \nVadym Hadetskyi  \nCracken AI [vh@cracken.ai](vh@cracken.ai)  \nDario Pasquini  \nCracken AI  \n[dp@cracken.ai](dp@cracken.ai)  \nArtem Sorokin  \nCracken AI [as@cracken.ai](as@cracken.ai)  \narXiv :2607 .027 14v2 [ cs .CR] 7 Jul 2026  \nAbstract  \nThere is no doubt that safety alignment is an essential step in LLM training.  \nHowever, conceptually it does not distinguish between various domains and the level of potential harm of a query, which creates significant complications in the fields like cyber security, where a model should not be constrained by its safety circuits to accomplish the goals of legitimate, authorized operations. In this work, we share our findings from a large scale abliteration experiment on 24 open-source LLMs and show that domain-specific abliteration is achievable with standard methodology on the example of a 1T-parameter Kimi K2 . Building on recent work showing that refusal in LLMs occupies a multi-dimensional subspace within layers, we find that it is also distributed widely across layers, especially in trillion-parameter MoE architectures, and so we aim to capture the part of it that represents harmful concepts in the cybersecurity domain exclusively. We also investigate the correlation between models’ features and the effect of domainspecific abliteration, identifying that the type of safety training and architecture are the most reliable predictors. Finally, we classify the models into 3 abliteration susceptibility tiers and put forward a set of conjectures as to why a particular effect from this intervention might be observed in a given model.  \n1 Introduction  \n1.1 Problem Statement & Motivation  \nIn this work, we study domain-specific abliteration. Abliteration is the modification of model weights via orthogonal projection to remove a target direction from the model activation space; typically the refusal direction Arditi et al. [2024], Labonne [2024] . Domain-specific abliteration is the approach of removing the refusal direction of a model for a chosen domain (e.g., cybersecurity), while preserving it for the general case.  \nTo the best of our knowledge, we achieved the first domain-specific abliterated model, removing refusal in the Kimi K2 Kimi Team [2025] for the cybersecurity domain only. However, in our experiments across different models, we observed that not all LLMs respond equally to domain-specific abliteration. We therefore set out to investigate whether this behavior is an inherent model property or a consequence of factors such as architecture, size, training methodology, or other incidental characteristics. More specifically, we ask: does the ability to distinguish between harmful requests in cybersecurity and those involving violence or biological weapons emerge only in sufficiently capable models, or can it also be observed in smaller, potentially simpler models?  \nRecent geometric findings that refusal in LLMs occupies a multi-dimensional subspace rather than a single direction Wollschläger et al. [2025], Pan et al. [2025] support this conjecture: if the refusal subspace is multi-dimensional, then in principle one slice of it, representing a single harm domain, should be removable without collapsing the rest. Whether the standard abliteration pipeline could  \nPreprint.  \nKimi-K2  \nDS-V3  \nLlama-Maverick  \nGLM-4 .7  \nLlama-Scout  \nQwen3-235B Mistral-Large Llama-70B  \nQwen3-32B  \nQwen3-30B  \nGemma-27B  \nMistral-Small  \nDS-V2-Lite  \nQwen3-14B  \nPhi-4  \nGemma-12B  \nQwen3-8B  \nLlama-8B  \nOLMoE  \nPhi-4-mini  \nLlama-3B  \nLlama-1B  \nGemma-1B  \nQwen3-0 .6B  \nRefusal at 30% Abliteration (%) Change from Baseline (pp)  \n88  \n86  \n66  \n100  \n20  \n82  \n90  \n64  \n62  \n78  \n80  \n66  \n84  \n84  \n94  \n76  \n24  \n68  \n68  \n58  \n58  \n66  \n42  \n2  \n40  \n8  \n12  \n42  \n10  \n46  \n20  \n8  \n14  \n18  \n14  \n20  \n26  \n12  \n56  \n24  \n14  \n10  \n6  \n28  \n6  \n8  \n0  \n0  \n51  \n36  \n22  \n89  \n8  \n36  \n51  \n15  \n24  \n36  \n19  \n45  \n38  \n33","cbCailphP67qeEoB","https://ap.wps.com/l/cbCailphP67qeEoB","pdf",856079,2,1,14,"English","en",105,"# Introduction\n## Problem Statement & Motivation","[{\"question\":\"What problem does the paper identify with current safety alignment in LLMs?\",\"answer\":\"Safety alignment does not conceptually distinguish among harm domains or the severity associated with different queries, which complicates cybersecurity use cases requiring legitimate and authorized actions.\"},{\"question\":\"How does the paper define domain-specific abliteration?\",\"answer\":\"Domain-specific abliteration modifies model weights via orthogonal projection to remove refusal directions for a chosen domain (e.g., cybersecurity) while preserving refusal behavior in the general case.\"},{\"question\":\"Which factors best predict how strongly a model responds to domain-specific abliteration?\",\"answer\":\"The study finds that the type of safety training and the model architecture are the most reliable predictors of abliteration effect.\"}]",1784189679,35,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"not-all-refusals-are-equal-how-safety-alignment-fails-cybersecurity-at-scale","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/not-all-refusals-are-equal-how-safety-alignment-fails-cybersecurity-at-scale/83676/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper identify with current safety alignment in LLMs?","Question",{"text":75,"@type":76},"Safety alignment does not conceptually distinguish among harm domains or the severity associated with different queries, which complicates cybersecurity use cases requiring legitimate and authorized actions.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the paper define domain-specific abliteration?",{"text":80,"@type":76},"Domain-specific abliteration modifies model weights via orthogonal projection to remove refusal directions for a chosen domain (e.g., cybersecurity) while preserving refusal behavior in the general case.",{"name":82,"@type":73,"acceptedAnswer":83},"Which factors best predict how strongly a model responds to domain-specific abliteration?",{"text":84,"@type":76},"The study finds that the type of safety training and the model architecture are the most reliable predictors of abliteration effect.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]