[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81778-en":3,"doc-seo-81778-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81778,549758252649,"Ivy","https://ap-avatar.wpscdn.com/avatar/8000253669c5317157?_k=1778319167496531819",8,"Research & Report","HARC Coupling Harmfulness and Refusal Directions for Robust Safety Alignment","Aligned LLMs encode safety in internal representations, and diagnosing alignment failures requires understanding how harmfulness and refusal are represented. Prior work links these concepts to separable directions in the residual stream at prompt positions. This study shows jailbreaks work by suppressing harmfulness and/or refusal before generation, and that during response generation the model recognizes harmful content yet may not translate it into refusal. It introduces HARC, a fine-tuning method coupling both directions across prompt and response positions within a harmfulness–refusal subspace, achieving strong robustness without degrading general capability or increasing over-refusal.","arXiv :2607 .00572v 3 [ cs .AI] 8 Jul 2026  \nHARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment  \nShei Pern Chua 1 ,3 Hao Wu2 ,3 Qianli Ma3 Fangzhao Wu3 ∗  \n1Tsinghua University 2 Southeast University 3Microsoft  \n[hi@sheichua.com](hi@sheichua.com) [fangzwu@microsoft.com](fangzwu@microsoft.com)  \nAbstract  \nUnderstanding how aligned LLMs internally represent safety is critical for diagnosing alignment vulnerabilities, as it explains why jailbreaks succeed and informs the design of robust alignment strategies. Prior work shows that aligned LLMs encode harmfulness and refusal as separable directions in the residual stream at prompt-side token positions. We show that jailbreaks succeed at prompt encoding by suppressing either the refusal or harmfulness direction before any token is generated, with distinct attack classes occupying separable regions of the harmfulnessrefusal plane. Extending the analysis to response-token positions, we find that the model recognizes harmful content while it is generating that content, even when it failed to recognize the input as harmful at the prompt side. Motivated by our findings, we introduce HARC (Harmfulness-And-Refusal Coupling), a fine-tuning method that pairs the two directions across both prompt and response positions.  \nSince the intervention is confined to the harmfulness-refusal subspace, it leaves the rest of the residual stream intact and does not degrade general capability or inflate over-refusal. Across extensive experiments, HARC achieves the strongest robustness-capability-usability trade-off among six baselines spanning the major training-time and inference-time safety methods. The harmfulness and refusal directions at prompt and response positions transfer across the five model families and two scales we tested without architecture-specific tuning. 2  \n1 Introduction  \nAligned large language models refuse harmful requests under direct prompting, but adversarial attacks such as adversarial suffixes [1, 2], persuasion-style rewrites [3, 4], iterative red-teaming [5, 6], multi-turn conversations [7–9], and obfuscation [10, 11] continue to bypass alignment on frontier models [12] . In response, a range of approaches have been proposed, including preference optimization [13], supervised refusal training [14], deliberative reasoning [15, 16], inference-time steering [17, 18], and representation-level interventions that reshape activations during training [19– 22] . Although these methods improve robustness, they offer little insight into the internal mechanisms by which LLMs encode safety, or why particular adversarial attacks succeed in bypassing these defenses.  \nRecent interpretability work begins to close this gap by showing that aligned models encode harmfulness (vharm ) and refusal (vref ) as distinct directions in the residual stream at prompt-side token positions [23, 24] . We therefore ask a sharper question: do jailbreaks succeed by exploiting this structure, and if so, how? By probing three attacks spanning distinct mechanism families, we find  \nthat successful jailbreaks suppress either direction, or both. When harmfulness is suppressed before ∗Corresponding author.  \n2 The code is available at [https://github.com/microsoft/HARC](https://github.com/microsoft/HARC)  \nPreprint.  \ngeneration, the model registers the jailbreak prompt as not harmful and shows no refusal intent. Yet it proceeds to produce a harmful response. This exposes a limitation of prompt-side analysis alone: does the model know what it is generating?  \nTo answer this, we extend the representational analysis to response-token positions by extracting harmfulness and refusal directions from residuals during generation (Section 3) . We find that the model recognizes harmful content while it is generating that content, even when it failed to recognize the input as harmful during prompt encoding. It knows what it is producing but fails to translate that knowledge into refusal. We fur","cbCaidU32KWQtJDt","https://ap.wps.com/l/cbCaidU32KWQtJDt","pdf",657397,4,1,27,"English","en",105,"# Introduction\n## Internal representations of safety in aligned LLMs\n## Prompt-side vs response-side harmfulness recognition\n## HARC: Harmfulness-And-Refusal Coupling\n## Experimental contributions and results","[{\"question\":\"How do jailbreaks succeed according to this paper’s representational analysis?\",\"answer\":\"Successful jailbreaks suppress the refusal direction and/or the harmfulness direction during prompt encoding before any token is generated, with different attack classes occupying separable regions of the harmfulness–refusal space.\"},{\"question\":\"What does the model’s behavior show when analyzing response-token positions?\",\"answer\":\"During response generation, the model can recognize harmful content even if it failed to recognize the input as harmful at the prompt side, indicating a gap between harmful recognition and refusal behavior.\"},{\"question\":\"What is HARC and how does it improve robust safety alignment?\",\"answer\":\"HARC couples the harmfulness and refusal directions at both prompt and response positions using representation-level fine-tuning confined to a harmfulness–refusal subspace. It yields a stronger robustness-capability-usability trade-off than multiple baseline safety methods while avoiding general capability degradation and minimizing over-refusal.\"}]",1784176088,68,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"harc-coupling-harmfulness-and-refusal-directions-for-robust-safety-alignment","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/harc-coupling-harmfulness-and-refusal-directions-for-robust-safety-alignment/81778/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"How do jailbreaks succeed according to this paper’s representational analysis?","Question",{"text":75,"@type":76},"Successful jailbreaks suppress the refusal direction and/or the harmfulness direction during prompt encoding before any token is generated, with different attack classes occupying separable regions of the harmfulness–refusal space.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What does the model’s behavior show when analyzing response-token positions?",{"text":80,"@type":76},"During response generation, the model can recognize harmful content even if it failed to recognize the input as harmful at the prompt side, indicating a gap between harmful recognition and refusal behavior.",{"name":82,"@type":73,"acceptedAnswer":83},"What is HARC and how does it improve robust safety alignment?",{"text":84,"@type":76},"HARC couples the harmfulness and refusal directions at both prompt and response positions using representation-level fine-tuning confined to a harmfulness–refusal subspace. It yields a stronger robustness-capability-usability trade-off than multiple baseline safety methods while avoiding general capability degradation and minimizing over-refusal.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]