[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82077-en":3,"doc-seo-82077-105":29,"detail-sidebar-cat-0-en-105":94},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":11,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82077,13056703019404,"Miles","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Optimizing Against Safety Representations: Activation-Guided Adversarial Suffixes and the Geometry of Refusal","Behavioral alignment in large language models often conceals fragile internal safety representations. Recent findings indicate refusal behavior is influenced by low-dimensional directions within activation space, prompting questions about their structure, locality, and accessibility to optimization. This work uses adversarial suffix attacks as a diagnostic probe and proposes Activation-Guided GCG, which optimizes internal refusal-direction projections rather than output-token objectives. Results show distributed refusal encoding across the forward pass, and Soft-GCG yields faster, stronger attacks while experiments assess scale-dependent vulnerability and resistance.","Optimizing Against Safety Representations: Activation-Guided Adversarial  \nSuffixes and the Geometry of Refusal  \nEge Cakar, Hannah Guan, Kayden Kehe  \nHarvard University  \nJohn A. Paulson School of Engineering and Applied Sciences  \n{ecakar, hguan, [kaydenkehe](kaydenkehe}@college.harvard.edu)[}](kaydenkehe}@college.harvard.edu)[@college.harvard.edu](kaydenkehe}@college.harvard.edu)  \narXiv :2607 .08883v 1 [ cs .LG] 9 Jul 2026  \nAbstract  \nBehavioral alignment in large language models often masks fragile internal safety representations. Recent work suggests that refusal behavior is mediated by low-dimensional directions in activation space. This raises questions about how such representations are structured, localized, and accessed by optimization. We study adversarial suffix attacks as a probe of representational alignment. We introduce Activation-Guided GCG, which replaces output-based objectives with losses that directly target a model’s internal refusal direction. Across several objective variants, we find that suppressing refusal globally across all layers and positions is more effective than targeting a single layer–position pair. This suggests that safety representations are distributed across the forward pass rather than causally localized toa single site. We further introduce Soft-GCG, a continuous relaxation of discrete suffix optimization using GumbelSoftmax. Soft-GCG achieves a 33 × speedup over standard GCG while improving attack success rates. Evaluating across model scales, we find that smaller models remain vulnerable while larger models resist both activation- and suffix-based attacks at our compute-constrained settings, consistent with larger and better safety trained models being harder to jailbreak. Together, our results clarify how safety mechanisms are encoded and can be broken in contemporary models. These insights provide concrete guidance for designing more robust and representation-aware alignment strategies.  \nCode—[github.com/Ege-Cakar/ImprovingGCG](github.com/Ege-Cakar/ImprovingGCG)  \nIntroduction  \nThe widespread deployment of Large Language Models (LLMs) has necessitated rigorous safety alignment techniques to prevent the generation of harmful or unethical content. At least in public chat interfaces, it seems like these alignment techniques have advanced significantly. However, recent research has demonstrated that these safety guardrails are surprisingly brittle. Two distinct lines of inquiry have emerged to analyze this fragility: automated adversarial attacks on the input space, and the analysis of refusal representations within the model’s residual streams.  \nThe Greedy Coordinate Gradient (GCG) method, proposed by (Zou et al. 2023), demonstrates that optimizing adversarial suffixes can reliably bypass alignment. However, GCG operates on a behavioral proxy: it optimizes inputs  \nto maximize the likelihood of specific target output tokens (e.g., forcing the model to output “Sure, here is...”) . While effective, this objective targets the symptom of the jailbreak (the output logits) rather than the mechanism of the refusal. Additionally, the GCG method is limited by the given initialization conditions and long waits until convergence. Its discrete optimization methods are significantly slower compared to continuous methods, and we hypothesize that a large portion of the optimization that is done by GCG can be done first via continuous optimization, and that discrete optimization is only necessary to overcome the “projection gap”(i.e., the loss in performance that comes from jumping to discrete tokens at the end after optimizing continuously and achieving a minimum in between real tokens that doesn’t generalize to hard tokens) .  \nConversely,(Arditi et al. 2024) investigates the model’s internal representations of refusal. They demonstrated that refusal behavior in open-source chat models is mediated by a single, one-dimensional subspace in the model’s activation space and that simply ablating ","cbCaioz0JcoTZ4xD","https://ap.wps.com/l/cbCaioz0JcoTZ4xD","pdf",499551,2,1,"English","en",105,"# Abstract\n# Introduction\n# Related Work\n## Adversarial Attacks on Aligned LLMs","[{\"question\":\"What problem does the paper address in LLM safety alignment?\",\"answer\":\"It examines how safety alignment can appear effective behaviorally while remaining brittle due to fragile internal refusal representations. The work studies how those internal mechanisms are encoded and breakable under optimization.\"},{\"question\":\"How does Activation-Guided GCG differ from standard GCG?\",\"answer\":\"Activation-Guided GCG replaces output-token-based jailbreak objectives with losses that directly minimize projections onto model internal refusal directions. This targets the mechanism rather than only suppressing the output symptom.\"},{\"question\":\"What does the paper conclude about where refusal representations are encoded?\",\"answer\":\"Suppressing refusal globally across all layers and positions is more effective than targeting a single layer–position pair. This suggests refusal representations are distributed across the forward pass rather than localized to one causal site.\"},{\"question\":\"What is Soft-GCG and what are its reported benefits?\",\"answer\":\"Soft-GCG relaxes discrete suffix optimization into a continuous formulation using Gumbel-Softmax. It provides a reported 33× speedup over standard GCG while improving attack success rates and helps analyze vulnerability across model sizes.\"}]",1784178090,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":89,"head_meta":91,"extra_data":93,"updated_unix":27},"optimizing-against-safety-representations-activation-guided-adversarial-suffixes-and-the-geometry-of-refusal","",{"@graph":35,"@context":88},[36,52,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,46,49],{"item":40,"name":41,"@type":42,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":20},"https://docshare.wps.com/document/","Document",{"item":47,"name":12,"@type":42,"position":48},"https://docshare.wps.com/document/research-report/",3,{"item":50,"name":13,"@type":42,"position":51},"https://docshare.wps.com/document/optimizing-against-safety-representations-activation-guided-adversarial-suffixes-and-the-geometry-of-refusal/82077/",4,{"url":50,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":23,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":40,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-20","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80,84],{"name":71,"@type":72,"acceptedAnswer":73},"What problem does the paper address in LLM safety alignment?","Question",{"text":74,"@type":75},"It examines how safety alignment can appear effective behaviorally while remaining brittle due to fragile internal refusal representations. The work studies how those internal mechanisms are encoded and breakable under optimization.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How does Activation-Guided GCG differ from standard GCG?",{"text":79,"@type":75},"Activation-Guided GCG replaces output-token-based jailbreak objectives with losses that directly minimize projections onto model internal refusal directions. This targets the mechanism rather than only suppressing the output symptom.",{"name":81,"@type":72,"acceptedAnswer":82},"What does the paper conclude about where refusal representations are encoded?",{"text":83,"@type":75},"Suppressing refusal globally across all layers and positions is more effective than targeting a single layer–position pair. This suggests refusal representations are distributed across the forward pass rather than localized to one causal site.",{"name":85,"@type":72,"acceptedAnswer":86},"What is Soft-GCG and what are its reported benefits?",{"text":87,"@type":75},"Soft-GCG relaxes discrete suffix optimization into a continuous formulation using Gumbel-Softmax. It provides a reported 33× speedup over standard GCG while improving attack success rates and helps analyze vulnerability across model sizes.","https://schema.org",{"og:url":50,"og:type":90,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":92,"canonical":50},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":95},[96,100,104,108,113,118,123,126,130,133,137],{"id":21,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Literature",80,"literature",{"id":51,"doc_module":4,"doc_module_name":45,"category_name":105,"show_sort_weight":106,"slug":107},"Exam",70,"exam",{"id":109,"doc_module":4,"doc_module_name":45,"category_name":110,"show_sort_weight":111,"slug":112},5,"Comic",60,"comic",{"id":114,"doc_module":4,"doc_module_name":45,"category_name":115,"show_sort_weight":116,"slug":117},6,"Technology",50,"technology",{"id":119,"doc_module":4,"doc_module_name":45,"category_name":120,"show_sort_weight":121,"slug":122},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":124,"slug":125},30,"research-report",{"id":127,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":28,"slug":129},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":28,"slug":132},"World Cup","world-cup",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":134,"slug":136},10,"Lifestyle","lifestyle",{"id":138,"doc_module":4,"doc_module_name":45,"category_name":139,"show_sort_weight":109,"slug":140},19,"General","general"]