[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85332-en":3,"doc-seo-85332-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85332,13056703020460,"Valentina","https://ap-avatar.wpscdn.com/avatar/be000253dac470eee5d?_k=1778207105932848923",8,"Research & Report","HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models","Safety alignment in large language models can weaken during fine-tuning, because even benign task adaptation may increase harmful compliance. Existing defenses are costly retraining/weight edits or model-agnostic safety classifiers that can miss checkpoint-specific failures. HyperSafe proposes a post-hoc, model-specific, non-invasive safety restoration framework. It generates a Safe Side Network per fine-tuned checkpoint using layer-wise activation fingerprints, enabling prompt-level routing to refusal for harmful inputs and original answering for safe inputs. Evaluations on Qwen2-7B and LLaMA-3-8B reduce harmful rates from 19–31% to below 1% while preserving average utility within 1% of baselines.","arXiv :2607 . 1 1475v 1 [ cs .LG] 13 Jul 2026  \nHyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models  \nAznaur Aliev Carlos Hinojosa Abdelrahman Eldesokey Bang An Bernard Ghanem Yibo Yang  \nKing Abdullah University of Science and Technology, Saudi Arabia  \n{aznaur.aliev,carlos.hinojosa,[yibo.yang}@kaust.edu.sa](yibo.yang}@kaust.edu.sa)  \nAbstract  \nSafety alignment in large language models can be fragile under fine-tuning, as even benign task adaptation may increase harmful compliance. Existing defenses mainly follow two directions: they either intervene during or after fine-tuning through retraining or weight modification, which can be costly and may hurt task performance, or they use model-agnostic safety classifiers, which may miss failures specific to a given fine-tuned checkpoint. These limitations motivate a post-hoc, model-specific, and non-invasive approach to safety restoration. To meet these requirements, we propose HyperSafe, a framework that restores safety behavior by generating a model-specific Safe Side Network (SSN) for each finetuned checkpoint. HyperSafe uses layer-wise activation fingerprints to capture how fine-tuning changes the model’s inner representations. With a small set of given calibration prompts, the hypernetwork maps these fingerprints to the parameters of the SSN in a single forward pass. The generated SSN runs alongside the frozen fine-tuned model and performs prompt-level safety classification: harmful prompts are routed to refusal, while safe prompts are answered by the original fine-tuned model. Thus, HyperSafe requires no gradient updates, no safety data at deployment time, and no modification to the deployed model weights. We evaluate HyperSafe on two model families, Qwen2-7B and LLaMA-3-8B, across multiple safety benchmarks. HyperSafe reduces harmful response rates from 19– 31% to below 1% on every held-out checkpoint, while keeping downstream task accuracy within 1% of the fine-tuned baseline on average. Code is available at  \n[https://github.com/nokronim/project-safety-remedy](https://github.com/nokronim/project-safety-remedy)  \n1 Introduction  \nFine-tuning safety-aligned LLMs is a common way to adapt them to downstream tasks, but it can also weaken their safety alignment. Prior work has shown that, even when the fine-tuning data is totally benign, the harmful-response rate of a LoRA fine-tuned model can climb from under 5%(the aligned baseline) to over 40% on standard safety benchmarks [1] . One possible reason is that safety alignment can depend heavily on the model’s behavior in the first few output tokens [2] . As a result, even benign fine-tuning updates may change the model’s refusal behavior and reduce its ability to reject harmful requests.  \nExisting defenses for this alignment degradation can be grouped into three main categories. The first category applies safety-preserving methods during fine-tuning. These methods add safety constraints during task adaptation, for example, through regularization, safety data, or constrained optimization (Lisa [3], RepNoise [4], Constrained-SFT [2]) . They aim to preserve the base model’s refusal behavior while learning the downstream task. However, they often introduce a trade-off between safety and utility: stronger safety constraints may reduce task performance, while weaker constraints may still  \nPreprint.  \nTable 1: Comparison with existing defenses against fine-tuning-induced safety degradation. Setup cost per checkpoint: compute to deploy on each new fine-tuned model. Model-aware: tailored to the target rather than a one-size-fits-all classifier. Preserves model weights: deployed weights kept unchanged. Standalone at deploy: requires only the fine-tuned model at inference, no aligned base M0. HyperSafe is the only approach that satisfies all four.  \n\n| Method | Setup cost\u003Cbr>per checkpoint | Model\u003Cbr>aware | Preserves\u003Cbr>model weights | Standalone\u003Cbr>at deploy |\n| --- | --- | --- | --- | --- |\n| Vaccine [11] | full ","cbCaihVjNyf06mLN","https://ap.wps.com/l/cbCaihVjNyf06mLN","pdf",4726360,3,1,17,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"Why can fine-tuning weaken safety alignment in large language models?\",\"answer\":\"Safety alignment may be fragile: even benign downstream adaptation can increase harmful compliance by changing refusal behavior, especially in early output tokens.\"},{\"question\":\"What are the main limitations of existing defenses against fine-tuning-induced safety degradation?\",\"answer\":\"They either intervene during/after fine-tuning via costly retraining or weight modifications, or they rely on model-agnostic external safety classifiers that may miss failures specific to a given fine-tuned checkpoint.\"},{\"question\":\"How does HyperSafe restore safety during inference without modifying model weights?\",\"answer\":\"HyperSafe generates a model-specific Safe Side Network using layer-wise activation fingerprints mapped from a small set of calibration prompts via a hypernetwork. The frozen fine-tuned model remains unchanged; harmful prompts are routed to refusal while safe prompts are answered by the original model.\"}]",1784202553,43,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"hypersafe-inference-time-safety-recovery-for-fine-tuned-language-models","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/hypersafe-inference-time-safety-recovery-for-fine-tuned-language-models/85332/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why can fine-tuning weaken safety alignment in large language models?","Question",{"text":75,"@type":76},"Safety alignment may be fragile: even benign downstream adaptation can increase harmful compliance by changing refusal behavior, especially in early output tokens.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What are the main limitations of existing defenses against fine-tuning-induced safety degradation?",{"text":80,"@type":76},"They either intervene during/after fine-tuning via costly retraining or weight modifications, or they rely on model-agnostic external safety classifiers that may miss failures specific to a given fine-tuned checkpoint.",{"name":82,"@type":73,"acceptedAnswer":83},"How does HyperSafe restore safety during inference without modifying model weights?",{"text":84,"@type":76},"HyperSafe generates a model-specific Safe Side Network using layer-wise activation fingerprints mapped from a small set of calibration prompts via a hypernetwork. The frozen fine-tuned model remains unchanged; harmful prompts are routed to refusal while safe prompts are answered by the original model.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]