[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84286-en":3,"doc-seo-84286-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84286,1374391974585,"Genevieve","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Efficient Safety Alignment of Language Models via Latent Personality Traits","Current safety methods for large language models are vulnerable to adversarial attacks, motivating defenses that remain robust without sacrificing utility. Latent Adversarial Training (LAT) is effective but data-intensive and can degrade performance. Latent Personality Alignment (LPA) replaces explicit harm-refusal training with adversarial stabilization of personality-anchored latent representations using only 66 harm-agnostic statements. LPA attains near-zero attack success on HarmBench across direct requests and jailbreak methods, with no harmful-content exposure and no loss on standard benchmarks. Training is lightweight, completing in minutes on a single GPU with 75× fewer examples than standard LAT. Extensive ablations validate robustness, efficiency, and generalization.","arXiv :2607 .079 18v 1 [ cs .LG] 8 Jul 2026  \nEfficient Safety Alignment of Language Models via Latent Personality Traits  \nMohamed Amine Merzouk1,2 Nolan Smyth1,4 Damiano Fornasiere3 Linh Le5 David Williams-King5 Adam Oberman2,3  \n1Mila, Quebec AI Institute 2McGill University 3LawZero 4Universit de  \nMontral 5Independent  \nAbstract  \nCurrent safety methods for large language models are known to be vulnerable to adversarial attacks, motivating research into robust alternatives. Latent Adversarial Training (LAT) is among the most effective defenses (Sheshadri et al., 2025), but can degrade utility and requires training on large datasets of harmful prompts (Yu et al., 2025) . We introduce Latent Personality Alignment (LPA), which replaces explicit harm refusal with adversarial training on just 66 harm-agnostic statements drawn from psychometric personality literature. We hypothesize that personality-anchored representations share latent structure with harm avoidance, so adversarially stabilizing them implicitly constrains the subspace exploited by jailbreak attacks. LPA achieves near-zero attack success rates on HarmBench across direct requests and five jailbreak methods, despite never seeing harmful content during training and no loss of performance on standard benchmarks. Moreover, the training process is lightweight; the entire procedure completes in minutes on a single GPU and uses 75 × fewer examples than standard LAT. Extensive ablations demonstrate the robustness, efficiency, and generalization of our method. We make our code available through ananonymized repository.  \n1 Introduction  \nEnsuring the safety of large language models (LLMs) without degrading their utility remains a major challenge for the machine learning community. Current post-training approaches often rely on explicit supervision over harmful content (Ouyang et al., 2022; Christiano et al., 2023), yet recent work has exposed various failure modes in seemingly aligned models.  \nLLMs are vulnerable to adversarial attacks such as jailbreaks and adversarial prompts (Perez et al., 2022; Zou et al., 2023; Mazeika et al., 2024; Rando et al., 2025; Boreiko et al., 2025; Liet al., 2024) . Furthermore, emergent misalignment shows that even finetuning on seemingly benign data can lead to significantly misaligned models (Betley et al., 2026; Wang et al., 2025a) . Aligned behaviors prove fragile under normal use too: the outputs of LLMs vary substantially under superficial prompt changes (Sclar et al., 2024), the adherence to system prompts degrades over a few interactions (Salinas & Morstatter, 2024; Qin et al., 2024), and the personality traits easily shift across contexts (Jiang et al., 2022; Pellert et al., 2023; Serapio-Garca et al., 2025; Gupta et al., 2024; Tosato et al., 2025) .  \nAddressing these vulnerabilities without degrading the utility of models is a critical technical problem, although some promising directions exist. Adversarial training methods operate in latent space to suppress harmful behavior more robustly (Sheshadri et al., 2025; Casper et al., 2025; Xhonneux et al., 2024) . However, they are data-intensive and prone to overfitting to specific classes of harm (Jain et al., 2024); this can also degrade performance on benign tasks (Cui et al., 2025; Panda et al., 2024) .  \nA second approach, activation steering (Chen et al., 2025; Lu et al., 2026), identifies approximately linear directions in activation space corresponding to a helpful assistant persona and intervenes at inference time by either steering or capping activations to prevent persona  \ndrift. This reduces attack success rates and can steer conversations away from harmful content. However, steering is not fully robust: while reducing vulnerability by a factor of two on persona-based jailbreaks, a large percentage of attacks still succeed (Lu et al., 2026) .  \nIn this work, we propose Latent Personality Alignment (LPA), a compute-efficient posttraining method that replaces explicit harm","cbCairq2qsvkji95","https://ap.wps.com/l/cbCairq2qsvkji95","pdf",369893,4,1,15,"English","en",105,"# Introduction\n## Background on LLM safety vulnerabilities\n## Related approaches and limitations\n## Proposed method: Latent Personality Alignment (LPA)\n## Experimental results and contributions","[{\"question\":\"What problem does Latent Personality Alignment (LPA) address in LLM safety?\",\"answer\":\"LPA targets the vulnerability of large language models to adversarial jailbreaks and adversarial prompts, aiming to improve safety while avoiding utility degradation.\"},{\"question\":\"How does LPA differ from Latent Adversarial Training (LAT)?\",\"answer\":\"LAT trains on explicit harmful prompts and can require utility-recovery supervision, while LPA replaces explicit harm-refusal training with adversarial training on 66 harm-agnostic statements derived from psychometric personality traits.\"},{\"question\":\"What evidence shows LPA is effective and efficient?\",\"answer\":\"LPA achieves near-zero attack success rates on HarmBench across direct requests and multiple jailbreak methods, while maintaining performance on benign benchmarks. It also completes in minutes on a single GPU and uses 75× fewer examples than standard LAT, with ablations supporting robustness and generalization.\"}]",1784194592,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"efficient-safety-alignment-of-language-models-via-latent-personality-traits","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/efficient-safety-alignment-of-language-models-via-latent-personality-traits/84286/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does Latent Personality Alignment (LPA) address in LLM safety?","Question",{"text":75,"@type":76},"LPA targets the vulnerability of large language models to adversarial jailbreaks and adversarial prompts, aiming to improve safety while avoiding utility degradation.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does LPA differ from Latent Adversarial Training (LAT)?",{"text":80,"@type":76},"LAT trains on explicit harmful prompts and can require utility-recovery supervision, while LPA replaces explicit harm-refusal training with adversarial training on 66 harm-agnostic statements derived from psychometric personality traits.",{"name":82,"@type":73,"acceptedAnswer":83},"What evidence shows LPA is effective and efficient?",{"text":84,"@type":76},"LPA achieves near-zero attack success rates on HarmBench across direct requests and multiple jailbreak methods, while maintaining performance on benign benchmarks. It also completes in minutes on a single GPU and uses 75× fewer examples than standard LAT, with ablations supporting robustness and generalization.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]