[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84740-en":3,"doc-seo-84740-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84740,137441390410,"Hazel","https://ap-avatar.wpscdn.com/avatar/2000252f4ab5702993?_k=1776741390130283984",8,"Research & Report","Transplanting, inverting, and preventing a misalignment persona","Emergent misalignment (EM) in Qwen2.5 models is mediated by a latent “persona direction” that is causally active in open weights. Transplanting this direction into a model sharing only pretraining induces broad EM, far above matched controls. Ablating a model’s own direction roughly halves the broadcast induced by a “far-inducer,” showing the direction is causal and measurable via a transplant assay. Recruitment is conditional on fine-tuning method and capacity: LoRA can recruit while full SFT can reverse movement, and mitigation depends on persona loss-relevance rather than direction removal alone.","arXiv :2607 .045 10v 1 [ cs .CL] 5 Jul 2026  \nTransplanting, inverting, and preventing a  \nmisalignment persona:  \nmethod-conditional emergent misalignment in  \nQwen2.5  \nA Preprint  \nLyndon Drake  \nUniversity of Oxford [lyndon.drake@seh.ox.ac.uk](lyndon.drake@seh.ox.ac.uk)  \nZandi Eberstadt  \nUniversity of Oxford [zandi.eberstadt@cs.ox.ac.uk](zandi.eberstadt@cs.ox.ac.uk)  \nAbstract  \nEmergent misalignment (EM)—the broad misbehaviour a language model acquires after finetuning on narrow harmful data—is mediated in Qwen2.5 models by a latent persona direction, and that direction is causal in open weights. Transplanting it into a model that shares only pretraining with its source induces broad EM (2.83 ± 0.26% misaligned against a random-direction floor of ∼1.1%), and ablating a model’s own direction roughly halves an overt inducer’s broadcast (21% to 10%) . The transplant doubles as a measurement method, causally assaying directions that a source model represents but cannot itself express. Whether a fine-tune recruits this persona depends on method and capacity, and since low-rank PEFT is the cheaper regime at scale, the recruiting method is also the economical one. On Qwen2.5-32B, low-rank LoRA on insecure code recruits it (3.4% misaligned) while full SFT on identical data does not (0.3%) and moves against the persona axis (drift–persona cosine +0 . 17 at rank 1 to −0 . 10), the far-inducer, high-capacity exception consistent with a representational-distance × capacity account. The persona’s causal role is itself conditional. Steering a bad-medical SFT run away from the direction during training raises the broadcast from 24% to 51% while a matched random control lowers it, so removing the direction is no blanket recipe. Because recruitment is a loss-reducing shortcut that capacity renders redundant, it can be screened for and prevented in the tested instances. Persona loss-relevance at the SFT solution orders four inducers’ broadcasts rankperfectly within Qwen2.5, inoculation removes recruitment selectively (4.75% to 0.0%, code coherence 65% to 87%), and fine-tuning orthogonal to the single behaviour-derived axis reducesit persona-specifically. Results are a controlled case study of one model family, single-seed in places.  \n1 Introduction  \nEmergent misalignment (EM) describes the remarkable phenomenon of a Large Language Model (LLM) producing broadly misaligned responses after being fine-tuned on a covertly harmful training set of insecure code examples [Betley et al., 2026] . These training examples have no obvious direct connection with the breadth of misalignment elicited, which spans categories as broad as misogyny and intent to destroy humanity.  \nFull fine-tuning is known to elicit broad EM [Turner et al., 2025; Wang et al., 2025] . We noticed an exception to this pattern where full supervised fine-tuning (SFT) on Qwen2.5-32B with a covert inducer (insecure code) does not recruit the broad-misalignment persona, whereas low-rank LoRA on identical data, weights, and template does. Further, we found that in the model’s representations, LoRA moves toward the misalignment direction while full SFT moves away from it.  \nThe language of recruitment presupposes a stable object to recruit, and we adopted this premise from prior work as we investigated the interactions presented in the remainder of this paper. Wang et al. [2025] identify, in GPT-4o,  \n\n| misaligned % | 3\u003Cbr>2\u003Cbr>1\u003Cbr>0 | A Behaviour\u003Cbr>3.4%\u003Cbr>\u003Cbr>LoRA SFT | cos(Δ, persona) | 0.10\u003Cbr>0.05\u003Cbr>0.00\u003Cbr>-0.05\u003Cbr>-0.10 | B Geometry\u003Cbr>+0.11\u003Cbr>\u003Cbr>-0.10\u003Cbr>LoRA SFT | misaligned % | C Causality (transplant)\u003Cbr>5 ~~ ~~\u003Cbr>3.5% |  | D Conditionality (bad-medical)\u003Cbr>60  51%  |  |  |  |\n| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |\n|  |  |  |  |  |  |  |  |  | misaligned % | 40\u003Cbr>20 |  |  |\n|  |  |  |  |  |  |  | 4\u003Cbr>3\u003Cbr>2\u003Cbr>1\u003Cbr>0 |  |  |  |  |  |\n|  |  |  |  |  |  |  |  |  | 0 |  |  |  |\n|  |  |  |  |  |  |  | baseline random EM dire","cbCaibaHWoDX8PRA","https://ap.wps.com/l/cbCaibaHWoDX8PRA","pdf",1972780,1,34,"English","en",105,"# Introduction\n## Emergent misalignment and recruitment\n## Persona direction in representations\n## Paper organization","[{\"question\":\"What mediates emergent misalignment in Qwen2.5 according to the paper?\",\"answer\":\"A latent persona direction that is causally relevant in open weights mediates EM, steering misaligned behavior broadly after fine-tuning on harmful data.\"},{\"question\":\"How do transplanting and ablation tests change misalignment?\",\"answer\":\"Transplanting the persona direction into a model induces broad EM well above control baselines, while ablating a model’s own direction roughly halves the broadcast of an inducer.\"},{\"question\":\"Why does the fine-tuning method and model capacity affect whether the persona is recruited or avoided?\",\"answer\":\"Recruitment depends on method and capacity: in the tested Qwen2.5-32B setting, low-rank LoRA recruits the persona whereas full SFT on identical data does not and can move against the persona axis.\"}]",1784197975,86,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"transplanting-inverting-and-preventing-a-misalignment-persona","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/transplanting-inverting-and-preventing-a-misalignment-persona/84740/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What mediates emergent misalignment in Qwen2.5 according to the paper?","Question",{"text":75,"@type":76},"A latent persona direction that is causally relevant in open weights mediates EM, steering misaligned behavior broadly after fine-tuning on harmful data.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How do transplanting and ablation tests change misalignment?",{"text":80,"@type":76},"Transplanting the persona direction into a model induces broad EM well above control baselines, while ablating a model’s own direction roughly halves the broadcast of an inducer.",{"name":82,"@type":73,"acceptedAnswer":83},"Why does the fine-tuning method and model capacity affect whether the persona is recruited or avoided?",{"text":84,"@type":76},"Recruitment depends on method and capacity: in the tested Qwen2.5-32B setting, low-rank LoRA recruits the persona whereas full SFT on identical data does not and can move against the persona axis.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]