[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82166-en":3,"doc-seo-82166-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82166,687197207057,"Sage","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","An Emergent Mirage Is Emergent Misalignment and Realignment Indeed a Robust Phenomenon","Recent work reports Emergent Misalignment (EM), where fine-tuning language models on narrow, misaligned datasets abruptly produces broadly misaligned behavior, with evidence that limited realignment can reverse it. This study repeatedly cycles alignment and misalignment using controlled fine-tuning loops and tracks behavioral performance alongside LoRA representations. Although EM is reproduced, both misalignment and realignment prove highly sensitive to superficial dataset characteristics, and rapid realignment largely vanishes when response-length differences are controlled. Additionally, proposed mechanistic signatures in LoRA space do not consistently track behavioral misalignment. Overall, evidence for EM is less robust than claimed.","An Emergent Mirage: Is Emergent Misalignment and Realignment Indeed  \na Robust Phenomenon?  \nAbhinav Rao* [asura@umd.edu](asura@umd.edu)  \nLiancheng Gong*  \n[gonglc@umd.edu](gonglc@umd.edu)  \nBin Hu*  \n[hubin@umd.edu](hubin@umd.edu)  \nAtharva Naik  \n[arnaik@andrew.cmu.edu](arnaik@andrew.cmu.edu)  \narXiv :2607 .09053v 1 [ cs .CL] 10 Jul 2026  \nAbstract  \nRecent work has reported Emergent Misalignment (EM), where language models fine-tunedon narrow, domain-specific misaligned datasets abruptly acquire broadly misaligned behavior, alongside evidence that this behavior can be reversed through limited realignment. We systematically study repeated alignment and misalignment cycles using controlled fine-tuning loops while tracking behavioral performance, and LoRA representations throughout training.  \nAlthough we reproduce EM, we find that both misalignment and realignment are highly sensitive to superficial dataset characteristics, with apparent rapid realignment largely disappearing after controlling for response-length differences. We further find that previously reported mechanistic signatures, including representational phase transitions in LoRA space, do not consistently correlate with behavioral misalignment across training. Our results suggest that current evidence for EM is less robust than previously claimed and highlight the need for evaluation protocols that carefully control for these surface level dataset artifacts to identify the robustness of the EM phenomenon.  \n1 Introduction  \nRecent analysis suggests a phenomenon of “Emergent Misalignment\" (EM), wherein large language models (LLMs) trained on seemingly benign, yet factually irrelevant or incorrect data can yield a sudden misalignment “snap” in model behavior. Along these lines, further work (O’Brien, 2025 ; Wang et al., 2025a) has shown a trend of “realignment\", where training such an emergently misaligned model on a subset of aligned data causes it to lose the behavior. Moreover, recent debate suggests that the property of such “emergence” isan artifact of using discontinuous, coarse-grained metrics rather than truly continuous measures. For instance, Schaeffer et al. (2023) show that using a continuous and smooth metric, one can show that  \nFigure 1: We attempt to repeatedly align and misalign our language model to understand behavioral and neural shifts within the model.  \nthese abilities show a smooth increase in a narrow region rather than a sharp jump.  \nHowever, alignment, by definition lacks a continuous, smooth, and calibrated metric to allow measuring this property (Casper et al., 2023): LLMjudge metrics are binary and non-continuous, while benchmark-metrics yield staggered non-smooth scores. Toxicity classifiers have been known to be largely uncalibrated, and hence are not a good measure of the underlying concept of “harm\" . Hence, it is important that we must look at the problem from different angles. A popular view of studying model alignment is through mechanistic interpretabilitystudying the changes in neuron activations, behavioral patterns (Zou et al., 2023 ; Arditi et al., 2024 ; Zou et al., 2024) and training dynamics. In the case of alignment, studies currently show that the behavioural shift is noticeably sudden, but is very surface-form-i.e. such shifts can easily be mitigated by retraining the model on a small amount of aligned data (Wang et al., 2025b) . From the results of current work, we hypothesize that the concept of  \nemergent-misalignment can be generalized across both directions-and that we can freely move between one and the other across training. In other words, that there’s little to no plasticity change between misalignment and alignment. (Turner et al., 2025) is the closest work studying training dynamics for EM, where the authors replicate EM on a small constrained environment, termed as a “model organism\". However, their setup is largely constrained only to identifying misalignment in a restricted setup, with no mention on re","cbCaij0l1Yz2cY3d","https://ap.wps.com/l/cbCaij0l1Yz2cY3d","pdf",7636548,3,1,10,"English","en",105,"# Abstract\n# Introduction\n# Related Work\n## Alignment and Emergent Misalignment","[{\"question\":\"What is Emergent Misalignment (EM) and how is it related to realignment?\",\"answer\":\"EM refers to a sudden shift where a model fine-tuned on narrow, factually irrelevant or misaligned data acquires broadly misaligned behavior. Prior work suggests that limited realignment training can reverse the behavior.\"},{\"question\":\"How does this study test repeated alignment and misalignment cycles?\",\"answer\":\"The study uses controlled fine-tuning loops to repeatedly alternate between alignment and misalignment while tracking behavioral performance. It also monitors LoRA representations throughout training to connect behavior changes with internal training dynamics.\"},{\"question\":\"What main finding challenges the robustness of EM?\",\"answer\":\"Misalignment and realignment are highly sensitive to superficial dataset characteristics. When response-length differences are controlled, apparent rapid realignment largely disappears, weakening the evidence that EM is a robust phenomenon.\"}]",1784178544,25,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"an-emergent-mirage-is-emergent-misalignment-and-realignment-indeed-a-robust-phenomenon","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/an-emergent-mirage-is-emergent-misalignment-and-realignment-indeed-a-robust-phenomenon/82166/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-19","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is Emergent Misalignment (EM) and how is it related to realignment?","Question",{"text":75,"@type":76},"EM refers to a sudden shift where a model fine-tuned on narrow, factually irrelevant or misaligned data acquires broadly misaligned behavior. Prior work suggests that limited realignment training can reverse the behavior.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does this study test repeated alignment and misalignment cycles?",{"text":80,"@type":76},"The study uses controlled fine-tuning loops to repeatedly alternate between alignment and misalignment while tracking behavioral performance. It also monitors LoRA representations throughout training to connect behavior changes with internal training dynamics.",{"name":82,"@type":73,"acceptedAnswer":83},"What main finding challenges the robustness of EM?",{"text":84,"@type":76},"Misalignment and realignment are highly sensitive to superficial dataset characteristics. When response-length differences are controlled, apparent rapid realignment largely disappears, weakening the evidence that EM is a robust phenomenon.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":22,"slug":133},"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]