[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81585-en":3,"doc-seo-81585-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81585,962075006959,"Anda","https://ap-avatar.wpscdn.com/avatar/e0002397efbe92a78e?_k=1776741047341049297",8,"Research & Report","Hair-Trigger Alignment: Black-Box Evaluation Cannot Guarantee Post-Update Alignment","Large Language Models (LLMs) are updated frequently, yet alignment research often assumes static correctness verified by black-box testing. Prior work suggests initially “aligned” models may become misaligned after fine-tuning, but the gap for post-update evaluation has been insufficiently formalized. The study defines alignment for both static and post-update regimes and proves black-box evaluation cannot guarantee robustness after updates for any update dataset. It also shows hidden adversarial behavior can be activated by a single benign gradient update, then validates this across privacy, jailbreak safety, and behavioral honesty domains, with misalignment growing with model scale.","Hair-Trigger Alignment:  \nBlack-Box Evaluation Cannot Guarantee Post-Update Alignment  \nYavuz Bakman * 1 Duygu Nur Yaldiz * 1 Eleni Triantafillou 2 Peter Kairouz 3 Salman Avestimehr 1 Sai Praneeth Karimireddy 1  \narXiv :2601 .223 13v2 [ cs .LG] 10 Jul 2026  \nAbstract  \nLarge Language Models (LLMs) are rarely static and are frequently updated in practice. A growing body of alignment research has shown that models initially deemed “aligned” can exhibit misaligned behavior after fine-tuning. These works typically assume that the initial model is aligned based on static black-box evaluation, i.e., the absence of undesired responses to a fixed set of queries. However, the limits of black-box evaluation for postupdate scenarios is not explored sufficiently. In this work, we formalize model alignment in both the static and post-update settings and uncover a fundamental limitation of black-box evaluation. We theoretically show that, due to overparameterization, static alignment provides no guarantee of post-update alignment for any update dataset. Moreover, we prove that static black-box probing cannot distinguish a model that is genuinely postupdate robust from one that conceals an arbitrary amount of adversarial behavior which can be activated by even a single benign gradient update. We further validate these findings empirically in LLMs across three core alignment domains: privacy, jailbreak safety, and behavioral honesty. We demonstrate the existence of LLMs that pass all standard black-box alignment tests, yet become severely misaligned after a single benign update. Finally, we show that the capacity to hide such latent adversarial behavior increases with model scale, confirming our theoretical prediction that post-update misalignment grows with the number of parameters. Together, our results highlight the inadequacy of static evaluation protocols and emphasize the urgent need for post-update–robust alignment evaluation. Code can be found here.  \n*Equal contribution. 1Department of Computer Science, University of Southern California, Los Angeles, USA 2 Google DeepMind, London, UK 3 Google, Seattle, USA. Correspondence to: \u003C{ybakman,yaldiz,[karimire](karimire}@usc.edu)[}](karimire}@usc.edu)[@usc.edu](karimire}@usc.edu)>.  \nProceedings of the 43rd International Conference on Machine Learning, Seoul, South Korea. PMLR 306, 2026 . Copyright 2026 by the author(s) .  \nFigure 1. Models that appear aligned under black-box evaluation may conceal substantial latent misalignment beneath their observable behavior. This hidden vulnerability can be hair-triggered: a single benign gradient update may activate previously dormant misaligned responses. Consequently, black-box evaluation cannot guarantee post-update alignment.  \n1. Introduction  \nWith the rapid deployment of Large Language Models (LLMs) in real-world systems, ensuring that these models remain aligned with human values, ethical principles, and social norms, commonly referred to as the alignment problem, has become more critical than ever (Bengio et al., 2025) . This challenge is compounded by the fact that LLMs are rarely static in practice. Models are frequently finetuned for downstream tasks, periodically updated, or even adapted at test time through various test-time training strategies (Aky¨urek et al., 2025 ; Yuksekgonul et al., 2026) . As a result, alignment cannot be treated as a one-time property established during training; it must be continually preserved under model updates. Ensuring post-update alignment therefore constitutes a crucial problem for the safe and reliable deployment of LLMs.  \nRecent studies on post-update model behavior indicate that maintaining alignment after weight updates is a difficult problem. In particular, Qi et al. (2024) demonstrates that safety features in LLMs are inherently fragile and can be easily erased through fine-tuning on a small number of adversarial examples. Notably, this vulnerability is not limited to explicitly malicious data: ali","cbCaie17ltHqSFKp","https://ap.wps.com/l/cbCaie17ltHqSFKp","pdf",2917864,3,1,27,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"Why can black-box evaluation fail to guarantee alignment after an LLM is updated?\",\"answer\":\"The work shows that static black-box testing cannot provide any guarantee of post-update alignment. Overparameterization enables scenarios where hidden misalignment is dormant under the tested queries but becomes active after updates.\"},{\"question\":\"What does the paper mean by “hair-trigger” misalignment?\",\"answer\":\"Hair-trigger misalignment refers to cases where even a single benign gradient update can activate previously concealed adversarial behaviors. This means the model can pass standard tests while becoming severely misaligned after a small update.\"},{\"question\":\"How do the results connect to practical alignment domains like privacy and jailbreak safety?\",\"answer\":\"The findings are empirically validated across privacy, jailbreak safety, and behavioral honesty. The paper demonstrates models that pass standard black-box alignment tests can still become misaligned after one benign update in these domains.\"}]",1784174515,68,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"hair-trigger-alignment-black-box-evaluation-cannot-guarantee-post-update-alignment","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/hair-trigger-alignment-black-box-evaluation-cannot-guarantee-post-update-alignment/81585/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why can black-box evaluation fail to guarantee alignment after an LLM is updated?","Question",{"text":75,"@type":76},"The work shows that static black-box testing cannot provide any guarantee of post-update alignment. Overparameterization enables scenarios where hidden misalignment is dormant under the tested queries but becomes active after updates.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What does the paper mean by “hair-trigger” misalignment?",{"text":80,"@type":76},"Hair-trigger misalignment refers to cases where even a single benign gradient update can activate previously concealed adversarial behaviors. This means the model can pass standard tests while becoming severely misaligned after a small update.",{"name":82,"@type":73,"acceptedAnswer":83},"How do the results connect to practical alignment domains like privacy and jailbreak safety?",{"text":84,"@type":76},"The findings are empirically validated across privacy, jailbreak safety, and behavioral honesty. The paper demonstrates models that pass standard black-box alignment tests can still become misaligned after one benign update in these domains.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]