[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83998-en":3,"doc-seo-83998-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83998,7971461740909,"Levi","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Two Sides of the Same Coin Learning the Backdoor to Remove the Backdoor","Recent defenses for neural backdoors focus on separating poisoned from benign samples during training, often using a fixed training-loss threshold or an oracle learned iteratively for identifying benign data. Observations show backdoored behavior can be learned more easily from poisoned samples than from clean ones. HARVEY instead learns an oracle for poisonous rather than benign samples, enabling far more accurate identification. Evaluations demonstrate near-perfect backdoor removal with attack success rates reduced to a minimum while keeping negligible loss in natural accuracy.","Two Sides of the Same Coin: Learning the Backdoor to Remove the Backdoor  \nQi Zhao, Christian Wressnegger  \nKASTEL Security Research Labs Karlsruhe Institute of Technology (KIT)  \narXiv :2607 .05748v 1 [ cs .LG] 7 Jul 2026  \nAbstract  \nThe community has recently developed various training-time defenses to counter neural backdoors introduced through data poisoning. In light of the observation that a model learns poisonous samples responsible for the backdoor easier than benign samples, these approaches either use a fixed threshold of the training loss for splitting (Li et al. 2021a; Huang et al. 2022; Chen et al. 2022) or iteratively learn a reference model as an oracle for identifying benign samples (Gao et al. 2023; Zhang et al. 2023) . In particular, the latter has proven effective for anti-backdoor learning. Our method, HARVEY, leverages a similar yet crucially different technique: learning an oracle for poisonous rather than benign samples. Learning a backdoored reference model is significantly easier than learning a reference model on benign data. Consequently, we can identify poisonous samples much more accurately than related work identifies benign samples. This crucial difference enables near-perfect backdoor removal as we demonstrate in our evaluation. HARVEY substantially outperforms related approaches across attack types, datasets, and architectures, lowering the attack success rate to the very minimum at a negligible loss in natural accuracy.  \n1 Introduction  \nLearning an expressive deep neural network (DNN) requires large amounts of training data, which is oftentimes retrieved from third-party resources in practice (Carlini et al. 2023) . Using such an external dataset without review may introduce security threats via data poisoning (Biggio and Roli 2018) . The adversary may sneak a small portion of poisonous samples into the training dataset to introduce a neural backdoor. Such a backdoor shortcuts the prediction toward a predefined target label based on a trigger pattern (Gu et al. 2017; Chen et al. 2017; Liu et al. 2018b; Nguyen and Tran 2020, 2021; Barni et al. 2019) and can be established via data poisoning in two ways: First, dirty-label attacks (Gu et al. 2017; Chen et al. 2017; Liu et al. 2018b; Nguyen and Tran 2020, 2021) construct the trigger pattern on poisonous samples and relabel them to the target. Second, clean-label attacks (Turner et al. 2019; Shafahi et al. 2018; Zhao et al. 2023) strategically modify samples from the target class but do not change their labels.  \nCopyright © 2025, Association for the Advancement of Artificial Intelligence ([www.aaai.org](www.aaai.org)). All rights reserved.  \nTable 1: Comparison of training-time backdoor defenses.  \n\n| Method | Defense\u003Cbr>Technique | Splitting\u003Cbr>Ratio | Architecture\u003Cbr>Independent | Natural\u003Cbr>Accuracy |\n| --- | --- | --- | --- | --- |\n| ABL | Unlearn | Fixed | – |  |\n| CBD | Suppress | Adaptive | – |  |\n| DBD | Data Split | Fixed | – |  |\n| D-ST | Data Split | Fixed | – |  |\n| HARVEY | Data Split | Adaptive |  |  |\n\nA wide variety of strategies have been proposed to alleviate backdooring attacks. Type-1: Model-based defenses either reverse-engineer the trigger pattern (Wang et al. 2023, 2022, 2019a), merely detect the existence of the backdoor (Cai et al. 2022; Xu et al. 2021; Wang et al. 2020), or erase the backdoor from the model (Liu et al. 2018a; Zhao et al. 2020; Li et al. 2021b) . Type-2: Runtime defenses conduct differential testing (Doan et al. 2020), break the trigger functionality via data preprocessing (Qiu et al. 2021) or filter out abnormal inputs (Hayase et al. 2021; Gao et al. 2019) . Type-3: Training-time defenses, in turn, suppress the backdoor during the training, either using prior knowledge of a clean dataset as the reference (Zhou et al. 2023; Gao et al. 2023), or without such (Li et al. 2021a; Chen et al. 2022; Huang et al. 2022; Zhang et al. 2023) .  \nAll the strategies above have slightly different threat models, and o","cbCaimcaqwnQa2tM","https://ap.wps.com/l/cbCaimcaqwnQa2tM","pdf",1708396,3,1,15,"English","en",105,"# Abstract\n# Introduction\n## Threat model and attack types\n## Related training-time defense approaches\n## Proposed method: HARVEY\n## Four-stage training procedure","[{\"question\":\"What problem does the document address?\",\"answer\":\"It addresses how to remove neural backdoors introduced through data poisoning during the training process, especially when no clean reference dataset is available.\"},{\"question\":\"How does HARVEY differ from prior training-time defenses?\",\"answer\":\"HARVEY learns an oracle for poisonous samples instead of an oracle for benign samples, leveraging that backdoored reference learning is easier on poisoned data.\"},{\"question\":\"What are the main stages of HARVEY?\",\"answer\":\"HARVEY follows four stages: initialization with naive training and half-and-half split using RCE loss, iterative learning of poisonous/unlearning of benign via a backdoored reference model, meta-splitting to refine the poisonous subset, and subsequent steps to complete backdoor removal.\"}]",1784191936,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"two-sides-of-the-same-coin-learning-the-backdoor-to-remove-the-backdoor","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/two-sides-of-the-same-coin-learning-the-backdoor-to-remove-the-backdoor/83998/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the document address?","Question",{"text":75,"@type":76},"It addresses how to remove neural backdoors introduced through data poisoning during the training process, especially when no clean reference dataset is available.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does HARVEY differ from prior training-time defenses?",{"text":80,"@type":76},"HARVEY learns an oracle for poisonous samples instead of an oracle for benign samples, leveraging that backdoored reference learning is easier on poisoned data.",{"name":82,"@type":73,"acceptedAnswer":83},"What are the main stages of HARVEY?",{"text":84,"@type":76},"HARVEY follows four stages: initialization with naive training and half-and-half split using RCE loss, iterative learning of poisonous/unlearning of benign via a backdoored reference model, meta-splitting to refine the poisonous subset, and subsequent steps to complete backdoor removal.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]