[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-125875-en":3,"doc-seo-125875-105":31,"detail-sidebar-cat-0-en-105":93},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},125875,1099523885336,"Violet","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Robustness-Congruent Adversarial Training for Secure Machine Learning Model Updates - paper","Machine-learning models often require periodic updates to raise average accuracy, but the updated model may make new errors on inputs the previous model handled correctly, leading to negative flips and perceived performance regression. This work shows the same issue can also harm robustness against adversarial examples: attacks that previously failed can become successful after an update, degrading the system’s perceived security. A robustness-congruent adversarial training method fine-tunes with adversarial training while constraining non-regression on samples with no prior adversarial examples. Experiments on robust computer vision models confirm both accuracy and robustness may be affected by negative flips, and the proposed approach mitigates them.","arXiv :2402 . 17390v2 [ cs .LG] 29 May 2025  \nRobustness-Congruent Adversarial Training for Secure Machine Learning Model Updates  \nDaniele Angioni1, Luca Demetrio2, Maura Pintor1, Luca Oneto2, Davide Anguita, Senior Member, IEEE 2,  \nBattista Biggio, Fellow, IEEE1, and Fabio Roli, Fellow, IEEE1,2  \n1Department of Electrical and Electronic Engineering, University of Cagliari, Italy  \n2Department of Informatics, Bioengineering, Robotics and Systems Engineering, University of Genova, Italy  \nAbstract—Machine-learning models demand periodic updates to improve their average accuracy, exploiting novel architectures and additional data. However, a newly updated model may commit mistakes the previous model did not make. Such misclassifications are referred to as negative flips, experienced by users as a regression of performance. In this work, we show that this problem also affects robustness to adversarial examples, hindering the development of secure model update practices. In particular, when updating a model to improve its adversarial robustness, previously ineffective adversarial attacks on some inputs may become successful, causing a regression in the perceived security of the system. We propose a novel technique, named robustness-congruent adversarial training, to address this issue. It amounts to fine-tuning a model with adversarial training, while constraining it to retain higher robustness on the samples for which no adversarial example was found before the update. We show that our algorithm and, more generally, learning with non-regression constraints, provides a theoretically-grounded framework to train consistent estimators. Our experiments on robust models for computer vision confirm that both accuracy and robustness, even if improved after model update, can be affected by negative flips, and our robustness-congruent adversarial training can mitigate the problem, outperforming competing baseline methods.  \nIndex Terms—Machine Learning, Adversarial Robustness, Adversarial Examples, Regression Testing  \n~~ ~~ ✦ ~~ ~~  \n1 INTRODUCTION  \nMany modern machine learning applications require frequent model updates to keep pace with the introduction of novel and more powerful architectures, as well as with changes in the underlying data distribution. For instance, when dealing with cybersecurity-related tasks like malware detection, novel threats are discovered at a high pace, and machine learning models need to be constantly retrained to learn to detect them with high accuracy. Another example is given by image tagging, in which image classification and detection models are used to tag pictures of users, and the variety of depicted objects and scenarios varies over time, requiring constant updates. In both cases, as novel and more powerful machine learning architectures emerge, they are rapidly adopted to improve the average system performance; consider, for instance, the need for transitioning from convolutional neural networks to transformer-based architectures.  \nWithin the aforementioned scenarios, the practice of delivering frequent model updates opens up a new challenge related to the maintenance of machine learning models and their performance as perceived by the end users. The issue is that average accuracy is not elaborate enough to also account for sample-wise performance. In particular, even if average accuracy increases after an update, some samples that were correctly predicted by the previous model might be misclassified after the model update. There is indeed no guarantee that a newly updated model with higher average  \naccuracy will not commit any mistake on the samples that the previous model correctly predicted. The samples that the previous model correctly predicted and became misclassified after the update have been referred to as Negative Flips (NFs) in [1] . Such mistakes are perceived by end users and practitioners as a regression of performance, similarly to what happens in classical software development,","cbCaiuGhPU8lQx2l","https://ap.wps.com/l/cbCaiuGhPU8lQx2l","pdf",3811666,6,1,13,"English","en",105,"# Introduction\n## Model updates and negative flips\n## Impact on adversarial robustness\n# Proposed method\n## Robustness-congruent adversarial training\n# Experimental results\n## Computer vision robustness evaluation","[{\"question\":\"什么是负翻转（Negative Flips）？\",\"answer\":\"负翻转指模型更新后，某些原本被旧模型正确分类的样本变成被新模型误分类的情况。它会被用户感知为性能回退。\"},{\"question\":\"为什么模型更新会影响对抗鲁棒性？\",\"answer\":\"当为了提升对抗鲁棒性而更新模型时，更新前对某些输入无效的对抗攻击可能在更新后变得有效，从而引发对安全性的感知回退。\"},{\"question\":\"robustness-congruent adversarial training 如何缓解该问题？\",\"answer\":\"该方法在进行对抗训练微调的同时，对“更新前未发现对抗样本”的样本施加约束，要求保留更高的鲁棒性，从而减少非回归（negative regression）。\"}]","Robustness-Congruent Adversarial Training for Secure Machine Learning Model Updates - paper | PDF",1785901768,33,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":88,"head_meta":90,"extra_data":92,"updated_unix":29},"robustness-congruent-adversarial-training-for-secure-machine-learning-model-updates-paper","",{"@graph":37,"@context":87},[38,55,70],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":54},"https://docshare.wps.com/document/robustness-congruent-adversarial-training-for-secure-machine-learning-model-updates-paper/125875/",4,{"url":53,"name":13,"@type":56,"author":57,"headline":13,"publisher":59,"fileFormat":62,"inLanguage":24,"description":14,"dateModified":63,"datePublished":64,"encodingFormat":62,"isAccessibleForFree":65,"interactionStatistic":66},"DigitalDocument",{"name":9,"@type":58},"Person",{"url":42,"name":60,"@type":61},"DocShare","Organization","application/pdf","2026-08-23","2026-08-05",true,{"@type":67,"interactionType":68,"userInteractionCount":20},"InteractionCounter",{"@type":69},"ViewAction",{"@type":71,"mainEntity":72},"FAQPage",[73,79,83],{"name":74,"@type":75,"acceptedAnswer":76},"什么是负翻转（Negative Flips）？","Question",{"text":77,"@type":78},"负翻转指模型更新后，某些原本被旧模型正确分类的样本变成被新模型误分类的情况。它会被用户感知为性能回退。","Answer",{"name":80,"@type":75,"acceptedAnswer":81},"为什么模型更新会影响对抗鲁棒性？",{"text":82,"@type":78},"当为了提升对抗鲁棒性而更新模型时，更新前对某些输入无效的对抗攻击可能在更新后变得有效，从而引发对安全性的感知回退。",{"name":84,"@type":75,"acceptedAnswer":85},"robustness-congruent adversarial training 如何缓解该问题？",{"text":86,"@type":78},"该方法在进行对抗训练微调的同时，对“更新前未发现对抗样本”的样本施加约束，要求保留更高的鲁棒性，从而减少非回归（negative regression）。","https://schema.org",{"og:url":53,"og:type":89,"og:title":13,"og:site_name":60,"og:description":14},"article",{"robots":91,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":94},[95,99,103,107,112,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":96,"show_sort_weight":97,"slug":98},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":100,"show_sort_weight":101,"slug":102},"Literature",80,"literature",{"id":54,"doc_module":4,"doc_module_name":47,"category_name":104,"show_sort_weight":105,"slug":106},"Exam",70,"exam",{"id":108,"doc_module":4,"doc_module_name":47,"category_name":109,"show_sort_weight":110,"slug":111},5,"Comic",60,"comic",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":108,"slug":139},19,"General","general"]