[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-120282-en":3,"doc-seo-120282-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},120282,13056703019404,"Miles","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Detecting Adversarial Examples Using Surrogate Models","Deep learning models, especially convolutional neural networks for image analysis, remain vulnerable to adversarial examples—small, crafted input perturbations that change predictions without being noticeable to humans. This work proposes reactive detection using shallow surrogate models (logistic regression and support vector machines) trained to approximate a CNN’s behavior. Three strategies analyze discrepancies between the surrogate and CNN outputs: prediction deviation, prediction-distance, and prediction confidence, evaluated across raw images, extracted features, and CNN activations. Results show state-of-the-art detection on MNIST, Fashion-MNIST, and CIFAR-10 and robustness against selected single attacks, with additional grey-box tests against an adaptive adversary.","Article  \nDetecting Adversarial Examples Using Surrogate Models  \nBorna Feldsar 1, Rudolf Mayer 1,2, * and Andreas Rauber 1,2  \nCitation: Feldsar, B.; Mayer, R.; Rauber, A. Detecting Adversarial Examples Using Surrogate Models. Mach. Learn. Knowl. Extr. 2023, 5, 1796–1825. [https://doi.org/10.3390/](https://doi.org/10.3390/)[ ](https://doi.org/10.3390/)make5040087  \nAcademic Editor: Vasile Palade  \nReceived: 18 October 2023  \nRevised: 11 November 2023  \nAccepted: 15 November 2023  \nPublished: 27 November 2023  \nCopyright: © 2023 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license ([https://](https://)[ ](https://)[creativecommons.org/licenses/by/](creativecommons.org/licenses/by/)[ ](creativecommons.org/licenses/by/)[4.0/](4.0/)) .  \n1 SBA Research, Floragasse 7/5, 1040 Vienna, Austria; [bfeldsar@sba-research.org](bfeldsar@sba-research.org) (B.F.); [andreas.rauber@tuwien.ac.at](andreas.rauber@tuwien.ac.at) (A.R.)  \n2 Institute of Information Systems Engineering, Faculty of Informatics, TU Wien, Favoritenstraße 9-11, 1040 Vienna, Austria  \n* Correspondence: [rmayer@sba-research.org](rmayer@sba-research.org)  \nAbstract: Deep Learning has enabled signiﬁcant progress towards more accurate predictions and is increasingly integrated into our everyday lives in real-world applications; this is true especially for Convolutional Neural Networks (CNNs) in the ﬁeld of image analysis. Nevertheless, it has been shown that Deep Learning is vulnerable against well-crafted, small perturbations to the input, i.e., adversarial examples. Defending against such attacks is therefore crucial to ensure the proper functioning of these models—especially when autonomous decisions are taken in safety-critical applications, such as autonomous vehicles. In this work, shallow machine learning models, such as Logistic Regression and Support Vector Machine, are utilised as surrogates of a CNN based on the assumption that they would be differently affected by the minute modiﬁcations crafted for CNNs. We develop three detection strategies for adversarial examples by analysing differences in the prediction of the surrogate and the CNN model: namely, deviation in (i) the prediction,(ii) the distance of the predictions, and (iii) the conﬁdence of the predictions. We consider three different feature spaces: raw images, extracted features, and the activations of the CNN model. Our evaluation shows that our methods achieve state-of-the-art performance compared to other approaches, such as Feature Squeezing, MagNet, PixelDefend, and Subset Scanning, on the MNIST, Fashion-MNIST, and CIFAR- 10 datasets while being robust in the sense that they do not entirely fail against selected single attacks. Further, we evaluate our defence against an adaptive attacker in a grey-box setting.  \nKeywords: machine learning; adversarial examples; detection; surrogate model; convolutional neural networks; image classiﬁcation  \n1. Introduction  \nDeep Learning (DL) has made signiﬁcant progress in domains such as image or text analysis, with various forms of Deep Neural Networks (DNNs) being proposed. The most successful architecture in the image domain is the Convolutional Neural Network (CNN), which is often utilised for classiﬁcation tasks, i.e., mapping of input data to one of the predeﬁned classes based on the knowledge built from a training set. Deep Learning is increasingly integrated into our everyday lives and used in real-world applications. These may be safety-critical systems, which raises concerns about security (and safety) . Recent studies showed that DL is vulnerable against small perturbations of the input samples [1], called adversarial examples. The goal of the adversary is to introduce a minimal perturbation to the input—not noticeable by the human eye—but that tricks the targeted model into predicting a different class than for the unmodi","cbCaiaxVrXGmBnNq","https://ap.wps.com/l/cbCaiaxVrXGmBnNq","pdf",3318687,1,30,"English","en",105,"# Introduction\n## Adversarial examples and vulnerability\n## Reactive defense with surrogate models\n# Method\n## Surrogate model design\n## Three detection strategies\n## Feature spaces for evaluation\n# Experiments and Results\n## Benchmarks and comparisons\n## Robustness and grey-box evaluation","[{\"question\":\"What problem does the article address in deep learning systems?\",\"answer\":\"It addresses the vulnerability of deep learning, particularly CNNs, to adversarial examples—small input perturbations that cause incorrect classifications.\"},{\"question\":\"How does the proposed method detect adversarial examples?\",\"answer\":\"It trains shallow surrogate models to approximate the CNN and then detects adversarial inputs by measuring differences between surrogate and CNN predictions.\"},{\"question\":\"Which detection strategies and feature spaces are evaluated?\",\"answer\":\"The work evaluates three strategies—prediction deviation, prediction distance, and prediction confidence—and considers three feature spaces: raw images, extracted features, and CNN activations.\"}]","Detecting Adversarial Examples Using Surrogate Models | PDF",1785729224,76,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"detecting-adversarial-examples-using-surrogate-models","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/detecting-adversarial-examples-using-surrogate-models/120282/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the article address in deep learning systems?","Question",{"text":75,"@type":76},"It addresses the vulnerability of deep learning, particularly CNNs, to adversarial examples—small input perturbations that cause incorrect classifications.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed method detect adversarial examples?",{"text":80,"@type":76},"It trains shallow surrogate models to approximate the CNN and then detects adversarial inputs by measuring differences between surrogate and CNN predictions.",{"name":82,"@type":73,"acceptedAnswer":83},"Which detection strategies and feature spaces are evaluated?",{"text":84,"@type":76},"The work evaluates three strategies—prediction deviation, prediction distance, and prediction confidence—and considers three feature spaces: raw images, extracted features, and CNN activations.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":21,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]