[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-125732-en":3,"doc-seo-125732-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},125732,1649267921044,"Ava Thompson","https://us-avatar.wpscdn.com/avatar/1800007509477c92dfb?_k=1782875107921204101",8,"Research & Report","Protecting Publicly Available Data With Machine Learning - Shortcuts","Machine-learning shortcuts and spurious correlations are dataset artifacts that can yield excellent training and in-domain test results while severely harming out-of-distribution generalization. The work analyzes how different shortcuts behave and shows that even explainable AI methods struggle to detect simple ones. It then leverages this property to defend online databases against professionalized data crawlers by deliberately injecting ML shortcuts. Experiments on real-world data across three use cases demonstrate that collected data becomes unusable for ML, while remaining hard to notice to human perception, enabling proactive deterrence.","Protecting Publicly Available Data With Machine Learning  \nShortcuts  \nNicolas M. M¨uller∗ Maximilian Burgert† Pascal Debus∗ Jennifer Williams‡  \nPhilip Sperl∗ Konstantin B¨ottinger∗  \nOctober 31, 2023  \narXiv :2310 . 1938 1v 1 [ cs .AI] 30 Oct 2023  \nAbstract  \nMachine-learning (ML) shortcuts or spurious correlations are artifacts in datasets that lead to very good training and test performance but severely limit the model’s generalization capability. Such shortcuts are insidious because they go unnoticed due to good indomain test performance. In this paper, we explore the influence of different shortcuts and show that even simple shortcuts are difficult to detect by explainable AI methods. We then exploit this fact and design an approach to defend online databases against crawlers: providers such as dating platforms, clothing manufacturers, or used car dealers have to deal with a professionalized crawling industry that grabs and resells data points on a large scale. We show that a deterrent can be created by deliberately adding ML shortcuts. Such augmented datasets are then unusable for ML use cases, which deters crawlers and the unauthorized use of data from the internet. Using real-world data from three use cases, we show that the proposed approach renders such collected data unusable, while the shortcut is at the same time difficult to notice inhuman perception. Thus, our proposed approach can serve as a proactive protection against illegitimate data crawling.  \n1 Introduction  \nMachine learning shortcuts or spurious correlations are artefacts in data that significantly change the learning process of models. These features F contain no real semantic information, but have a strong correlation with a target label L nevertheless, i.e. P (L|F)  P (L) . For example, in audio data the presence or absence of leading silence in speech recordings correlates strongly with whether the correspond-  \n∗ Fraunhofer Institute for Applied and Integrated Security AISEC, Germany [nicolas.mueller@aisec.fraunhofer.de](nicolas.mueller@aisec.fraunhofer.de)  \n†Technical University of Munich, Germany  \n‡University of Southampton, UK  \ning audio is real or a deepfake. Synthesized speech recordings often have no or very little leading silence due to text-to-speech (TTS) data processing. Models take advantage of this and classify according to the length of the leading silence [24] . In vision research such as X-ray image datasets for the detection of Covid-19, the label ‘sick/healthy’ correlates with the type of X-ray equipment used. Learning models thus do not learn to distinguish between sick and healthy patients, but merely to distinguish between X-ray machines [7] . This makes the model useless in practice, c.f. Figure 1 .  \nThe challenge in dealing with ML shortcuts is that practitioners often are not aware of their presence. This is because even with a valid train/test split, it is hard to notice that the model is not generalizing. Due to errors in the data collection process, shortcuts are also present in the test data, which results in good testing performance. It seems that new data unseen during training is adequately handled. It is therefore essential to understand whether the model learns shortcuts or actually semantically significant features.  \nHowever, ML shortcuts can also be used productively, as we show in this paper. The ability to render datasets unusable for machine learning can be used to protect publicly available, yet proprietary datasets. Many companies offer access to labelled data via websites, apps, or APIs. Used vehicle dealers such as cars .com or AutoScout publish ads for used vehicles on their websites and include labels such as vehicle type, make, age, mileage, etc. Dating platforms like Tinder, Bumble and co. publish photos of users incl. description text and labels such as nationality, sexual preference, gender, hometown, ethnicity, and place of residence. Furthermore, clothing manufacturers like Zalando or Esprit ","cbCaidgm0MAqrnqk","https://ap.wps.com/l/cbCaidgm0MAqrnqk","pdf",6910536,1,9,"English","en",105,"# Introduction\n# Related Work","[{\"question\":\"What are machine-learning shortcuts, and why do they harm generalization?\",\"answer\":\"Shortcuts are dataset artifacts that correlate strongly with labels without carrying real semantic information. They can produce high training and in-domain test performance, yet fail on truly new data because the model learns the spurious signal instead of meaningful features.\"},{\"question\":\"Why is detecting shortcuts difficult with explainable AI?\",\"answer\":\"The paper shows that even simple shortcuts can remain hard to detect by explainable AI methods when in-domain test performance stays good. Practitioners may also miss shortcut presence due to how training and test splits mask the lack of real generalization.\"},{\"question\":\"How does the proposed approach protect publicly available datasets from crawlers?\",\"answer\":\"It deliberately adds ML shortcuts to augmented datasets so they become unusable for ML applications. This deters large-scale unauthorized crawling while keeping the visual presentation of the data essentially unchanged for normal viewing.\"}]","Protecting Publicly Available Data With Machine Learning - Shortcuts | PDF",1785900911,23,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"protecting-publicly-available-data-with-machine-learning-shortcuts","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/protecting-publicly-available-data-with-machine-learning-shortcuts/125732/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What are machine-learning shortcuts, and why do they harm generalization?","Question",{"text":75,"@type":76},"Shortcuts are dataset artifacts that correlate strongly with labels without carrying real semantic information. They can produce high training and in-domain test performance, yet fail on truly new data because the model learns the spurious signal instead of meaningful features.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why is detecting shortcuts difficult with explainable AI?",{"text":80,"@type":76},"The paper shows that even simple shortcuts can remain hard to detect by explainable AI methods when in-domain test performance stays good. Practitioners may also miss shortcut presence due to how training and test splits mask the lack of real generalization.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the proposed approach protect publicly available datasets from crawlers?",{"text":84,"@type":76},"It deliberately adds ML shortcuts to augmented datasets so they become unusable for ML applications. This deters large-scale unauthorized crawling while keeping the visual presentation of the data essentially unchanged for normal viewing.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]