[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-123661-en":3,"doc-seo-123661-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},123661,962075006959,"Anda","https://ap-avatar.wpscdn.com/avatar/e0002397efbe92a78e?_k=1776741047341049297",8,"Research & Report","A Generic Machine Learning Framework for Fully-Unsupervised Anomaly Detection with Contaminated Data - Research","Anomaly detection models commonly rely on residual-based learning from normal data and score unseen samples by dissimilarity to the learned normal regime. In real operations, training sets are often contaminated by a fraction of abnormal samples, which degrades residual-based anomaly detection performance. This paper proposes a generic, fully unsupervised refinement framework that removes candidate anomalies from contaminated training data in a single step. Experiments on multivariate time series datasets show clear improvements over training on contaminated data without refinement and strong competitiveness against an ideal anomaly-free reference.","A Generic Machine Learning Framework for Fully-Unsupervised Anomaly Detection with Contaminated Data  \narXiv :2308 . 13352v2 [ cs .LG] 7 Sep 2023  \nMarkus Ulmer 1 , Jannik Zgraggen2 , and Lilach Goren Huber3  \n1,2,3 Zurich University of Applied Sciences, Winterthur 8401, Switzerland [markus.ulmer@zhaw.ch](markus.ulmer@zhaw.ch)  \n[jannik.zgraggen@zhaw.ch](jannik.zgraggen@zhaw.ch)  \n[lilach.gorenhuber@zhaw.ch](lilach.gorenhuber@zhaw.ch)  \nABSTRACT  \nAnomaly detection (AD) tasks have been solved using machine learning algorithms in various domains and applications. The great majority of these algorithms use normal data to train a residual-based model, and assign anomaly scores to unseen samples based on their dissimilarity with the learned normal regime. The underlying assumption of these approaches is that anomaly-free data is available for training. This is, however, often not the case in real-world operational settings, where the training data may be contaminated with a certain fraction of abnormal samples. Training with contaminated data, in turn, inevitably leads to a deteriorated AD performance of the residual-based algorithms.  \nIn this paper we introduce a framework for a fully unsupervised refinement of contaminated training data for AD tasks. The framework is generic and can be applied to any residualbased machine learning model. We demonstrate the application of the framework to two public datasets of multivariate time series machine data from different application fields. We show its clear superiority over the naive approach of training with contaminated data without refinement. Moreover, we compare it to the ideal, unrealistic reference in which anomaly-free data would be available for training. Since the approach exploits information from the anomalies, and not only from the normal regime, it is comparable and often outperforms the ideal baseline as well.  \n1. INTRODUCTION  \nAnomaly detection (AD) tasks are common in very diverse fields, including medical image processing, autonomous driving, fraud detection, and fault detection in industrial machines. An inherent property of AD tasks is that very few or no labelled examples of anomalous behavior are provided in advance.  \nTherefore, the most common machine learning approaches to solve these tasks are based on using exclusively normal  \ndata to train a selected prediction algorithm, which is subsequently used to infer on unseen data. The hidden assumption here is that the algorithm’s prediction errors (residuals) will be higher whenever the input sample does not belong to the learned distribution. This family of models can be referred to as residual-based models, irrespective of whether they use regression or reconstruction residuals to detect anomalies, with the most common models being reconstruction models such as principal component analysis (PCA) or different types of autoencoder (AE) neural networks. It is worth noting that these models are often termed ”unsupervised”. However, in the context of AD, they should be referred to as ”semi supervised” since they assume the availability of labeled normal data for training.  \nIn fact, such models tend to perform rather poorly whenever contamination in the form of anomalous samples is introduced into the training data. In real world applications, however, the assumption of having anomaly-free training data does not always hold, as data contamination cannot be avoided. In this case, truly unsupervised methods, based on clustering or one-class classification (Schlkopf, Williamson, Smola, Shawe-Taylor, & Platt, 1999), are required. Recently there has been a growing effort to develop deep unsupervised algorithms for AD, whose performance is not severely damaged by data contamination (Munir, Siddiqui, Dengel, & Ahmed, 2018) .  \nDespite the high practical relevance of the problem, systematic solutions for AD with contaminated training data are rather rare. Recently, several papers have suggested useful approaches based on d","cbCaifWQrjhuI3Cl","https://ap.wps.com/l/cbCaifWQrjhuI3Cl","pdf",1723341,1,11,"English","en",105,"# Abstract\n# Introduction\n## Residual-based anomaly detection and its assumptions\n## Limitations under contaminated training data\n## Fully unsupervised anomaly detection approaches\n# Proposed framework for refinement\n## Single-step refinement strategy\n## Application to residual-based models\n# Experimental evaluation\n## Industrial time series datasets\n## Comparison against naive and ideal baselines","[{\"question\":\"Why do residual-based anomaly detection methods often fail with contaminated training data?\",\"answer\":\"They assume training data is normal and learn a residual pattern for the normal regime; abnormal contamination breaks this assumption and deteriorates anomaly scores.\"},{\"question\":\"What does the proposed framework do differently?\",\"answer\":\"It performs a single-step refinement on contaminated training data by proposing candidate anomalies to remove, enabling fully unsupervised training for residual-based models.\"},{\"question\":\"How is the framework evaluated and what baselines are used?\",\"answer\":\"It is tested on two multivariate time series datasets and compared against training a residual-based model on contaminated data without refinement, and against an ideal anomaly-free training reference.\"}]","A Generic Machine Learning Framework for Fully-Unsupervised Anomaly Detection with Contaminated Data - Research | PDF",1785817889,28,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"a-generic-machine-learning-framework-for-fully-unsupervised-anomaly-detection-with-contaminated-data-research","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/a-generic-machine-learning-framework-for-fully-unsupervised-anomaly-detection-with-contaminated-data-research/123661/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05","2026-08-04",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why do residual-based anomaly detection methods often fail with contaminated training data?","Question",{"text":76,"@type":77},"They assume training data is normal and learn a residual pattern for the normal regime; abnormal contamination breaks this assumption and deteriorates anomaly scores.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What does the proposed framework do differently?",{"text":81,"@type":77},"It performs a single-step refinement on contaminated training data by proposing candidate anomalies to remove, enabling fully unsupervised training for residual-based models.",{"name":83,"@type":74,"acceptedAnswer":84},"How is the framework evaluated and what baselines are used?",{"text":85,"@type":77},"It is tested on two multivariate time series datasets and compared against training a residual-based model on contaminated data without refinement, and against an ideal anomaly-free training reference.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]