[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118347-en":3,"doc-seo-118347-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},118347,13056703019404,"Miles","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Tolerant Machine Learning for Deficient Training Data - Doctor of Philosophy Thesis","Supervised machine learning can automate manual data classification, but degraded training datasets often reduce model accuracy. This thesis develops tolerant machine learning methods that remain effective despite common training data deficiencies, enabling users to obtain value proportional to data quality and making new applications more cost-effective. It addresses missing class labels via the WITAN algorithm, dataset shift through a Gain-Some-Lose-Some (GSLS) quantification model with decision-tree selection, and class noise using a model-agnostic null-labelling rejection approach that improves the error–rejection trade-off. The work also provides an integrated software architecture.","TOLERANT MACHINE LEARNING FOR DEFICIENT TRAINING DATA  \nA THESIS SUBMITTED TO AUCKLAND UNIVERSITY OF TECHNOLOGY  \nIN PARTIAL FULFILMENT OF THE REQUIREMENTS FOR THE DEGREE OF  \nDOCTOR OF PHILOSOPHY  \nSupervisors  \nProfessor Edmund M-K Lai  \nAssociate Professor Roopak Sinha  \nProfessor Muhammad Asif Naeem  \nNovember 2022  \nBy  \nBenjamin James Denham  \nSchool of Engineering, Computer and Mathematical Sciences  \nAbstract  \nWhile supervised machine learning methods have shown great potential for automating time-consuming manual data classiﬁcation tasks, their application is often hindered by deﬁciencies in training datasets that degrade model accuracy. This thesis proposes tolerant machine learning methods that can be applied despite common training data deﬁciencies. By enabling users to derive value proportional to data quality, tolerant machine learning makes exploring new machine learning applications more cost-effective.  \nThe ﬁrst training data deﬁciency addressed by this thesis is a lack of class labels for training data, including when the set of possible classes is unknown. Prior work in the weak supervision paradigm of data programming has sought to provide large quantities of data labels through user-deﬁned heuristic labelling functions. Despite the development of methods to assist users in deﬁning such functions, users must still have a small labelled dataset or at least upfront knowledge of the set of possible classes. The WITAN algorithm proposed in this thesis can suggest labelling functions without any initial supervision, supporting the user to discover classes progressively. Experiments with binary and multi-class datasets demonstrate WITAN's competitive efﬁciency and  \naccuracy compared to alternative labelling methods, despite its lack of supervision.  \nThe second training data deﬁciency addressed by this thesis is the problem of dataset shift, where the data distribution of the training dataset differs from that of a target population. Dataset shift is a challenging yet expected problem when estimating the prevalences of classes in different target samples—so-called quantiﬁcation. Existing  \nquantiﬁcation methods make strong assumptions on the nature of dataset shift that may not hold in practice. This thesis proposes a Gain-Some-Lose-Some (GSLS) quantiﬁcation model that is experimentally demonstrated to provide more reliable quantiﬁcation prediction intervals than existing methods under more general conditions of shift. GSLS is integrated into a decision tree for dynamically selecting an appropriate quantiﬁcation method for a given target sample. Selection by a Kolmogorov-Smirnov test for any shift followed by a newly proposed “Adjusted Kolmogorov-Smirnov” test for non-prior shift is found to best balance quantiﬁcation and runtime performance. Additionally, a framework is presented for constraining quantiﬁcation prediction intervals to user-speciﬁed limits by requesting class labels from the user for smaller sets of instances than would be required with rejection of classiﬁcations based on conﬁdence scores alone.  \nThe third and ﬁnal training data deﬁciency addressed by this thesis is the problem of class noise. When a class noise process distorts the relationships between input features and class labels or true class values, a classiﬁer should reject instances for which it cannot provide a conﬁdent classiﬁcation. This thesis demonstrates how the popular model-agnostic conﬁdence-thresholding rejection method does not leverage relationships between input features and class noise. A novel model-agnostic null-labelling rejection method is proposed to learn such relationships, and an experimental evaluation demonstrates its ability to achieve a signiﬁcantly better trade-off between classiﬁcation error and the rate of rejection under an evaluation framework that uniﬁes prior theories for combining rejecting-classiﬁers. Additionally, null-labelling is demonstrated to enable users to understand relationships between i","cbCaivD9WNodVDIq","https://ap.wps.com/l/cbCaivD9WNodVDIq","pdf",2381975,1,194,"English","en",105,"# Abstract\n# Attestation of Authorship\n# Publications\n# Acknowledgements\n# Intellectual Property Rights\n# 1 Introduction\n## Motivation\n## Research Objectives\n## Contributions\n## Thesis Structure\n# 2 Suggesting Labelling Functions without Supervision for Low-Effort Data Programming\n## Introduction\n## Literature Review\n## The Proposed Algorithm\n## Multi-class Extension of IWS\n## Experimental Study\n## Conclusion","[{\"question\":\"What training data deficiencies does the thesis focus on?\",\"answer\":\"The thesis addresses three core deficiencies: missing class labels, dataset shift, and class noise that corrupts the relationship between features and labels.\"},{\"question\":\"How does WITAN help when no labels or known class sets are available?\",\"answer\":\"WITAN suggests labelling functions without any initial supervision, allowing users to discover classes progressively while supporting low-effort data programming.\"},{\"question\":\"What is the GSLS approach for dataset shift and quantification?\",\"answer\":\"GSLS provides quantification prediction intervals that are more reliable under more general shift conditions, and a framework integrates method selection using statistical testing, with optional user-constrained interval refinement.\"}]","Tolerant Machine Learning for Deficient Training Data - Doctor of Philosophy Thesis | PDF",1785683210,489,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"tolerant-machine-learning-for-deficient-training-data-doctor-of-philosophy-thesis","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/tolerant-machine-learning-for-deficient-training-data-doctor-of-philosophy-thesis/118347/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What training data deficiencies does the thesis focus on?","Question",{"text":75,"@type":76},"The thesis addresses three core deficiencies: missing class labels, dataset shift, and class noise that corrupts the relationship between features and labels.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does WITAN help when no labels or known class sets are available?",{"text":80,"@type":76},"WITAN suggests labelling functions without any initial supervision, allowing users to discover classes progressively while supporting low-effort data programming.",{"name":82,"@type":73,"acceptedAnswer":83},"What is the GSLS approach for dataset shift and quantification?",{"text":84,"@type":76},"GSLS provides quantification prediction intervals that are more reliable under more general shift conditions, and a framework integrates method selection using statistical testing, with optional user-constrained interval refinement.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]