[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-125680-en":3,"doc-seo-125680-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},125680,8796095360427,"Lucas Martin","https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d",8,"Research & Report","Novel Algorithm-Level Approaches for Class-Imbalanced Machine Learning - Doctoral dissertation","Machine learning classifiers typically assume balanced class instance counts, yet real-world data often show severe imbalance. This dissertation investigates neural-network adaptations that remain robust on imbalanced datasets without requiring data manipulation and that can integrate with any model architecture or framework. Two complementary algorithm-level approaches are proposed: novel loss functions based on an approximated confusion matrix and a thresholded output activation mechanism. Methods target reduced false negatives, and are extended from binary to multi-class settings.","Novel Algorithm-Level Approaches for Class-Imbalanced Machine  \nLearning  \nDavid Twomey  \nA dissertation submitted in partial fulfillment of the requirements for the degree of  \nDoctor of Philosophy  \nof  \nUniversity College London.  \nDepartment of Computer Science  \nUniversity College London  \nMarch 27, 2023  \n2  \nI, David Twomey, confirm that the work presented in this thesis is my own. Where information has been derived from other sources, I confirm that this has been  \nindicated in the work.  \nAbstract  \nMachine learning classifiers are designed with the underlying assumption of a roughly balanced number of instances per class. However, in many real-world applications this is far from true. This thesis explores adaptations of neural networks which are robust to class imbalanced datasets, do not involve data manipulation, and are flexible enough to be used with any model architecture or framework. The thesis explores two complementary approaches to the problem of class imbalance. The first exchanges conventional choices of classification loss function, which are fundamentally measures of how far network outputs are from desired ones, for ones that instead primarily register whether outputs are right or wrong. The construction of these novel loss functions involves the concept of an approximated confusion matrix, another use of which is to generate new performance metrics, especially useful for monitoring validation behaviour for imbalanced datasets. The second approach changes the form of the output layer activation function to one with a threshold which can be learned so as to more easily classify the more difficult minority class. These two approaches can be used together or separately, with the combined technique being a promising approach for cases of extreme class imbalance. While the methods are developed primarily for binary classification scenarios, as these are the most numerous in the applications literature, the novel loss functions introduced here are also demonstrated to be extensible to a multi-class scenario.  \nImpact Statement  \nClassification is arguably the most common current application area of machine learning. While in many cases, such as recommender systems (“is this movie going to appeal to the subscriber, or not?”), errors are not of much importance, there are other, such as medical diagnosis and screening, where errors, in particular ones in which a positive (in general, having the property that the classifier seeks to identify) instance that is missed (referred to as a false negative or Type-II error) may have serious consequences. However, classifier systems tend to frequently display these types of errors when the number of positive cases is small compared to the number of negative ones. This thesis addresses the problem of classification in such imbalanced scenarios, and presents novel methods that help avoid false negativesin these situations. The novel methods of this thesis are of two kinds: the first changes the representation of the classification problem so that learning is more aggressively directed toward the correct classification of both negative (majority) and positive (minority) examples, and the second helps the system to more easily correctly classify the minority type. Taken together, these novel tools are a significant extension to the current toolkit for classification in machine learning that has also, as has been emphasised above, the potential for a substantial practical value.  \nAcknowledgements  \nI would like to thank Dr. Denise Gorse who has taught me so much regarding the skill of research and the coherent presentation of ideas. She has been incredibly generous and supportive in her role as my supervisor and has shown a genuine interest in helping me to achieve my goals.  \nI would also like to express my sincere gratitude to my parents, aunt Pat, and partner Lorna, who, among many others, have shown me remarkable support and  \nencouragement throughout this process.  \nCon","cbCaifWnAsWI1LNT","https://ap.wps.com/l/cbCaifWnAsWI1LNT","pdf",2826364,1,153,"English","en",105,"# Contents\n## 1 Introduction\n## 2 Background\n## 3 Geometric Mean as a Classification-Focused Training Loss\n## 4 Classification-Focused Training Metrics: Alternatives & Extension to Multi-Class Scenarios","[{\"question\":\"What problem does this dissertation focus on?\",\"answer\":\"It focuses on classification performance when classes are imbalanced, especially scenarios where positive cases are rare compared to negatives and false negatives can have serious consequences.\"},{\"question\":\"What are the two algorithm-level approaches proposed?\",\"answer\":\"The first replaces conventional classification losses with new ones that primarily reflect whether outputs are right or wrong, using an approximated confusion matrix and related metrics. The second modifies the output layer activation with a learnable threshold to better separate the minority class.\"},{\"question\":\"Does the work apply only to binary classification?\",\"answer\":\"Although the methods are developed and demonstrated mainly for binary classification, the introduced loss functions are also shown to be extensible to multi-class scenarios.\"}]","Novel Algorithm-Level Approaches for Class-Imbalanced Machine Learning - Doctoral dissertation | PDF",1785900616,386,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"novel-algorithm-level-approaches-for-class-imbalanced-machine-learning-doctoral-dissertation","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/novel-algorithm-level-approaches-for-class-imbalanced-machine-learning-doctoral-dissertation/125680/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does this dissertation focus on?","Question",{"text":75,"@type":76},"It focuses on classification performance when classes are imbalanced, especially scenarios where positive cases are rare compared to negatives and false negatives can have serious consequences.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What are the two algorithm-level approaches proposed?",{"text":80,"@type":76},"The first replaces conventional classification losses with new ones that primarily reflect whether outputs are right or wrong, using an approximated confusion matrix and related metrics. The second modifies the output layer activation with a learnable threshold to better separate the minority class.",{"name":82,"@type":73,"acceptedAnswer":83},"Does the work apply only to binary classification?",{"text":84,"@type":76},"Although the methods are developed and demonstrated mainly for binary classification, the introduced loss functions are also shown to be extensible to multi-class scenarios.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]