[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-122198-en":3,"doc-seo-122198-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},122198,7971461740909,"Levi","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Tabular Machine Learning on Small-Size and High-Dimensional Data - Thesis Summary","The thesis presents four novel methods to improve the generalisation of machine learning models on small-size and high-dimensional tabular datasets. It addresses key obstacles caused by limited samples and the curse of dimensionality, which increase overfitting and weaken performance. Two model-centric approaches constrain parameters via shared auxiliary networks (WPFS and GCondNet) to reduce degrees of freedom. Two data-centric augmentation methods generate synthetic data (TabEBM and TabMDA) to enlarge diversity without additional training, improving classification on small datasets.","Tabular Machine Learning on Small-Size and High-Dimensional Data  \nAndrei Margeloiu  \nSelwyn College  \nThis thesis is submitted in November 2024 for the degree of Doctor of Philosophy  \nDeclaration  \nThis thesis is the result of my own work and includes nothing which is the outcome of work done in collaboration except as declared in the preface and specified in the text. It is not substantially the same as any work that has already been submitted, or is being concurrently submitted, for any degree, diploma or other qualification at the University of Cambridge or any other University or similar institution except as declared in the preface and specified in the text. It does not exceed the prescribed word limit for the relevant Degree Committee.  \nAndrei Margeloiu November 2024  \nTabular Machine Learning on Small-Size and High-Dimensional Data Andrei Margeloiu  \nThis thesis introduces four novel methods to improve the generalisation of machine learning models on small-size and high-dimensional tabular datasets. Tabular data – tables where each row represents an individual record and each column represents features – is ubiquitous in critical fields such as medicine, scientific research and finance. However, these areas often face data scarcity due to difficulties in data acquisition, making it challenging to gather large sample sizes. Conversely, new data collection technologies enable the collection of high-dimensional data, leading to datasets where the number of features greatly exceeds the number of samples. Data scarcity and high dimensionality present significant challenges for machine learning models, primarily due to the increased risk of overfitting arising from the curse of dimensionality and the limited data available to adequately represent the underlying distribution. Existing approaches often struggle to generalise effectively in such scenarios, resulting in suboptimal performance. As a result, training models on small and high-dimensional datasets requires methods designed to address these limitations and generalise more effectively from limited data.  \nWe introduce two new model-centric approaches to address overfitting in neural networks trained on small-size and high-dimensional data. Our key innovation is to mitigate overfitting by constraining model parameters through shared auxiliary networks, which capture latent relationships in tabular data to partially determine the predictor model’s parameters, thereby reducing its degrees of freedom. Firstly, we introduce WPFS, a parameter-efficient architecture that imposes hard parameter-sharing on the model’s parameters using weight predictor networks. Secondly, we introduce GCondNet, a method that uses Graph Neural Networks (GNNs) to facilitate soft parameter-sharing in an underlying predictor model. When applied to biomedical tabular datasets, these methods demonstrate improved predictive performance, primarily by reducing overfitting.  \nWhile relying solely on model-centric approaches is common, integrating data-centric methods can provide additional performance gains, particularly in data-scarce tasks. To this end, we introduce two novel data augmentation methods that generate synthetic data to increase the size and diversity of the training set, capturing more variability of the underlying data distribution. Our key innovation is transforming pre-trained tabular classifiers into data generators, leveraging their pre-training information in two  \nnovel ways. The first method, TabEBM, constructs dedicated class-specific Energy-Based Models (EBMs) to approximate class-conditional distributions, which are then used to generate additional training data. The second method, TabMDA, introduces in-context subsetting (ICS), a technique that enables label-invariant transformations within the manifold space learned by pre-trained in-context classifiers, effectively enlarging the training dataset. Both methods are general, fast, require no additional training, and can be ap","cbCaiuYIU01XuVaQ","https://ap.wps.com/l/cbCaiuYIU01XuVaQ","pdf",7625408,1,212,"English","en",105,"# Introduction\n## Background: tabular data and key challenges\n## Thesis contributions: four proposed methods\n# Model-centric approaches\n## WPFS: parameter-efficient hard sharing\n## GCondNet: soft sharing with graph neural networks\n# Data-centric augmentation methods\n## TabEBM: class-specific energy-based models\n## TabMDA: in-context subsetting for label-invariant transformations\n# Applications and impact","[{\"question\":\"What problem does the thesis focus on for tabular datasets?\",\"answer\":\"It focuses on improving generalisation when tabular data are both small in size and high-dimensional, where overfitting is more likely and performance often becomes suboptimal.\"},{\"question\":\"How do the model-centric methods reduce overfitting?\",\"answer\":\"They mitigate overfitting by constraining model parameters through shared auxiliary networks, which capture latent relationships and reduce the model’s effective degrees of freedom (WPFS, GCondNet).\"},{\"question\":\"How do the data-centric methods improve performance without additional training?\",\"answer\":\"They generate synthetic training data by transforming pre-trained tabular classifiers into data generators (TabEBM and TabMDA). These methods are general, fast, require no extra training, and consistently boost classification on small datasets.\"}]","Tabular Machine Learning on Small-Size and High-Dimensional Data - Thesis Summary | PDF",1785809310,534,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"tabular-machine-learning-on-small-size-and-high-dimensional-data-thesis-summary","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/tabular-machine-learning-on-small-size-and-high-dimensional-data-thesis-summary/122198/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the thesis focus on for tabular datasets?","Question",{"text":75,"@type":76},"It focuses on improving generalisation when tabular data are both small in size and high-dimensional, where overfitting is more likely and performance often becomes suboptimal.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How do the model-centric methods reduce overfitting?",{"text":80,"@type":76},"They mitigate overfitting by constraining model parameters through shared auxiliary networks, which capture latent relationships and reduce the model’s effective degrees of freedom (WPFS, GCondNet).",{"name":82,"@type":73,"acceptedAnswer":83},"How do the data-centric methods improve performance without additional training?",{"text":84,"@type":76},"They generate synthetic training data by transforming pre-trained tabular classifiers into data generators (TabEBM and TabMDA). These methods are general, fast, require no extra training, and consistently boost classification on small datasets.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]