[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-123894-en":3,"doc-seo-123894-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},123894,137441390410,"Hazel","https://ap-avatar.wpscdn.com/avatar/2000252f4ab5702993?_k=1776741390130283984",8,"Research & Report","Data augmentation with automated machine learning - approaches and performance comparison with classical data augmentation methods - survey","Data augmentation is a key regularization technique for improving machine learning generalization by applying transformations that create new samples with desired properties. Manual exploration of candidate augmentations and their hyperparameters is often time-consuming and labor-intensive. AutoML-based data augmentation addresses this challenge by automating the design, optimization, and evaluation of augmentation strategies. The work surveys AutoML-driven approaches, emphasizing image augmentation but also covering other modalities such as tabular data via suitable data integration methods, and analyzes performance against classical augmentation.","arXiv :2403 .08352v 3 [ cs .LG] 5 Mar 2025  \nData augmentation with automated machine learning: approaches and performance comparison with classical data augmentation  \nmethods  \nAlhassan Mumuni 1 ∗ and Fuseini Mumuni2  \nAbstract—  \nData augmentation is arguably the most important regularization technique commonly used to improve generalization performance of machine learning models. It primarily involves the application of appropriate data transformation operations to create new data samples with desired properties. Despite its effectiveness, the process is often challenging because of the time-consuming trial and error procedures for creating and testing different candidate augmentations and their hyperparameters manually. State-of-the-art approaches are increasingly relying on automated machine learning (AutoML) principles. This work presents a comprehensive survey of AutoMLbased data augmentation techniques. We discuss various approaches for accomplishing data augmentation with AutoML, including data manipulation, data integration and data synthesis techniques. The focus of this work is on image data augmentation methods. Nonetheless, we cover other data modalities, especially in cases where the specific data augmentations techniques being discussed are more suitable for these other modalities. For instance, since automated data integration methods are more suitable for tabular data, we cover tabular data in the discussion of data integration methods. The work also presents extensive discussion of techniques for accomplishing each of the major subtasks of the image data augmentation process: search space design, hyperparameter optimization and model evaluation. Finally, we carried out an extensive comparison and analysis of the performance of automated data augmentation techniques and state-of-the-art methods based on classical augmentation approaches. The results show that AutoML methods for data augmentation currently outperform state-of-the-art techniques based on conventional approaches.  \nIndex Terms—Data augmentation, AutoML, automated machine learning, machine learning, data preparation, image augmentation.  \n~~ ~~ ✦ ~~ ~~  \n1 INTRODUCTION  \n1.1 Background  \nPractical implementations of machine learning systems require large data samples to produce satisfactory results. Since data is often not available in sufficient quantities, regularization techniques are critical for achieving good performance. These techniques commonly entail tweaking the machine learning model configuration or applying data augmentation—a range of methods for extending the available data by applying appropriate transformations. The basic idea is to modify training datasets by applying suitable transformations in ways that increase the quantity, representation quality and variability of the original data.  \nThe most commonly used data augmentation techniques include geometric transformations – particularly, rotation, flipping, shearing and scaling–and photometric transformations such as color jittering, solarizaion, brightness, contrast adjustment, noise addition, denoising, and color space conversion. Data augmentation can also involve creating completely new data from scratch [1],[2] . This approach can be useful when the training data for the target application  \n• 1Alhassan Mumuni: Department of Electrical and Electronics Engineering, Cape Coast Technical University, Cape Coast, Ghana.  \n∗ Corresponding author, E-mail: [alhassan.mumuni@cctu.edu.gh](alhassan.mumuni@cctu.edu.gh)  \n• 2Fuseini Mumuni: University of Mines and Technology, UMaT, Tarkwa, [Ghana. E-mail: fmumuni@umat.edu.gh](Ghana. E-mail: fmumuni@umat.edu.gh)  \nis inaccessible [3] . Methods for synthetic data generation include explicitly creating samples with desired data distribution using computer graphics tools ([2]) or algorithmically generating artificial data with the aid of special deep learning models (e.g., with techniques such as differential neural rendering [4], [5] . ","cbCaignIydfovjGe","https://ap.wps.com/l/cbCaignIydfovjGe","pdf",2625565,1,29,"English","en",105,"# Introduction\n## Background\n## Limitations of classical data augmentation approaches","[{\"question\":\"Why is data augmentation important in machine learning?\",\"answer\":\"Data augmentation improves generalization by increasing data quantity, representation quality, and variability through appropriate transformations.\"},{\"question\":\"What makes classical data augmentation methods difficult to use effectively?\",\"answer\":\"Classical methods rely on laborious manual trial-and-error, involve a combinatorial search over augmentation settings, and results may not transfer well across datasets or tasks.\"},{\"question\":\"How does AutoML improve the data augmentation process?\",\"answer\":\"AutoML automates selecting augmentation strategies by addressing major subtasks such as search space design, hyperparameter optimization, and model evaluation, and it can outperform conventional methods.\"}]","Data augmentation with automated machine learning - approaches and performance comparison with classical data augmentation methods - survey | PDF",1785819110,73,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"data-augmentation-with-automated-machine-learning-approaches-and-performance-comparison-with-classical-data-augmentation-methods-survey","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/data-augmentation-with-automated-machine-learning-approaches-and-performance-comparison-with-classical-data-augmentation-methods-survey/123894/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is data augmentation important in machine learning?","Question",{"text":75,"@type":76},"Data augmentation improves generalization by increasing data quantity, representation quality, and variability through appropriate transformations.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What makes classical data augmentation methods difficult to use effectively?",{"text":80,"@type":76},"Classical methods rely on laborious manual trial-and-error, involve a combinatorial search over augmentation settings, and results may not transfer well across datasets or tasks.",{"name":82,"@type":73,"acceptedAnswer":83},"How does AutoML improve the data augmentation process?",{"text":84,"@type":76},"AutoML automates selecting augmentation strategies by addressing major subtasks such as search space design, hyperparameter optimization, and model evaluation, and it can outperform conventional methods.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]