[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-122125-en":3,"doc-seo-122125-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},122125,1099514068035,"Ezra","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Toward the Creation of Precise Synthetic Datasets for Trustworthy and Privacy-preserving Machine Learning - Master Thesis","With the development of artificial intelligence, the importance of data in machine learning keeps increasing, while restrictions on disclosure and rising privacy concerns create a growing contradiction. Synthetic data, artificially generated to resemble real data distributions, helps protect sensitive information and supports research needs. The work targets efficient generation of precise, privacy-preserving synthetic datasets by improving augmentation for balanced data and designing generation from limited statistics. It proposes BCTGAN and BPCA, validated on five datasets.","E ES – Commun icat ion Systems Group, Prof. Dr. Burkhard Stiller  \nToward the Creation of Precise Synthetic Datasets for Trustworthy and Privacy-preserving Machine  \nLearning  \nJingjing Li  \nZurich, Switzerland  \nStudent ID: 21-738-091  \nSupervisor: Dr. Alberto Huertas Celdran, Weijie Niu Date of Submission: March 11, 2024  \nUniversity of Zurich  \nDepartment of Informatics (IFI) Binzmühlestrasse 14, CH-8050 Zürich, Switzerland  \nifi  \nMaster Thesis  \nCommunication Systems Group (CSG)  \nDepartment of Informatics (IFI) University of Zurich  \nBinzmühlestrasse 14, CH-8050 Zürich, Switzerland URL: [http://www.csg.uzh.ch/](http://www.csg.uzh.ch/)  \nAbstract  \nWith the development of artificial intelligence, the importance of data in machine learning is gradually increasing. Data plays an important role in machine learning models’ training and testing. Besides, concerns about data privacy are also gaining more attention. The contradiction between the demand for data and the restrictions on data disclosure is increasing. In this case, synthetic data is a reliable way to resolve this problem. Synthetic data is artificially created data that resembles the distribution of real data. It not only protects the privacy of real data, but also creates large amounts of data that satisfy research requirements. But how to efficiently generate precise synthetic data with privacy protection is still a problem.  \nThis paper aims to achieve two goals. The first goal is to optimize existing data augmentation algorithms to augment balanced synthetic data. The second goal is to design a data generation algorithm that can generate data using limited information.  \nThis paper ultimately developed two algorithms: BCTGAN and BPCA. The BCTGAN algorithm effectively achieves the first goal by generating balanced data. The BPCA algorithm accomplishes the second goal by generating data using only mean, variance, and covariance. Experimental validation is conducted on five datasets, assessing the quality of synthetic data in terms of machine learning usability and privacy preservability.  \nThe synthetic dataset augmented by BCTGAN performs better than CTGAN on classifying minority classes and reaches an average F1-score of 0 .705. The synthetic dataset generated by BPCA performs an average 0.683 F1-score, outperforming the baseline dataset generated by the Cholesky Decomposition method by 0 .011.  \nFor future work, enhancing the stability of the algorithms, optimizing the time and space complexity, and trying to ensure privacy when performing data augmentation are all directions worth trying.  \nii  \nZusammenfassung  \nMit der Entwicklung der ku¨nstlichen Intelligenz nimmt die Bedeutung von Daten beim maschinellen Lernen allma¨hlich zu. Daten spielen eine wichtige Rolle beim Trainieren und Testen von Modellen des maschinellen Lernens. Außerdem gewinnen Bedenken hinsichtlich des Datenschutzes immer mehr an Bedeutung. Der Widerspruch zwischen der Nachfrage nach Daten und den Beschra¨nkungen fu¨r die Offenlegung von Daten wird immer gro¨ßer. In diesem Fall sind synthetische Daten ein zuverla¨ssiger Weg, dieses Problem zu lo¨sen. Synthetische Daten sind ku¨nstlich erzeugte Daten, die der Verteilung realer Daten a¨hneln. Dadurch wird nicht nur die Privatspha¨re echter Daten geschu¨tzt, sondern es werden auch große Datenmengen erzeugt, die den Anforderungen der Forschung entsprechen. Ein Problem bleibt jedoch die effiziente Erzeugung pra¨ziser synthetischer Daten unter Wahrung der Privatspha¨re.  \nMit dieser Arbeit sollen zwei Ziele erreicht werden. Das erste Ziel ist die Optimierung bestehender Algorithmen zur Datenerweiterung, um ausgewogene synthetische Daten zu erweitern. Das zweite Ziel besteht darin, einen Algorithmus zur Datengenerierung zu entwickeln, der Daten mit begrenzten Informationen erzeugen kann.  \nIn dieser Arbeit wurden schließlich zwei Algorithmen entwickelt: BCTGAN und BPCA. Der BCTGAN-Algorithmus erreicht das erste Ziel effektiv, indem er ausgewogene","cbCaimYL0SOQ7u39","https://ap.wps.com/l/cbCaimYL0SOQ7u39","pdf",7919905,1,92,"English","en",105,"# Introduction\n## Motivation\n## Description of Work\n## Thesis Outline\n# Background\n## Data Types\n## Comparison of data augmentation","[{\"question\":\"Why are synthetic datasets used in machine learning?\",\"answer\":\"Synthetic datasets resemble the distribution of real data, helping protect the privacy of sensitive information while still providing enough data for training and evaluation.\"},{\"question\":\"What are the two main goals of the thesis?\",\"answer\":\"The thesis aims to optimize existing data augmentation methods to produce balanced synthetic data, and to design a data generation algorithm that works using limited information.\"},{\"question\":\"How do BCTGAN and BPCA contribute to the goals?\",\"answer\":\"BCTGAN generates balanced synthetic data to improve performance on minority-class classification, while BPCA generates data using mean, variance, and covariance to enable generation with limited statistical inputs.\"}]","Toward the Creation of Precise Synthetic Datasets for Trustworthy and Privacy-preserving Machine Learning - Master Thesis | PDF",1785808932,232,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"toward-the-creation-of-precise-synthetic-datasets-for-trustworthy-and-privacy-preserving-machine-learning-master-thesis","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/toward-the-creation-of-precise-synthetic-datasets-for-trustworthy-and-privacy-preserving-machine-learning-master-thesis/122125/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why are synthetic datasets used in machine learning?","Question",{"text":75,"@type":76},"Synthetic datasets resemble the distribution of real data, helping protect the privacy of sensitive information while still providing enough data for training and evaluation.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What are the two main goals of the thesis?",{"text":80,"@type":76},"The thesis aims to optimize existing data augmentation methods to produce balanced synthetic data, and to design a data generation algorithm that works using limited information.",{"name":82,"@type":73,"acceptedAnswer":83},"How do BCTGAN and BPCA contribute to the goals?",{"text":84,"@type":76},"BCTGAN generates balanced synthetic data to improve performance on minority-class classification, while BPCA generates data using mean, variance, and covariance to enable generation with limited statistical inputs.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]