[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-121910-en":3,"doc-seo-121910-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},121910,549758252649,"Ivy","https://ap-avatar.wpscdn.com/avatar/8000253669c5317157?_k=1778319167496531819",8,"Research & Report","Robust Machine Learning - by Transforming and Augmenting Imperfect Training Data - A Thesis","Machine Learning (ML) converts data into programs for prediction or optimal control, but real-world training datasets are often imperfect and induce fragile behaviors after deployment. This thesis examines multiple data sensitivities behind such failures and proposes remedies. It addresses preventing ML from encoding discriminatory signals via fair representation learning, handling spurious features by identifying partitions that reveal inconsistency, and reinforcement learning under insufficient state-action coverage using causal priors for counterfactual data augmentation.","Robust Machine Learning by Transforming and Augmenting Imperfect Training Data  \narXiv :2312 . 12597v1 [ cs .LG] 19 Dec 2023  \nby  \nElliot Creager  \nA thesis submitted in conformity with the requirements for the degree of Doctor of Philosophy Department of Computer Science University of Toronto  \n© Copyright 2023 by Elliot Creager  \nRobust Machine Learning by Transforming and Augmenting Imperfect Training Data  \nElliot Creager  \nDoctor of Philosophy  \nDepartment of Computer Science  \nUniversity of Toronto  \n2023  \nAbstract  \nMachine Learning (ML) is an expressive framework for turning data into computer programs. Across many problem domains—both in industry and policy settings—the types of computer programs needed for accurate prediction or optimal control are difficult to write by hand. On the other hand, collecting instances of desired system behavior may be relatively more feasible. This makes ML broadly appealing, but also induces data sensitivities that often manifest as unexpected failure modes during deployment. In this sense, the training data available tend to be imperfect for the task at hand. This thesis explores several data sensitivities of modern machine learning and how to address them.  \nWe begin by discussing how to prevent ML from codifying prior human discrimination measured in the training data, where we take a fair representation learning approach. We then discuss the problem of learning from data containing spurious features, which provide predictive fidelity during training but are unreliable upon deployment. Here we observe that insofar as standard training methods tend to learn such features, this propensity can be leveraged to search for partitions of training data that expose this inconsistency, ultimately promoting learning algorithms invariant to spurious features. Finally, we turn our attention to reinforcement learning from data with insufficient coverage over all possible states and actions. To address the coverage issue, we discuss how causal priors can be used to model the single-step dynamics of the setting where data are collected. This enables a new type of data augmentation where observed trajectories are stitched together to produce new but plausible counterfactual trajectories.  \nThis manuscript is dedicated to the memory of William B. Hooper.  \nAcknowledgements  \nThe research described in this thesis was conducted in the city now known as Toronto, Canada. The Mohawk name for this place,“Tkaronto”, indicates a meeting point between the trees and the water. The land, water, and air that make up the surrounding region are governed by the Dish with One Spoon Wampum; throughout the generations, they have been cared for by many Indigenous peoples, who in turn have been subjected to material and cultural dispossession through a (still ongoing) project of colonization. This is a human tragedy, and also an intellectual tragedy. For example, Leanne Betasamosake Simpson argues that, by failing to engage with Indigenous ontologies that emphasize relationality, Western Scholars often unwittingly use concepts like “abstraction” (the bread and butter of Computer Science) towards extractive ends [Simpson, 2017] . I have seen this pattern emerge frequently in research on technical approaches to mitigate harms in socio-technical systems, which typically exhibit nuanced and context-specific complexities. In the Academy and beyond, much remains to be done when it comes to engaging with Indigenous (and other non-Western) ways of producing and preserving knowledge. These efforts can be understood as a small part of the broader movement towards reconciliation.  \nOn a personal level, I am grateful to this land, and the animal and plant life it supports, for providing a continual source of inspiration throughout my studies here.  \nWriting about gratitude is difficult for me because the readily available language tends to evoke an outstanding debt or obligation. But these sentiments fail to capture my experience","cbCailx4cy3p5zN6","https://ap.wps.com/l/cbCailx4cy3p5zN6","pdf",11103370,1,137,"English","en",105,"# Abstract\n# Acknowledgements\n## Dedicated to William B. Hooper","[{\"question\":\"为什么训练数据的不完美会导致机器学习部署失败？\",\"answer\":\"训练数据常包含数据敏感性，使得模型在部署时出现意外故障模式。文中指出，这类问题与训练数据相对任务目标存在不匹配有关。\"},{\"question\":\"论文如何应对训练数据中可能存在的歧视性信号？\",\"answer\":\"通过公平表征学习的方法，旨在防止 ML 将训练数据中测得的先验人类歧视直接编码到模型中。\"},{\"question\":\"面对强化学习中状态与动作覆盖不足，论文提出了什么思路？\",\"answer\":\"论文利用因果先验来建模数据采集环境的单步动力学，并将观察到的轨迹拼接，生成新的但仍合理的反事实轨迹，从而进行数据增强。\"}]","Robust Machine Learning - by Transforming and Augmenting Imperfect Training Data - A Thesis | PDF",1785807698,345,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"robust-machine-learning-by-transforming-and-augmenting-imperfect-training-data-a-thesis","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/robust-machine-learning-by-transforming-and-augmenting-imperfect-training-data-a-thesis/121910/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"为什么训练数据的不完美会导致机器学习部署失败？","Question",{"text":75,"@type":76},"训练数据常包含数据敏感性，使得模型在部署时出现意外故障模式。文中指出，这类问题与训练数据相对任务目标存在不匹配有关。","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"论文如何应对训练数据中可能存在的歧视性信号？",{"text":80,"@type":76},"通过公平表征学习的方法，旨在防止 ML 将训练数据中测得的先验人类歧视直接编码到模型中。",{"name":82,"@type":73,"acceptedAnswer":83},"面对强化学习中状态与动作覆盖不足，论文提出了什么思路？",{"text":84,"@type":76},"论文利用因果先验来建模数据采集环境的单步动力学，并将观察到的轨迹拼接，生成新的但仍合理的反事实轨迹，从而进行数据增强。","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]