[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-116877-en":3,"doc-seo-116877-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},116877,687197100911,"Himbo","https://ap-avatar.wpscdn.com/avatar/a000239b6f1da00475?x-image-process=image/resize,m_fixed,w_180,h_180&k=1785132997149421697",8,"Research & Report","Recent Advances in Optimal Transport for Machine Learning","Optimal Transport (OT) has emerged in machine learning as a probabilistic framework for comparing and transforming probability distributions. Rooted in the classical theory of Monge and Kantorovich, OT provides both a principled metric perspective—often using Wasserstein distance—and a practical toolbox for manipulating distributions. This survey reviews key contributions from 2012 to 2022 across supervised, unsupervised, transfer, and reinforcement learning, while also emphasizing recent progress in computational OT and its interaction with ML practice.","arXiv :2306 . 16156v1 [ cs .LG] 28 Jun 2023  \nRecent Advances in Optimal Transport for  \nMachine Learning  \nEduardo Fernandes Montesuma, Fred Ngol Mboula, and Antoine Souloumiac,  \nAbstract—Recently, Optimal Transport has been proposed as a probabilistic framework in Machine Learning for comparing and manipulating probability distributions. This is rooted in its rich history and theory, and has offered new solutions to different problems in machine learning, such as generative modeling and transfer learning. In this survey we explore contributions of Optimal Transport for Machine Learning over the period 2012 – 2022, focusing on four sub-ﬁelds of Machine Learning: supervised, unsupervised, transfer and reinforcement learning. We further highlight the recent development in computational Optimal Transport, and its interplay with Machine Learning practice.  \nIndex Terms—Optimal Transport, Wasserstein Distance, Machine Learning  \n~~ ~~ F ~~ ~~  \n1 INTRODUCTION  \nOPTIMAL transportation theory is a well-established  \nﬁeld of mathematics founded by the works of Gaspard Monge [1] and economist Leonid Kantorovich [2] . Since its genesis, this theory has made signiﬁcant contributions to mathematics [3], physics [4], and computer science [5] . Here, we study how Optimal Transport (OT) contributes to different problems within Machine Learning (ML) . Optimal Transport for Machine Learning (OTML) is a growing research subject in the ML community. Indeed, OT is useful for ML through at least two viewpoints: (i) as a loss function, and (ii) for manipulating probability distributions.  \nFirst, OT deﬁnes a metric between distributions, known by different names, such as Wasserstein distance, Dudley metric, Kantorovich metric, or Earth Mover Distance (EMD) . This metric belongs to the family of Integral Probability Metrics (IPMs) (see section 2.1.1) . In many problems (e.g., generative modeling), the Wasserstein distance is preferable over other notions of dissimilarity between distributions, such as the Kullback-Leibler (KL) divergence, due its topological and statistical properties.  \nSecond, OT presents itself as a toolkit or framework for ML practitioners to manipulate probability distributions. For instance, researchers can use OT for aggregating or interpolating between probability distributions, through Wasserstein barycenters and geodesics. Hence, OT is a principled way to study the space of probability distributions.  \nThis survey provides an updated view of how OTML in the recent years. Even though previous surveys exist [6]–[11], the rapid growth of the ﬁeld justiﬁes a closer look at OTML. The rest of this paper is organized as follows. Section 2 discusses the fundamentals of different areas of ML, as well as an overview of OT. Section 3 reviews recent developments in computational optimal transport. The further sections explore OT for 4 ML problems: supervised (section 4), unsupervised (section 5), transfer (section 6), and  \n􀀏 The authors are with the Universit´e Paris-Saclay, CEA, LIST, F-91120, Palaiseau, France.E-mail: [eduardo.fernandesmontesuma@cea.fr](eduardo.fernandesmontesuma@cea.fr)  \nreinforcement learning (section 7) . Section 8 concludes this paper with general remarks and future research directions.  \n1.1 Related Work  \nBefore presenting the recent contributions of OT for ML, we brieﬂy review previous surveys on related themes. Concerning OT as a ﬁeld of mathematics, a broad literature is available. For instance,[12] and [13], present OT theory from a continuous point of view. The most complete text on the ﬁeld remains [3], but it is far from an introductory text.  \nAs we cover in this survey, the contributions of OT for ML are computational in nature. In this sense, practitioners will be focused on discrete formulations and computational aspects of how to implement OT. This is the case of surveys [9] and [10] . While [9] serves as an introduction to OT and its discrete formulation, [10] gives a more broad view on the co","cbCaiarsnI7EzA4q","https://ap.wps.com/l/cbCaiarsnI7EzA4q","pdf",4086550,1,20,"English","en",105,"# Introduction\n## Related Work\n# Background\n## Optimal Transport\n# Fundamentals and Overview\n## Computational Optimal Transport\n# Optimal Transport for Machine Learning Tasks\n## Supervised Learning\n## Unsupervised Learning\n## Transfer Learning\n## Reinforcement Learning\n# Conclusion and Future Directions","[{\"question\":\"How is Optimal Transport used in machine learning?\",\"answer\":\"Optimal Transport is used either as a loss function or as a framework to manipulate probability distributions. It provides comparisons through metrics such as Wasserstein distance and supports operations like aggregation and interpolation.\"},{\"question\":\"Why is Wasserstein distance often preferred to KL divergence in generative modeling?\",\"answer\":\"The survey highlights that Wasserstein distance has favorable topological and statistical properties. This makes it preferable to other dissimilarity notions such as Kullback-Leibler divergence in many settings.\"},{\"question\":\"Which learning paradigms does the survey cover between 2012 and 2022?\",\"answer\":\"The survey focuses on four sub-fields: supervised, unsupervised, transfer, and reinforcement learning. It also reviews recent developments in computational optimal transport and its relevance to ML practice.\"}]","Recent Advances in Optimal Transport for Machine Learning | PDF",1785672184,50,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"recent-advances-in-optimal-transport-for-machine-learning","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/recent-advances-in-optimal-transport-for-machine-learning/116877/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"How is Optimal Transport used in machine learning?","Question",{"text":75,"@type":76},"Optimal Transport is used either as a loss function or as a framework to manipulate probability distributions. It provides comparisons through metrics such as Wasserstein distance and supports operations like aggregation and interpolation.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why is Wasserstein distance often preferred to KL divergence in generative modeling?",{"text":80,"@type":76},"The survey highlights that Wasserstein distance has favorable topological and statistical properties. This makes it preferable to other dissimilarity notions such as Kullback-Leibler divergence in many settings.",{"name":82,"@type":73,"acceptedAnswer":83},"Which learning paradigms does the survey cover between 2012 and 2022?",{"text":84,"@type":76},"The survey focuses on four sub-fields: supervised, unsupervised, transfer, and reinforcement learning. It also reviews recent developments in computational optimal transport and its relevance to ML practice.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,126,129,133],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":29,"slug":113},6,"Technology","technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":21,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":127,"show_sort_weight":21,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":46,"category_name":135,"show_sort_weight":106,"slug":136},19,"General","general"]