[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-seo-128843-105":3,"detail-sidebar-cat-0-en-105":81,"doc-detail-128843-en":130},{"code":4,"msg":5,"data":6},0,"ok",{"site_id":7,"language":8,"slug":9,"title":10,"keywords":11,"description":12,"schema_data":13,"social_meta":74,"head_meta":76,"extra_data":78,"updated_unix":80},105,"en","contributions-to-the-trustworthy-machine-learning-pipeline-data-selection-training-and-post-training-verification-through-convexity-and-the-wasserstein-distance","Contributions to the Trustworthy Machine Learning Pipeline - Data Selection, Training, and Post-training Verification through Convexity and the Wasserstein Distance","","The thesis addresses challenges and risks in deploying complex machine learning models in safety-critical applications. It reviews methods designed to improve model reliability and introduces the concept of a trustworthy machine learning pipeline. The work focuses on the needs and constraints of virtual power plants, where predictive performance must be strengthened and operational errors can have severe consequences. It proposes contributions across the model lifecycle without sacrificing performance. First, it presents a robust distributional training method using Wasserstein distance for shallow convex neural networks trained with corrupted data, providing theoretical out-of-sample guarantees, physical convex constraint adaptation, and a post-training verification problem for stability. Second, it develops an unsupervised anomaly detection and data filtering approach based on truncated sliced-Wasserstein distance, with computational approximations and experiments including an open dataset for local demand management via localized critical peak rebates, supporting a benchmark for training data selection.",{"@graph":14,"@context":73},[15,34,56],{"@type":16,"itemListElement":17},"BreadcrumbList",[18,23,27,31],{"item":19,"name":20,"@type":21,"position":22},"https://docshare.wps.com","Home","ListItem",1,{"item":24,"name":25,"@type":21,"position":26},"https://docshare.wps.com/document/","Document",2,{"item":28,"name":29,"@type":21,"position":30},"https://docshare.wps.com/document/research-report/","Research & Report",3,{"item":32,"name":10,"@type":21,"position":33},"https://docshare.wps.com/document/contributions-to-the-trustworthy-machine-learning-pipeline-data-selection-training-and-post-training-verification-through-convexity-and-the-wasserstein-distance/128843/",4,{"url":32,"name":10,"@type":35,"image":36,"author":41,"headline":10,"publisher":44,"fileFormat":47,"inLanguage":8,"description":12,"dateModified":48,"datePublished":49,"encodingFormat":47,"isAccessibleForFree":50,"interactionStatistic":51},"DigitalDocument",{"url":37,"@type":38,"width":39,"height":40},"https://docshare.wps.com/thumbnails/contributions-to-the-trustworthy-machine-learning-pipeline-data-selection-training-and-post-training-verification-through-convexity-and-the-wasserstein-distance/128843.png","ImageObject",300,407,{"name":42,"@type":43},"Aria","Person",{"url":19,"name":45,"@type":46},"DocShare","Organization","application/pdf","2026-09-20","2026-08-06",true,{"@type":52,"interactionType":53,"userInteractionCount":55},"InteractionCounter",{"@type":54},"ViewAction",8,{"@type":57,"mainEntity":58},"FAQPage",[59,65,69],{"name":60,"@type":61,"acceptedAnswer":62},"What is the main goal of the thesis on trustworthy machine learning pipelines?","Question",{"text":63,"@type":64},"To improve the reliability of complex machine learning models when deployed in safety-critical settings by strengthening methods for data selection, training, and post-training verification across the model lifecycle.","Answer",{"name":66,"@type":61,"acceptedAnswer":67},"How does the proposed training method improve robustness?",{"text":68,"@type":64},"It introduces a Wasserstein-distance-based robust distributional training approach for shallow convex neural networks under corrupted data, framed as a convex optimization problem with theoretical out-of-sample performance guarantees.",{"name":70,"@type":61,"acceptedAnswer":71},"What role does Wasserstein distance play in anomaly detection and data selection?",{"text":72,"@type":64},"The thesis proposes an unsupervised data filtering and anomaly detection method using truncated sliced-Wasserstein distance, including computational approximations, and uses the filtered data to support training-data selection and benchmarking.","https://schema.org",{"og:url":32,"og:type":75,"og:title":10,"og:site_name":45,"og:description":12},"article",{"robots":77,"canonical":32},"index,follow",{"doc_id":79,"site_id":7},128843,1786003835,{"code":4,"msg":82,"data":83},"success",[84,88,92,96,101,106,111,114,119,122,126],{"id":22,"doc_module":4,"doc_module_name":25,"category_name":85,"show_sort_weight":86,"slug":87},"Story & Novel",90,"story-novel",{"id":26,"doc_module":4,"doc_module_name":25,"category_name":89,"show_sort_weight":90,"slug":91},"Literature",80,"literature",{"id":33,"doc_module":4,"doc_module_name":25,"category_name":93,"show_sort_weight":94,"slug":95},"Exam",70,"exam",{"id":97,"doc_module":4,"doc_module_name":25,"category_name":98,"show_sort_weight":99,"slug":100},5,"Comic",60,"comic",{"id":102,"doc_module":4,"doc_module_name":25,"category_name":103,"show_sort_weight":104,"slug":105},6,"Technology",50,"technology",{"id":107,"doc_module":4,"doc_module_name":25,"category_name":108,"show_sort_weight":109,"slug":110},7,"Healthcare",40,"healthcare",{"id":55,"doc_module":4,"doc_module_name":25,"category_name":29,"show_sort_weight":112,"slug":113},30,"research-report",{"id":115,"doc_module":4,"doc_module_name":25,"category_name":116,"show_sort_weight":117,"slug":118},9,"Religion & Spirituality",20,"religion-spirituality",{"id":117,"doc_module":4,"doc_module_name":25,"category_name":120,"show_sort_weight":117,"slug":121},"World Cup","world-cup",{"id":123,"doc_module":4,"doc_module_name":25,"category_name":124,"show_sort_weight":123,"slug":125},10,"Lifestyle","lifestyle",{"id":127,"doc_module":4,"doc_module_name":25,"category_name":128,"show_sort_weight":97,"slug":129},19,"General","general",{"code":4,"msg":82,"data":131},{"doc_id":79,"user_id":132,"nickname":42,"user_avatar":133,"doc_module":4,"category_id":55,"category_name":29,"doc_title":10,"doc_description":12,"doc_content":134,"file_id":135,"file_url":136,"file_type":137,"file_size":138,"view_count":55,"is_deleted":4,"is_public":22,"is_downloadable":22,"audit_status":22,"page_count":139,"language":140,"language_code":8,"site_id":7,"html_lang":8,"table_of_contents":141,"faqs":142,"seo_title":143,"seo_description":12,"update_tm":80,"read_time":144},2336474459895,"https://ap-avatar.wpscdn.com/avatar/22000baeef7a5ed0655?x-image-process=image/resize,m_fixed,w_180,h_180&k=1786071322749376916","Document en libre accès dans PolyPublie  \nOpen Access document in PolyPublie  \n| URL de PolyPublie:\u003Cbr>PolyPublie URL: | [https://publications.polymtl.ca/64736/](https://publications.polymtl.ca/64736/) |\n| --- | --- |\n| Directeurs de recherche:\u003Cbr>Advisors: | Antoine Lesage-Landry |\n| Programme:\u003Cbr>  Program: | Génie électrique |\n\nCe fichier a été téléchargé à partir de PolyPublie, le dépôt institutionnel de Polytechnique Montréal  \nThis file has been downloaded from PolyPublie, the institutional repository of Polytechnique Montréal  \n[https://publications.polymtl.ca](https://publications.polymtl.ca)  \nPOLYTECHNIQUE MONTRÉAL  \naffiliée à l’Université de Montréal  \nContributions to the Trustworthy Machine Learning Pipeline: Data Selection, Training, and Post-training Verification through Convexity and the Wasserstein Distance  \nJULIEN PALLAGE  \nDépartement de génie électrique  \nMémoire présenté en vue de l’obtention du diplôme de Maîtrise ès sciences appliquées  \nGénie électrique  \nAvril 2025  \n© Julien Pallage, 2025 .  \nPOLYTECHNIQUE MONTRÉAL  \naffiliée à l’Université de Montréal  \nCe mémoire intitulé :  \nContributions to the Trustworthy Machine Learning Pipeline: Data Selection, Training, and Post-training Verification through Convexity and the Wasserstein Distance  \nprésenté par Julien PALLAGE  \nen vue de l’obtention du diplôme de Maîtrise ès sciences appliquéesa été dûment accepté par le jury d’examen constitué de :  \nShuang GAO, président  \nAntoine LESAGE-LANDRY, membre et directeur de recherche Bowen YI, membre  \niii  \nDEDICATION  \nÀ Georgette qui a su me transmettre sa curiosité .  \niv  \nACKNOWLEDGEMENTS  \nThis work would not have been possible without the support of so many! I want to extend my gratitude for the thoughtful and sincere guidance of Prof. Antoine Lesage-Landry; the day-to-day help in the realm of data science from Dr. Salma Naccache and Dr. Bertrand Scherrer; the endorsement of Odile Noël, Ahmed Abdellatif, and Steve Boursiquot from Hilo; the financial support of the Natural Sciences and Engineering Research Council of Canada (NSERC), the Fonds de Recherche du Québec (FRQ), MITACS, and Hilo; as well as the fun discussions and good time spent with friends and colleagues, Olivier, Jean-Luc, Matthias, Abraham, Loreley, William, Samuel, Xavier, Étienne, Christina, Julien, and Tudor. Finally, I want to thank my partner, Nohaila, for always being there forme, it would have been a lot harder without you.  \nOn another note, the works of Prof. Spyros Chatzivasileiadis and Prof. Peyman Mohajerin Esfahani were instrumental in motivating this thesis. I heartfully invite interested readers to consult their respective literature.  \nv  \nRÉSUMÉ  \nDans ce travail, nous étudions les défis et risques liés au déploiement de modèles complexes d’apprentissage machine (machine learning, ML) dans des applications critiques. Nous débutonspar un survol de différentes méthodes de la litérature utilisées pour renforcer la fiabilité de ces modèles afin d’introduire le concept de pipeline d’apprentissage machine fiable (trustworhty) . Nous étudions spécifiquement les besoins et les contraintes associées aux centrales virtuelles étant donné leurbesoin de meilleurs algorithmes de prédiction et la sévérité potentielle d’erreurs opérationnelles. Nous proposons plusieurs contributions visant à améliorer la sûreté des modèles d’apprentissage, et ce tout au long de leur cycle de vie, sans compromettre leur performance.  \nComme première contribution, nous présentons une nouvelle méthode d’entraînement robuste en distribution, sous la distance de Wasserstein, pour les réseaux de neurones convexes peu profond ( shallow convex neural networks, SCNNs) soumis à des jeux de données corrompues. Notre approche repose sur un nouveau problème d’optimisation d’entraînement convexe permettant de faire le pont entre les réseaux de neurones ReLU convexes optimaux et les réseaux de neurones ReLU non-convexes. Nous adaptons ce programme d’entraînement sous sa","cbCaimnDkOFHyh0C","https://ap.wps.com/l/cbCaimnDkOFHyh0C","pdf",17540382,113,"English","# Introduction\n## Trustworthy ML pipeline and safety-critical deployment\n## Virtual power plants and operational error risks\n# Robust distributional training\n## Convex neural networks with corrupted datasets\n## Wasserstein-robust optimization and theoretical guarantees\n## Post-training verification for stability\n## Physical convex constraints adaptation\n# Data filtering and anomaly detection\n## Truncated sliced-Wasserstein unsupervised method\n## Computational approximations\n## Open dataset benchmark for demand management\n# Experiments and evaluations","[{\"question\":\"What is the main goal of the thesis on trustworthy machine learning pipelines?\",\"answer\":\"To improve the reliability of complex machine learning models when deployed in safety-critical settings by strengthening methods for data selection, training, and post-training verification across the model lifecycle.\"},{\"question\":\"How does the proposed training method improve robustness?\",\"answer\":\"It introduces a Wasserstein-distance-based robust distributional training approach for shallow convex neural networks under corrupted data, framed as a convex optimization problem with theoretical out-of-sample performance guarantees.\"},{\"question\":\"What role does Wasserstein distance play in anomaly detection and data selection?\",\"answer\":\"The thesis proposes an unsupervised data filtering and anomaly detection method using truncated sliced-Wasserstein distance, including computational approximations, and uses the filtered data to support training-data selection and benchmarking.\"}]","Contributions to the Trustworthy Machine Learning Pipeline - Data Selection, Training, and Post-training Verification through Convexity and the Wasserstein Distance | PDF",285]