[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-122258-en":3,"doc-seo-122258-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},122258,1099514067415,"Rowan","https://ap-avatar.wpscdn.com/avatar/100002539d78ffe74a7?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779092875211072502",8,"Research & Report","Mitigating Hidden Technical Debt in Machine Learning Systems - master project, spring 2023","The paper “Sculley et al. [2015]” is widely recognized in the Machine Learning Operations community, and its identified technical debts have improved ML deployments through better artifact version control, reproducibility, and infrastructure practices. Yet, most tools focus on model and related software, leaving data-related debts insufficiently addressed, as interviews indicate. This project proposes a data-centric method using an automatic metadata compiler and an implementation compiler to capture dependencies, standardize feature engineering, optimize systems, and reduce development cost.","Mats Eikeland Mollestad  \nMitigating Hidden Technical Debt in Machine Learning Systems  \nmaster project, spring 2023  \nArtificial Intelligence Group  \nDepartment of Computer and Information Science  \nFaculty of Information Technology, Mathematics and Electrical Engineering  \ni  \nAbstract  \nThe paper Sculley et al. [2015] is recognized as one of the most influential works within the Machine Learning Operations community. The technical debts identified therein have led to significant improvements in the deployment of ML systems, particularly concerning artifact version control, reproducibility, and infrastructure. However, existing tools have predominantly focused on the model and the associated software, leaving data related debts largely unaddressed, as evidenced by conducted interviews.  \nThis project introduces a novel method specifically targeting data-related debt. This approach results in improved ML products and reduces the time spent understanding the system’s data flow. This is achieved by offering a metadata compiler that automatically captures data dependencies, describes data schemas, standardise feature engineering, and simplifies the understanding of model logic. The collection of dependency metadata would otherwise be too inconvenient to manually define. The obtained metadata can then be empowered by an implementation compiler that optimize the system and decrease development costs by generating common components essential for Machine Learning Operations (MLOps) .  \nii  \nPreface  \nI would like to thank my supervisor Anders Kofod-Petersen for making this master thesis possible, and providing feedback. Furthermore, I would like to thank everyone that took the time to be interview about their problems when deploying ML, but also for testing the proposed solution.  \nLastly, I would like to thank Otovo ASA for allowing me to write a master thesis, while working part time on practical problems related to the master thesis. Therefore, making it clearer why such research would be useful.  \nMats Eikeland Mollestad Trondheim, June 13, 2023  \nContents  \n1 Introduction 1  \n1.1 Background and Motivation ...................... 2  \n1.2 Goals and Research Questions ..................... 2  \n1.3 Research Method ............................ 2  \n1.4 Contributions .............................. 3  \n1.5 Thesis Structure ............................ 3  \n2 Background Theory and Motivation 5  \n2.1 Background Theory .......................... 5  \n2.1.1 Software 2.0 ........................... 5  \n2.1.2 Types of ML .......................... 6  \n2.1.3 ML research vs. deployment .................. 7  \n2.1.4 MLOps ............................. 9  \n2.1.5 Model development life-cycle ................. 10  \n2.1.6 Inference behavior ....................... 16  \n2.1.7 Data Drift ............................ 17  \n2.1.8 Data management ....................... 19  \n2.1.9 Automation ........................... 24  \n2.1.10 Versioning ............................ 26  \n2.1.11 Declarative design ....................... 26  \n2.1.12 Model cards ........................... 27  \n2.1.13 Vector database ........................ 27  \n2.1.14 Data centric AI ......................... 28  \n2.1.15 LLMs .............................. 31  \n2.1.16 Insights from Interviews .................... 32  \n2.2 Hidden technical Debt in ML Systems ................ 33  \n2.2.1 Entanglement .......................... 33  \n2.2.2 Correction Cascades ...................... 34  \n2.2.3 Unstable Data Dependencies ................. 34  \niv CONTENTS  \n2.2.4 Underutilized Data Dependencies ............... 34  \n2.2.5 Static Analysis of Data Dependencies ............ 35  \n2.2.6 Glue Code ............................ 35  \n2.2.7 Pipeline Jungles ........................ 35  \n2.2.8 Dead Experimental Codepaths ................ 36  \n2.2.9 Plain-Old-Data Type Smell .................. 36  \n2.2.10 Multiple-Language Smell ................... 36  \n2.2.11 Prediction Bias ......................... 37  \n","cbCairx9ZoHe8QJ6","https://ap.wps.com/l/cbCairx9ZoHe8QJ6","pdf",7179633,1,94,"English","en",105,"# Introduction\n## Background and Motivation\n## Goals and Research Questions\n## Research Method\n## Contributions\n## Thesis Structure\n# Background Theory and Motivation\n## Software 2.0\n## Types of ML\n## ML research vs. deployment\n## MLOps\n## Model development life-cycle\n## Inference behavior\n## Data Drift\n## Data management\n## Automation\n## Versioning\n## Declarative design\n## Model cards\n## Vector database\n## Data centric AI\n## LLMs\n## Insights from Interviews\n## Hidden technical Debt in ML Systems\n## Motivation\n# Proposed System\n## Core Design Choice\n## Metadata Compiler\n## Implementation Compiler\n# Experiments and Results\n## Experimental Plan\n## Results\n# Evaluation and Conclusion\n## Evaluation","[{\"question\":\"Why focus on data-related technical debt in ML systems?\",\"answer\":\"Existing work mainly improves model and software aspects like version control and reproducibility, while data-related issues remain largely unaddressed. Interviews during deployment highlighted this gap.\"},{\"question\":\"How does the proposed approach capture data dependencies automatically?\",\"answer\":\"It introduces a metadata compiler that automatically collects dependency metadata, describes data schemas, standardizes feature engineering, and makes the model logic easier to understand.\"},{\"question\":\"What benefits are expected from generating dependency metadata and compiling an implementation?\",\"answer\":\"The dependency metadata enables an implementation compiler to optimize the system, generate common MLOps components, improve ML product quality, and reduce development costs and understanding time.\"}]","Mitigating Hidden Technical Debt in Machine Learning Systems - master project, spring 2023 | PDF",1785809700,237,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"mitigating-hidden-technical-debt-in-machine-learning-systems-master-project-spring-2023","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/mitigating-hidden-technical-debt-in-machine-learning-systems-master-project-spring-2023/122258/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why focus on data-related technical debt in ML systems?","Question",{"text":75,"@type":76},"Existing work mainly improves model and software aspects like version control and reproducibility, while data-related issues remain largely unaddressed. Interviews during deployment highlighted this gap.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed approach capture data dependencies automatically?",{"text":80,"@type":76},"It introduces a metadata compiler that automatically collects dependency metadata, describes data schemas, standardizes feature engineering, and makes the model logic easier to understand.",{"name":82,"@type":73,"acceptedAnswer":83},"What benefits are expected from generating dependency metadata and compiling an implementation?",{"text":84,"@type":76},"The dependency metadata enables an implementation compiler to optimize the system, generate common MLOps components, improve ML product quality, and reduce development costs and understanding time.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]