[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-123353-en":3,"doc-seo-123353-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},123353,962075114765,"Quinn","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Machine Learning Systems in Constrained Environments - Dissertation","Machine learning training and inference systems face major constraints in modern computation environments as ML models grow in size and deployment becomes more diverse. This dissertation presents system designs and algorithmic techniques that enable efficient ML execution under limited memory GPUs, on-premises clusters, and serverless platforms. It introduces multi-device training with fast, high-quality placements, inference resource sharing via rapid near-optimal autoscaling decisions, and cost-efficient distributed GNN training through analytic offline optimization and gray-box online tuning.","© 2024 Beomyeol Jeon  \nMACHINE LEARNING SYSTEMS IN CONSTRAINED ENVIRONMENTS  \nBY  \nBEOMYEOL JEON  \nDISSERTATION  \nSubmitted in partial fulfillment of the requirements for the degree of Doctor of Philosophy in Computer Science in the Graduate College of the  \nUniversity of Illinois Urbana-Champaign, 2024  \nUrbana, Illinois  \nDoctoral Committee:  \nProfessor Indranil Gupta, Chair  \nProfessor Matthew Caesar  \nAssistant Professor Yongjoo Park  \nDoctor Chen Wang, IBM Research  \nABSTRACT  \nMachine learning (ML) training and inference systems encounter constraints in current computation environments due to increased ML model sizes, the fast-growing popularity of ML/AI, etc. In this thesis, we show how machine learning training and inference systems can be executed successfully and efficiently in constrained computation environments, such as limited-memory GPUs, on-premises clusters, and serverless environments, by using a novel combination of algorithms, optimizations, and well-reasoned system designs. Concretely, we propose (i) a system that enables large ML model training over multiple memory-constrained GPU devices via algorithms and system designs that achieve fast placements with a quality comparable to expert-designed placements, (ii) a system that enables efficient resource sharing among ML inference jobs in fixed-size on-premises clusters by making close-to-optimal autoscaling decisions quickly via several relaxation methods in optimization and prediction, and (iii) a system that enables cost-efficient distributed GNN training on constrained serverless execution environments by auto-tuning configuration via analytic model-based offline optimization and gray-box heuristic-based online optimization.  \nTo my parents, my wife, and my lovely daughter  \niii  \nACKNOWLEDGMENTS  \nI would like to express my sincere gratitude to my advisor, Prof. Indranil Gupta, for his support and guidance. I have enjoyed countless intellectual discussions with him and learned a lot about developing ideas, conveying them crisply and professionally, and many other aspects. I also appreciated his valuable advice on research directions, career paths after the PhD program, and even life.  \nMy profound gratitude goes to my PhD thesis committee members, Prof. Matthew Caesar, Prof. Yongjoo Park, and Dr. Chen Wang, for asking me perceptive questions and providing me with insightful feedback. They allowed me to think more deeply about my research from different angles and guided me on how to improve my work and thesis.  \nI want to acknowledge all the research collaborations involved in this thesis. I was privileged to work with Chen Wang, Diana Arroyo, and Alaa Youssef at IBM Research to extend my research interest to containerized clusters. Prof. Yongjoo Park shared his brilliant ideas and guided me toward promising directions. The contributions of Linda Cai, Pallavi Srivastava, Jintao Jiang, Xiaolan Ke, Yitao Meng, and Cong Xie were essential to complete the work in this thesis.  \nI would like to thank all DRPG folks including Shegufta Ahsan, Rui Yang, Faria Kalim, Cong Xie, Mainak Ghosh, Shadi Noghabi, Le Xu, Anna Karanika, Ali Zaidi, Xiaojuan Ma, Maleeha Masood, Chirag Shetty, and many others. They have been good friends and helped me finish this long PhD journey by sharing their thoughts both academically and personally. I also would like to acknowledge the Department of Computer Science at the University of Illinois Urbana-Champaign for providing me with the opportunity to pursue a doctorate degree, with wonderful academic support from faculty members and staff. I wish to acknowledge the National Science Foundation, IBM Research, and Schlumberger for their financial support for the work in this thesis.  \nI am also grateful for internship opportunities at Google with Hector Gonzalez and Mohsen Vakilian, and at Nokia Bell Labs with Muntasir Rahman and Anwar Walid. I have learned a lot about engineering and research work in the industry, and I have grown my s","cbCaikMcLE6M8bfh","https://ap.wps.com/l/cbCaikMcLE6M8bfh","pdf",2590740,2,1,136,"English","en",105,"# Chapter 1 Introduction\n## Thesis Contributions\n## Broader Impact\n## Thesis Organization\n# Chapter 2 Fast Device Placement of Machine Learning Graphs via Baechi\n## Introduction\n## Related Work\n## New Algorithms for Memory-Constrained Placement\n## Baechi Design\n## Implementation\n## Evaluation\n## Conclusions\n# Chapter 3 SLO-Awareness for On-Premises Containerized ML Inference Clusters via Faro\n## Introduction\n## Motivation: Challenges, Contributions\n## Related Work\n## Faro Building Blocks\n## Faro Autoscaler Design\n## Implementation\n## Evaluation","[{\"question\":\"How does the dissertation address ML constraints in limited-memory environments?\",\"answer\":\"It proposes system and algorithm combinations that execute training and inference successfully in constrained compute settings, including limited-memory GPUs, on-premises clusters, and serverless execution.\"},{\"question\":\"What is the contribution for training large ML models across multiple memory-constrained GPUs?\",\"answer\":\"It presents a system that enables large-model training over multiple constrained GPU devices by performing fast placements with quality comparable to expert-designed placements.\"},{\"question\":\"How does the dissertation improve resource sharing and scaling for on-premises inference clusters?\",\"answer\":\"It introduces an approach that makes near-optimal autoscaling decisions quickly by using relaxation methods for optimization and prediction to share resources efficiently across inference jobs.\"}]","Machine Learning Systems in Constrained Environments - Dissertation | PDF",1785816086,343,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"machine-learning-systems-in-constrained-environments-dissertation","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,48,51],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":20},"https://docshare.wps.com/document/","Document",{"item":49,"name":12,"@type":44,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":44,"position":53},"https://docshare.wps.com/document/machine-learning-systems-in-constrained-environments-dissertation/123353/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-06","2026-08-04",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"How does the dissertation address ML constraints in limited-memory environments?","Question",{"text":76,"@type":77},"It proposes system and algorithm combinations that execute training and inference successfully in constrained compute settings, including limited-memory GPUs, on-premises clusters, and serverless execution.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What is the contribution for training large ML models across multiple memory-constrained GPUs?",{"text":81,"@type":77},"It presents a system that enables large-model training over multiple constrained GPU devices by performing fast placements with quality comparable to expert-designed placements.",{"name":83,"@type":74,"acceptedAnswer":84},"How does the dissertation improve resource sharing and scaling for on-premises inference clusters?",{"text":85,"@type":77},"It introduces an approach that makes near-optimal autoscaling decisions quickly by using relaxation methods for optimization and prediction to share resources efficiently across inference jobs.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]