[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-127458-en":3,"doc-seo-127458-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},127458,962084925290,"Ophelia","https://ap-avatar.wpscdn.com/davatar_085a072bc5b1113ac321206ff7593b45",8,"Research & Report","Towards Efficient and Reliable Infrastructure for Machine Learning - Dissertation","The rapid advancement of machine learning, especially large language and generative models, delivers transformative capabilities while imposing extreme computational costs. Training and serving require large numbers of specialized accelerators, driving high power consumption and multi-hundred-million-dollar expenditures. This thesis targets infrastructure efficiency through three linked dimensions: core optimization, workload adaptation, and reliability enhancement, presenting five contributions for improved resource utilization and resilience across distributed systems.","© 2025 Archit Patke  \nTOWARDS EFFICIENT AND RELIABLE INFRASTRUCTURE FOR MACHINE  \nLEARNING  \nBY  \nARCHIT PATKE  \nDISSERTATION  \nSubmitted in partial fulfillment of the requirements for the degree of Doctor of Philosophy in Electrical and Computer Engineering  \nin the Graduate College of the  \nUniversity of Illinois Urbana-Champaign, 2025  \nUrbana, Illinois  \nDoctoral Committee:  \nProfessor Ravishankar K. Iyer, Chair  \nProfessor Nam Sung Kim  \nProfessor Jian Huang  \nDr. Mudhakar Srivatsa, IBM Research  \nAbstract  \nThe rapid advancement of machine learning, particularly large language and generative models, has enabled transformative applications but at immense computational cost. Training and serving these models requires tens of thousands of specialized accelerators that consume megawatts of power and costs hundreds of millions of dollars. The efficiency of the underlying compute infrastructure is, therefore, critical for making machine learning more accessible and sustainable.  \nThis thesis addresses infrastructure efficiency challenges through three interconnected dimensions: core infrastructure optimization, workload adaptation, and reliability enhancement. We present five contributions that introduce novel methodologies and system designs for improving efficiency, resource utilization, and reliability.  \nFirst, in core infrastructure optimization, we address resource fragmentation by enabling resource disaggregation, and resolving network contention that arises in disaggregated systems. INDIGO addresses memory disaggregation challenges through network-aware page migration. The system uses contextual multi-armed bandits trained on historical application data to make migration decisions that account for network transfer costs and memory access locality benefits. Netscope introduces a delay sensitivity-driven congestion mitigation framework that quantifies how applications are affected by network congestion using probabilistic regression models. The framework dynamically adjusts congestion control parameters based on estimated delay sensitivity, and selectively throttles applications with low sensitivity while protecting delay-sensitive ones.  \nSecond, in workload adaptation, we address the unique characteristics of modern ML workloads. QLM improves efficiency of distributed inference for large language models by multiplexing interactive and batch requests. QLM leverages statistical properties of continuous batching to estimate request waiting times in queues and groups requests with similar performance characteristics to enable efficient decision-making and orchestrates request pulling, eviction, load balancing, and model swapping operations. Complementary to QLM, Chiron introduces hierarchical autoscaling that employs multi-level backpressure mechanisms that distinguishes between interactive and batch requests. The framework dynamically adjusts batch sizes atthe local level using reactive backpressure and makes global scaling decisions based on request waiting time estimation to enable meeting the service-level objectives while maximizing efficiency.  \nThird, in reliability enhancement, we conduct a comprehensive characterization of GPU failures in modern AI accelerators through analysis of failure data from large-scale production clusters. Our methodology examines failure patterns across hardware components, failure types, propagation mechanisms, and systemwide impacts, thus providing insights into the interplay between hardware failures, resilience mechanisms, and application-level fault tolerance.  \nTogether, these contributions provide a holistic approach towards making AI infrastructure more efficient, adaptive, and resilient.  \nTo my family for their faith, love, and support.  \niii  \nAcknowledgments  \nFirst and foremost, I am grateful to my advisor, Prof. Ravi Iyer, for his constant support throughout my journey. Ravi has mentored me since my visiting-student days at UIUC as an undergrad and has been instrumental in s","cbCaitiaToz9ltIQ","https://ap.wps.com/l/cbCaitiaToz9ltIQ","pdf",3665227,1,131,"English","en",105,"# Abstract\n## Core Infrastructure Optimization\n## Workload Adaptation\n## Reliability Enhancement\n# Acknowledgments","[{\"question\":\"Why does the thesis focus on infrastructure efficiency for machine learning?\",\"answer\":\"Machine learning models achieve major breakthroughs but require vast compute resources, specialized accelerators, and high power and cost. Improving infrastructure efficiency is key to making these systems more accessible and sustainable.\"},{\"question\":\"What are the three interconnected dimensions addressed in the thesis?\",\"answer\":\"The thesis addresses core infrastructure optimization, workload adaptation, and reliability enhancement. Each dimension targets different bottlenecks in efficiency, utilization, and resilience.\"},{\"question\":\"How does the thesis improve reliability in AI infrastructure?\",\"answer\":\"It characterizes GPU failures using failure data from large-scale production clusters, analyzing patterns across hardware components, failure types, propagation mechanisms, and systemwide impacts to connect resilience mechanisms with application-level fault tolerance.\"}]","Towards Efficient and Reliable Infrastructure for Machine Learning - Dissertation | PDF",1785938996,330,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"towards-efficient-and-reliable-infrastructure-for-machine-learning-dissertation","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/towards-efficient-and-reliable-infrastructure-for-machine-learning-dissertation/127458/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why does the thesis focus on infrastructure efficiency for machine learning?","Question",{"text":75,"@type":76},"Machine learning models achieve major breakthroughs but require vast compute resources, specialized accelerators, and high power and cost. Improving infrastructure efficiency is key to making these systems more accessible and sustainable.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What are the three interconnected dimensions addressed in the thesis?",{"text":80,"@type":76},"The thesis addresses core infrastructure optimization, workload adaptation, and reliability enhancement. Each dimension targets different bottlenecks in efficiency, utilization, and resilience.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the thesis improve reliability in AI infrastructure?",{"text":84,"@type":76},"It characterizes GPU failures using failure data from large-scale production clusters, analyzing patterns across hardware components, failure types, propagation mechanisms, and systemwide impacts to connect resilience mechanisms with application-level fault tolerance.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]