[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-117523-en":3,"doc-seo-117523-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},117523,687197207057,"Sage","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Toward Efficient Machine Learning Systems with Sampling and Compression","Machine learning growth in model size, dataset volume, and task complexity introduces major computational and memory efficiency constraints. This dissertation presents sampling and compression techniques that improve efficiency across multiple application areas. It introduces CacheSample for efficient graph neural networks, FastSR-NeRF for accelerated 3D image rendering, and TeleRAG for optimized retrieval-augmented generation. It also proposes SPIN for convolutional neural networks, Atom for large language models, and Palu for key-value cache compression. Results reduce compute and memory while preserving accuracy, enabling scalable, accessible ML systems.","Toward Efficient Machine Learning Systems with Sampling and Compression  \nChien-Yu Lin  \nA dissertation submitted in partial fulfillment of the requirements for the degree of  \nDoctor of Philosophy  \nUniversity of Washington  \n2025  \nReading Committee:  \nLuis Ceze, Chair  \nBaris Kasikci  \nArvind Krishnamurthy  \nProgram Authorized to Offer Degree: Computer Science and Engineering  \n© Copyright 2025  \nChien-Yu Lin  \nUniversity of Washington  \nAbstract  \nToward Efficient Machine Learning Systems with  \nSampling and Compression  \nChien-Yu Lin  \nChair of the Supervisory Committee:  \nProfessor Luis Ceze  \nComputer Science and Engineering  \nThe rapid growth of machine learning-in terms of model size, dataset volume, and task complexity-has created significant computational and memory efficiency challenges. This thesis addresses these challenges by developing sampling and compression techniques across a variety of machine learning applications. Specifically, we introduce sampling strategies for efficient graph neural networks (CacheSample), accelerated 3D image rendering (FastSR-NeRF), and optimized retrieval-augmented generation systems (TeleRAG) . We further propose novel compression methods targeting convolutional neural networks (SPIN), large language models (Atom), and their key-value caches (Palu) . Collectively, these techniques substantially reduce computational and memory requirements while preserving model accuracy, facilitating the scalability and accessibility of machine learning systems. Finally, I present my vision for future efficiency innovations to ensure continued scalability and robustness as machine learning models continue to grow in complexity.  \nAcknowledgements  \nWhen I started my PhD in 2018, I had heard it would be hard—but I never imagined it would take nearly seven years and bring so many challenges. Along the way, I’ve been incredibly fortunate to have the support of many people. Without them, I would not have made it to the finish line.  \nFirst and foremost, I want to sincerely thank my advisor, Luis Ceze. Thank you for your guidance and support throughout my PhD, and for taking me on when I decided to focus more on machine learning systems after my first year. That decision opened the door to some of the most exciting research I’ve ever done. Thank you for believing in me—even during the times when I was stuck. Your support meant a lot and helped me keep going.  \nI’m also grateful to Baris Kasikci. Though I’m not officially your student, thank you for your generosity in mentoring me and offering invaluable feedback in our discussions and paper writing.  \nThank you to Arvind Krishnamurthy for being on my committee from the start. I really appreciated the opportunity to help design the ML systems course in Fall 2024—it was a fun and formative teaching experience. I also thank Luke Zettlemoyer and Arka Majumdar for bringing valuable perspectives from different domains to my committee, and Stephanie Wang for your mentorship in navigating collaborationsand helping organize our research group.  \nTo my mentors and collaborators at Apple—Carlo, Anish, Thomas, Qichen, Karren, Anurag, Sachin, and Max—thank you for working with me and devoting time to our joint projects. Those experiences gave me confidence that I can contribute meaningfully to impactful research.  \nThanks to all my collaborators in the Sampl Lab. A special shout-out to Zihao—your ability to tackle complex problems continues to amaze me. Working with you expanded my perspective and inspired me to take on harder challenges. To the early Sampl Lab crew—Liang Luo, Luis Vega, Thierry Moreau,  \nMegan, and Eddie—thank you for making school life fun and welcoming. And to the post-COVID lab members—Yile, Dedong, Liangyu, Kan, and Rohan, Zechou, Haoran and more—I’ve really appreciated your friendliness, whether collaborating or just chatting about life.  \nI’m also thankful for the interns I’ve had the pleasure of working closely with: Yilong Zhao, Chi-Chih ","cbCaifOvbCqSe9MI","https://ap.wps.com/l/cbCaifOvbCqSe9MI","pdf",15435133,1,200,"English","en",105,"# Chapter 1 Introduction\n## Efficient Machine Learning Systems with Sampling\n# Chapter 2 Accelerating Graph Neural Networks Inference with Edge Sampling\n## Edge Sampling for GNN’s Inference\n## CacheSample Kernel Design\n## Evaluation\n## Discussion and Future Work","[{\"question\":\"What efficiency challenges does the dissertation focus on?\",\"answer\":\"It focuses on computational and memory efficiency problems caused by increasing model size, dataset volume, and task complexity in machine learning.\"},{\"question\":\"Which sampling techniques are proposed, and for what applications?\",\"answer\":\"It proposes CacheSample for efficient graph neural network inference, FastSR-NeRF for accelerated 3D image rendering, and TeleRAG for optimized retrieval-augmented generation systems.\"},{\"question\":\"What compression methods are introduced and what do they target?\",\"answer\":\"It introduces SPIN targeting convolutional neural networks, Atom targeting large language models, and Palu targeting key-value caches to reduce compute and memory while preserving accuracy.\"}]","Toward Efficient Machine Learning Systems with Sampling and Compression | PDF",1785676634,504,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"toward-efficient-machine-learning-systems-with-sampling-and-compression","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/toward-efficient-machine-learning-systems-with-sampling-and-compression/117523/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What efficiency challenges does the dissertation focus on?","Question",{"text":75,"@type":76},"It focuses on computational and memory efficiency problems caused by increasing model size, dataset volume, and task complexity in machine learning.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which sampling techniques are proposed, and for what applications?",{"text":80,"@type":76},"It proposes CacheSample for efficient graph neural network inference, FastSR-NeRF for accelerated 3D image rendering, and TeleRAG for optimized retrieval-augmented generation systems.",{"name":82,"@type":73,"acceptedAnswer":83},"What compression methods are introduced and what do they target?",{"text":84,"@type":76},"It introduces SPIN targeting convolutional neural networks, Atom targeting large language models, and Palu targeting key-value caches to reduce compute and memory while preserving accuracy.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]