[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-121292-en":3,"doc-seo-121292-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},121292,1099514067415,"Rowan","https://ap-avatar.wpscdn.com/avatar/100002539d78ffe74a7?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779092875211072502",8,"Research & Report","High-Performance, Energy-Efficient, and Scalable Accelerator Design for Emerging Machine Learning Applications","Machine learning (ML) is widely used across domains such as autonomous driving, scientific computing, and robotics, yet rapid growth in model complexity and data size is driving unprecedented computation and communication demands on today’s systems, especially for large language models. These demands are intensified by unstructured data and platform constraints. This dissertation studies accelerator designs for diverse ML workloads, including flexible chiplet-based communication fabrics for deep learning, architectures for irregular sparsity in GCNs, and analysis of intermediate feature reuse and its communication bottlenecks in GNNs.","Graduate Thesis and Dissertation post-2024  \n2024  \nHigh-Performance, Energy-Efficient, and Scalable Accelerator Design for Emerging Machine Learning Applications  \nLingxiang Yin  \nUniversity of Central Florida  \nFind similar works at: [https://stars.library.ucf.edu/etd2024](https://stars.library.ucf.edu/etd2024)  \nUniversity of Central Florida Libraries [http://library.ucf.edu](http://library.ucf.edu)  \nThis Dissertation is brought to you for free and open access by STARS. It has been accepted for inclusion in Graduate Thesis and Dissertation post-2024 by an authorized administrator of STARS. For more information, please contact [STARS@ucf.edu](STARS@ucf.edu).  \nSTARS Citation  \nYin, Lingxiang, \"High-Performance, Energy-Efficient, and Scalable Accelerator Design for Emerging Machine Learning Applications\" (2024) . Graduate Thesis and Dissertation post-2024. 86.  \n[https://stars.library.ucf.edu/etd2024/86](https://stars.library.ucf.edu/etd2024/86)  \nHIGH-PERFORMANCE, ENERGY-EFFICIENT, AND SCALABLE ACCELERATOR DESIGN FOR EMERGING MACHINE LEARNING APPLICATIONS  \nby  \nLINGXIANG YIN  \nB.S. Nanjing University of Science and Technology, 2011  \nM.S. Nanjing University of Science and Technology, 2014  \nM.S. The University of Texas at Dallas, 2016  \nA dissertation submitted in partial fulfilment of the requirements for the degree of Doctor of Philosophy in the Department of Electrical and Computer Engineering in the College of Engineering and Computer Science at the University of Central Florida  \nOrlando, Florida  \nFall Term  \n2024  \nMajor Professor: Hao Zheng  \n© 2024 Lingxiang Yin  \nii  \nABSTRACT  \nThe use of machine learning (ML) is pervasive in numerous application domains, such as autonomous driving, scientific computing, robotics, and among others. However, the continuous growth of ML model complexity and data size is posing unprecedented computation and communication demands on current computing systems, especially in the era of large language models. The problem is further compounded by unstructured data and technology limitations.  \nTo address these challenges, this dissertation research explores novel accelerator designs tailored for a wide range of machine learning applications. First, this research investigates a flexible communication fabric that can enable efficient training for deep learning applications in chiplet-based accelerators. Furthermore, this research explores an efficient accelerator architecture that can dynamically handle irregular sparsity in accelerating graph convolutional neural networks. Furthermore, the dissertation uncovers the extensive intermediate feature data reuse opportunities and their communication bottlenecks in complex graph neural network models. These innovative accelerator designs can deliver high-performance, energy-efficient, and scalable solutions for emerging machine learning workloads and advance their practical deployment at an unprecedented scale.  \nTo my beloved family, whose love has guided me through every step of this journey.  \nACKNOWLEDGMENTS  \nI would like to express my heartfelt gratitude to my advisor, Dr. Zheng, for his unwavering support and guidance throughout my Ph.D. journey. His mentorship has gone beyond academic guidance, as he has always cared about me as an individual, not just as a student. During challenging moments, his encouragement and patience helped me navigate difficulties and move forward. His thoughtful advice and genuine care have had a lasting impact on both my academic progress and personal growth. I feel truly fortunate to have had a mentor who exemplifies not only academic excellence but also kindness, humility, and integrity—qualities that I will carry with me in both my career and life.  \nI am deeply grateful to my committee members—Dr. Mingjie Lin, Dr. Rickard Ewetz, Dr. Fan Yao, and Dr. Qian Lou—for their invaluable dedication and expertise. I extend special thanks to Dr. Lin, whose insightful perspectives consistently pushed me to broaden the","cbCaivxavnchSp5E","https://ap.wps.com/l/cbCaivxavnchSp5E","pdf",4481190,1,153,"English","en",105,"# Chapter 1: Introduction\n# Chapter 2: Background and Motivation\n## Prevalent Machine Learning Models\n## Deep Neural Network (DNN) Models\n## Graph Neural Network (GNN) Models\n## Challenges in Distributed DNN Training\n## Parallelism in Distributed DNN Training\n## Model Parallelism\n## Data Parallelism\n## AllReduce in Gradient Synchronization\n## Challenges in GNN Acceleration\n## Irregularity and Workload Imbalance\n## Data Reuse and Communication Overhead\n## Heterogeneous Operations and Architectural Challenges\n# Chapter 3: ARIES: Accelerating Distributed Training in Chiplet-Based Systems via Flexible Interconnects\n## Introduction\n## Background\n## Chiplet-based Architectures\n## Proposed ARIES Design","[{\"question\":\"What key problem does the dissertation address in modern machine learning systems?\",\"answer\":\"It addresses the escalating computation and communication demands caused by increasing ML model complexity and data size, particularly in the era of large language models, along with complications from unstructured data and technology limits.\"},{\"question\":\"Which accelerator concepts are explored for deep learning training on chiplet-based systems?\",\"answer\":\"The research investigates a flexible communication fabric designed to enable efficient training for deep learning applications in chiplet-based accelerators.\"},{\"question\":\"How does the dissertation handle irregular sparsity and data movement in graph neural networks?\",\"answer\":\"It studies an efficient accelerator architecture that dynamically supports irregular sparsity in graph convolutional neural networks and examines intermediate feature reuse opportunities and their communication bottlenecks in complex GNN models.\"}]","High-Performance, Energy-Efficient, and Scalable Accelerator Design for Emerging Machine Learning Applications | PDF",1785734943,386,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"high-performance-energy-efficient-and-scalable-accelerator-design-for-emerging-machine-learning-applications","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/high-performance-energy-efficient-and-scalable-accelerator-design-for-emerging-machine-learning-applications/121292/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What key problem does the dissertation address in modern machine learning systems?","Question",{"text":75,"@type":76},"It addresses the escalating computation and communication demands caused by increasing ML model complexity and data size, particularly in the era of large language models, along with complications from unstructured data and technology limits.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which accelerator concepts are explored for deep learning training on chiplet-based systems?",{"text":80,"@type":76},"The research investigates a flexible communication fabric designed to enable efficient training for deep learning applications in chiplet-based accelerators.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the dissertation handle irregular sparsity and data movement in graph neural networks?",{"text":84,"@type":76},"It studies an efficient accelerator architecture that dynamically supports irregular sparsity in graph convolutional neural networks and examines intermediate feature reuse opportunities and their communication bottlenecks in complex GNN models.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]