[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-122310-en":3,"doc-seo-122310-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},122310,687197207057,"Sage","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Hardware Acceleration of Machine Learning - Evaluation and comparison of hardware-aware optimization techniques","The thesis addresses the computational demands of Transformer fine-tuning and focuses on optimizing and accelerating Transformer-based model training under multiple evaluation criteria. It evaluates training time, energy consumption, cost, and hardware utilization, comparing GPU training configurations against specialized AI accelerators such as Google TPUs. The work develops an optimized kernel for the Adan optimizer, applies LightSeq to accelerate Transformer components, and introduces mixed precision training. It further analyzes distributed multi-GPU training, includes backpropagation time estimation, and benchmarks workloads across V100, A100, A10, and T4 to compare performance, usability, reliability, and stability.","Hardware Acceleration of Machine Learning  \nEvaluation and comparison of different hardware-aware optimization techniques  \nMaster’s thesis in Computer Science and Engineering  \nFangzhou Chen William Sköld  \nDepartment of Computer Science and Engineering CHALMERS UNIVERSITY OF TECHNOLOGY UNIVERSITY OF GOTHENBURG  \nGothenburg, Sweden 2023  \nMaster’s thesis 2023  \nHardware Acceleration of Machine Learning  \nEvaluation and comparison of diﬀerent hardware-aware optimization  \ntechniques  \nFangzhou Chen  \nWilliam Sköld  \nDepartment of Computer Science and Engineering Chalmers University of Technology University of Gothenburg Gothenburg, Sweden 2023  \nHardware Acceleration of Machine Learning  \nEvaluation and comparison of diﬀerent hardware-aware optimization techniques Fangzhou Chen, William Sköld  \n© Fangzhou Chen, William Sköld, 2023 .  \nSupervisor: Pedro Petersen Moura Trancoso, Department of Computer Science and Engineering  \nAdvisor: Evangelos Siminos, Volvo Group (SML)  \nExaminer: Pedro Petersen Moura Trancoso, Department of Computer Science and Engineering  \nMaster’s Thesis 2023  \nDepartment of Computer Science and Engineering  \nChalmers University of Technology and University of Gothenburg SE-412 96 Gothenburg  \nTelephone +46 31 772 1000  \nCover: A running megatron, generated by DALLE2 .  \nTypeset in LATEX  \nGothenburg, Sweden 2023  \nHardware Acceleration of Machine Learning Fangzhou Chen  \nWilliam Skold  \nDepartment of Computer Science and Engineering  \nChalmers University of Technology and University of Gothenburg  \nAbstract  \nThe Transformer architecture has been widely used in various ﬁelds, as demonstrated by GPT-3, a large language model that shows impressive performance. However, achieving such excellent performance requires high computational capabilities. Therefore, improving the computational power of current machine learning systems is of great importance.  \nThis thesis aims to optimize and accelerate ﬁne-tuning of Transformer-based models while taking into account several evaluation criteria, such as training time, energy consumption, cost, and hardware utilization. Additionally, a comparison is made between GPU training settings and specialized AI accelerators, such as TPU training settings.  \nIn our study, a high-performance kernel for the Adan optimizer was introduced, and the LightSeq library is applied to accelerate existing Transformer components. We also introduce mixed precision training into our workﬂow and compare all these optimization techniques step by step with baseline performance. In addition, our analysis includes distributed training with multiple GPUs, and a backpropagation time estimation algorithm is introduced. Next, Google’s TPU accelerator is used to run our task, and its performance is compared to the similar GPU setup used in our study. Finally, the advantages and disadvantages of diﬀerent methods are systematically analyzed, while training on V100, A100, A10 and T4 with diﬀerent conﬁgurations. Meanwhile, the workﬂow between GPUs and TPUs is analyzed, illustrating the pros and cons of diﬀerent accelerators.  \nVarious weights for measuring optimization methods based on time, energy consumption, cost, and hardware utilization are proposed. Our analysis shows that optimal scores in all metrics can be achieved by implementing the optimized LightSeq model, kernel fusion for the Adan optimizer, and enabling mixed precision training. While training with TPU oﬀers certain advantages, such as large batch sizes when loading training data, the ease of use, reliability, and software stability of GPU training surpasses that of TPU training.  \nKeywords: Transformer, GPU, Distributed, Energy consumption, Fine-tuning, TPU.  \nAcknowledgements  \nWe would like to express our sincere gratitude to our supervisors, Evangelos and Pedro for their guidance, expertise and unwavering support throughout this thesis. Their mentorship, advice and constructive feedback have been instrumental in shaping the direction ","cbCaifiusDihaIsT","https://ap.wps.com/l/cbCaifiusDihaIsT","pdf",4525724,1,106,"English","en",105,"# Introduction\n## Machine Learning Hardware\n## Machine Learning Framework\n## Transformer usage at Volvo\n## Problem statement\n## Related Work\n## Aim and Objectives\n## Delimitations\n## Thesis Outline\n# Theory\n## Transformer\n## Deep Learning Optimizer","[{\"question\":\"What goal does the thesis pursue for Transformer models?\",\"answer\":\"To optimize and accelerate fine-tuning of Transformer-based models while considering training time, energy consumption, cost, and hardware utilization.\"},{\"question\":\"Which hardware platforms are compared in the experiments?\",\"answer\":\"GPU training settings are compared with specialized AI accelerators, especially Google TPU training, and results are benchmarked using devices such as V100, A100, A10, and T4.\"},{\"question\":\"What optimization techniques are introduced and how are they evaluated?\",\"answer\":\"A high-performance kernel for the Adan optimizer, LightSeq acceleration for Transformer components, and mixed precision training are introduced, then systematically compared step by step against baseline performance and other approaches.\"}]","Hardware Acceleration of Machine Learning - Evaluation and comparison of hardware-aware optimization techniques | PDF",1785809947,267,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"hardware-acceleration-of-machine-learning-evaluation-and-comparison-of-hardware-aware-optimization-techniques","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/hardware-acceleration-of-machine-learning-evaluation-and-comparison-of-hardware-aware-optimization-techniques/122310/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What goal does the thesis pursue for Transformer models?","Question",{"text":75,"@type":76},"To optimize and accelerate fine-tuning of Transformer-based models while considering training time, energy consumption, cost, and hardware utilization.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which hardware platforms are compared in the experiments?",{"text":80,"@type":76},"GPU training settings are compared with specialized AI accelerators, especially Google TPU training, and results are benchmarked using devices such as V100, A100, A10, and T4.",{"name":82,"@type":73,"acceptedAnswer":83},"What optimization techniques are introduced and how are they evaluated?",{"text":84,"@type":76},"A high-performance kernel for the Adan optimizer, LightSeq acceleration for Transformer components, and mixed precision training are introduced, then systematically compared step by step against baseline performance and other approaches.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]