[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119454-en":3,"doc-seo-119454-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},119454,1374391974468,"Eden","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Parallel algorithms for GPU in machine learning","This master’s thesis extends the Template Numerical Library (TNL), a modern C++ framework for high-performance numerical computing, by implementing two widely used machine learning components: K-means clustering and reverse-mode automatic differentiation (AD). The designs integrate with TNL abstractions to support CPU and GPU execution. The K-means work includes a baseline, a TNL::Segments-based variant, and a GPU-optimized implementation inspired by prior research. Experiments compare performance against cuML and TensorFlow, assess dataset-dependent strengths, and highlight current limitations such as missing higher-order differentiation support.","PARALLEL ALGORITHMS FOR GPU IN MACHINE LEARNING  \nBc. Samuel Križan  \nMaster’s thesis  \nFaculty of Information Technology Czech Technical University in Prague Department of Computer Systems Study program: Informatics  \nSpecialisation: Computer Systems and Networks  \nSupervisor: doc. Ing. Tomáš Oberhuber, Ph.D. May 9, 2025  \nAssignment of master’s thesis  \nTitle: Parallel algorithms for GPU in machine learning  \nStudent: Bc. Samuel Križan  \nSupervisor: doc. Ing. Tomáš Oberhuber, Ph. D.  \nStudy program: Informatics  \nBranch / specialization: Computer Systems and Networks Department: Department of Computer Systems  \nValidity: until the end of summer semester 2025/2026  \nInstructions  \n1. Get familiar with the TNL library for developing parallel algorithms for GPUs and multicore CPUs.  \n2. Learn about the data abstraction segments in the TNL library.  \n3. Using this abstraction, implement a parallel algorithm for k-means clustering suitable for execution on a GPU and compare it with another k-means clustering algorithm on a GPU, e.g., [1] .  \n4. Furthermore, in the TNL library, create a basic implementation of automatic  \ndifferentiation using the reverse mode [2,3] on a GPU and compare it with the TensorFlow library.  \n[1] Kruliš M., Kratochvíl M., Detailed analysis and optimization of CUDA k-means algorithm. In Proceedings of the 49th International Conference on Parallel Processing 2020, pp. 1-11, 2020.  \n[2] Nocedal J., Wright S. J., Numerical Optimization, Springer, 2006.  \n[3] Pevný T., Šmídl V., Zorek M., Heim N., Scientiﬁc programming in Julia, [https://](https://)[ ](https://)[juliateachingctu.github.io/Scienti](juliateachingctu.github.io/Scienti) ﬁc-Programming-in-Julia/dev/lecture   08/lecture/ .  \nElectronically approved by prof. Ing. Pavel Tvrdík, CSc. on 1 January 2025 in Prague.  \nCzech Technical University in Prague Faculty of Information Technology  \n© 2025 Bc. Samuel Križan. All rights reserved.  \nThis thesis is school work as defined by Copyright Act of the Czech Republic. It has been submitted at Czech Technical University in Prague, Faculty of Information Technology. The thesis is protected by the Copyright Act and its usage without author’s permission is prohibited (with exceptions defined by the Copyright Act) .  \nCitation of this thesis: Križan Samuel. Parallel algorithms for GPU in machine learning. Master’s thesis. Czech Technical University in Prague, Faculty of Information Technology, 2025 .  \nFirst, I would like to thank my supervisor, doc . Ing. Tomáš Oberhuber, Ph.D. , for his invaluable guidance, support, and the very pleasant and productive cooperation throughout the writing of this thesis.  \nSecondly, I would like to thank my significant other, Alica, who was my constant motivation and source of mental support during my studies.  \nLastly, I would like to thank doc . Ing. Ivan Šimeček, Ph.D. for his excellent introduction to the world of GPU and parallel programming, which has since grown on me .  \nDeclaration  \nI hereby declare that the presented thesis is my own work and that I have cited all sources of information in accordance with the Guideline for adhering to ethical principles when elaborating an academic final thesis. I acknowledge that my thesis is subject to the rights and obligations stipulated by the Act No. 121/2000 Coll., the Copyright Act, as amended, in particular the fact that the Czech Technical University in Prague has the right to conclude a licence agreement on the utilization of this thesis as a school work pursuant of Section 60 (1) of the Act. I declare that I have used AI tools during the preparation and writing of my thesis. I have verified the generated content. I confirm that I am aware that I am fully responsible for the content of the thesis.  \nIn Prague on May 9, 2025  \nAbstract  \nThis thesis extends the Template Numerical Library (TNL), a modern C++ framework for high-performance numerical computing, by implementing two core components widely used in machine learning:","cbCaiejlTtAAwxeq","https://ap.wps.com/l/cbCaiejlTtAAwxeq","pdf",583565,1,75,"English","en",105,"# Assignment of master’s thesis\n## TNL library and parallel algorithm tasks\n## Implement K-means on GPU and compare\n## Implement reverse-mode automatic differentiation and compare\n## Declaration and academic notes","[{\"question\":\"What two main components does the thesis add to TNL?\",\"answer\":\"The thesis adds a K-means clustering implementation and a reverse-mode automatic differentiation (AD) system integrated with TNL.\"},{\"question\":\"How is K-means implemented and evaluated for GPU execution?\",\"answer\":\"K-means is implemented as a baseline using core TNL structures, a variant using TNL::Segments, and a GPU-optimized version inspired by prior research. Experimental evaluation compares results to cuML and discusses dataset-size suitability.\"},{\"question\":\"How does the reverse-mode AD implementation compare with TensorFlow?\",\"answer\":\"The AD system is implemented as a graph-based reverse-mode engine compatible with TNL data types. Compared to TensorFlow, it shows advantages on smaller datasets, while TensorFlow dominates on larger workloads.\"}]","Parallel algorithms for GPU in machine learning | PDF",1785724364,189,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"parallel-algorithms-for-gpu-in-machine-learning","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/parallel-algorithms-for-gpu-in-machine-learning/119454/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05","2026-08-03",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What two main components does the thesis add to TNL?","Question",{"text":76,"@type":77},"The thesis adds a K-means clustering implementation and a reverse-mode automatic differentiation (AD) system integrated with TNL.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How is K-means implemented and evaluated for GPU execution?",{"text":81,"@type":77},"K-means is implemented as a baseline using core TNL structures, a variant using TNL::Segments, and a GPU-optimized version inspired by prior research. Experimental evaluation compares results to cuML and discusses dataset-size suitability.",{"name":83,"@type":74,"acceptedAnswer":84},"How does the reverse-mode AD implementation compare with TensorFlow?",{"text":85,"@type":77},"The AD system is implemented as a graph-based reverse-mode engine compatible with TNL data types. Compared to TensorFlow, it shows advantages on smaller datasets, while TensorFlow dominates on larger workloads.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]