[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-117407-en":3,"doc-seo-117407-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},117407,7971461740909,"Levi","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","TDML-A Trustworthy Distributed Machine Learning Framework","Deep learning progress has accelerated alongside large generative models, but training large models increasingly bottlenecks on scarce and uneven GPU resources. Distributed machine learning methods such as federated learning reduce dependence on centralized data, yet practical deployment remains hard, especially when combining parallelism optimizations with flexible remote training and ensuring trustworthy execution. The proposed TDML framework coordinates remote trainers and validates workloads using blockchain, targeting privacy, transparency, and efficient model training across public distributed computing resources. Experiments confirm improved performance limits and effective malicious node detection.","TDML-A Trustworthy Distributed Machine Learning Framework  \nZhen Wang, Qin Wang, Guangsheng Yu, Shiping Chen  \nCSIRO Data61, Australia  \narXiv :2407 .07339v 1 [ cs .CR] 10 Jul 2024  \nABSTRACT  \nRecent years have witnessed a surge in deep learning research, marked by the introduction of expansive generative models like OpenAI’s SORA and GPT, Meta AI’s LLAMA series, and Google’s FLAN, BART, and Gemini models. However, the rapid advancement of large models (LM) has intensified the demand for computing resources, particularly GPUs, which are crucial for their parallel processing capabilities. This demand is exacerbated by limited GPU availability due to supply chain delays and monopolistic acquisition by major tech firms. Distributed Machine Learning (DML) methods, such as Federated Learning (FL), mitigate these challenges by partitioning data and models across multiple servers, though implementing optimizations like tensor and pipeline parallelism remains complex. Blockchain technology emerges as a promising solution, ensuring data integrity, scalability, and trust in distributed computing environments, but still lacks guidance on building practical DML systems. In this paper, we propose a trustworthy distributed machine learning (TDML) framework that leverages blockchain to coordinate remote trainers and validate workloads, achieving privacy, transparency, and efficient model training across public remote computing resources. Experimental validation demonstrates TDML’s efficacy in overcoming performance limitations and malicious node detection, positioning it as a robust solution for scalable and secure distributed machine learning.  \nKEYWORDS  \nFederated learning, Distributed, Blockchain, Trust, Large Model  \n1 INTRODUCTION  \nThere has been a remarkable surge in deep learning research and its practical applications. Leading tech giants have unveiled their expansive generative models, with examples like OpenAI’s SORA and GPT. Meta AI unveiled the LLAMA series, the world’s first open-source large language model. Google introduced its language models, such as FLAN, BART, and Gemini.  \nWith the rapid advancement of large models (LM), computing resources have become a critical bottleneck in the AI domain. Graphics Processing Units (GPUs) are favored for AI computing due to their parallel infrastructure and ability to process data simultaneously, making them indispensable for machine learning tasks. However, the limited number of companies involved in GPU development and distribution creates significant delays in the manufacturing supply chain. Moreover, major tech companies in cloud computing exacerbate the shortage of computing resources for smaller organizations by prioritizing the acquisition of the majority of GPUs. For example, OpenAI and Microsoft plan to invest USD 100bin GPUs by 2027 to enhance their data center capabilities. Meta’s Llama 3 models are trained on two clusters, each with 24,576 H100 GPUs, and Meta intends to acquire an additional 350,000 Nvidia H100 GPUs for over USD 10b. This unequal competitive landscape  \nhampers the ability of AI startups to construct large deep learning models and compete on an even playing field.  \nDistributed Machine Learning (DML) integrates distributed computing resources to provide fast learning capabilities, particularly for tasks involving large-scale data or extensive model parameters. This method partitions the training data and the model, with parameter servers coordinating multiple clients to learn each partition asa subtask. Federated Learning (FL) [10] exemplifies data parallelism in DML, coordinating the distributed training process using local data and aggregating a global model on a central server. Model parallelism training methods have been applied to solve many real problems that deal with large model-distributed training systems. The training network systems involve numerous connected computing and storage units. In order to reduce the model training complexity, variou","cbCaipAGDP2T9Iq1","https://ap.wps.com/l/cbCaipAGDP2T9Iq1","pdf",1853195,1,10,"English","en",105,"# Introduction\n# Background: GPU Bottlenecks and DML\n## Federated Learning and Model Parallelism\n## Challenges in Existing Frameworks\n# Blockchain for Trust in Distributed ML\n## Blockchain Fundamentals and Immutability\n## Smart Contracts and Verifiable Coordination\n## Blockchain-based Federated Learning","[{\"question\":\"Why is computing resources, especially GPUs, a bottleneck for large model training?\",\"answer\":\"GPU availability is limited by supply chain delays and by major cloud providers prioritizing GPU allocation for large players, which constrains smaller organizations and startups.\"},{\"question\":\"What problem does TDML aim to solve compared with traditional distributed ML frameworks?\",\"answer\":\"TDML targets the lack of guidance for building practical DML systems that support open and flexible remote training while ensuring privacy, transparency, and trustworthy execution.\"},{\"question\":\"How does TDML use blockchain to improve distributed machine learning trustworthiness?\",\"answer\":\"TDML leverages blockchain to coordinate remote trainers and validate workloads, using cryptographic immutability and smart contracts to protect against tampering and to enable verifiable transactions without a central intermediary.\"}]","TDML-A Trustworthy Distributed Machine Learning Framework | PDF",1785675711,25,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"tdml-a-trustworthy-distributed-machine-learning-framework","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/tdml-a-trustworthy-distributed-machine-learning-framework/117407/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is computing resources, especially GPUs, a bottleneck for large model training?","Question",{"text":75,"@type":76},"GPU availability is limited by supply chain delays and by major cloud providers prioritizing GPU allocation for large players, which constrains smaller organizations and startups.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What problem does TDML aim to solve compared with traditional distributed ML frameworks?",{"text":80,"@type":76},"TDML targets the lack of guidance for building practical DML systems that support open and flexible remote training while ensuring privacy, transparency, and trustworthy execution.",{"name":82,"@type":73,"acceptedAnswer":83},"How does TDML use blockchain to improve distributed machine learning trustworthiness?",{"text":84,"@type":76},"TDML leverages blockchain to coordinate remote trainers and validate workloads, using cryptographic immutability and smart contracts to protect against tampering and to enable verifiable transactions without a central intermediary.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":21,"slug":133},"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]