[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-128131-en":3,"doc-seo-128131-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},128131,3985741905716,"Rowan","https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d",8,"Research & Report","Personalized and Efficient Distributed Machine Learning","Modern machine learning scales in two directions at once: edge data grows across heterogeneous devices, and model capacity expands to billions of parameters, stressing memory bandwidth and computation while risking numerical instability. This dissertation advances both personalization and efficiency for distributed learning. It develops statistical tools for client heterogeneity in supervised and unsupervised settings, and proposes meta-optimizers plus lower-precision training methods for large language models. The work includes QuPeD, hierarchical-Bayes generalizations, adaptive generation and dimensionality reduction, MADA, and stochastic rounding for BF16 training, improving throughput and validation perplexity.","UCLA  \nUCLA Electronic Theses and Dissertations  \nTitle  \nPersonalized and Efficient Distributed Machine Learning  \nPermalink  \n[https://escholarship.org/uc/item/5hn344n1](https://escholarship.org/uc/item/5hn344n1)  \nAuthor  \nOzkara, Kaan  \nPublication Date  \n2025  \nPeer reviewed|Thesis/dissertation  \n[eScholarship.org](eScholarship.org) Powered by the California Digital Library  \nUniversity of California  \nUNIVERSITY OF CALIFORNIA Los Angeles  \nPersonalized and Efficient Distributed Machine Learning  \nA dissertation submitted in partial satisfaction of the requirements for the degree Doctor of Philosophy in Electrical and Computer Engineering  \nby  \nKaan Ozkara  \n2025  \n© Copyright by Kaan Ozkara 2025  \nABSTRACT OF THE DISSERTATION  \nPersonalized and Efficient Distributed Machine Learning  \nby  \nKaan Ozkara  \nDoctor of Philosophy in Electrical and Computer Engineering University of California, Los Angeles, 2025  \nProfessor Suhas Diggavi, Chair  \nModern machine learning faces two simultaneous explosions of scale: i) vast amount of data generated at the edge, ii) number of model parameters have scaled up to billions, particularly for large language models (LLMs) . Firstly, data is now generated from an ecosystem of devices whose statistical characteristics and hardware capabilities differ widely, demanding a personalized distributed learning that enables collaboration of heterogeneous clients and is tailored according to specific needs. Secondly, model capacity has grown to billions of parameters, stretching the limits of memory bandwidth and computational capabilities, which necessitates efficient training methodologies while preserving numerical stability. The goal of this thesis is to advance both fronts. For personalization, we develop a statistical framework that enables a fundamental understanding and characterization of client heterogeneity both in supervised and unsupervised learning settings. Our methodologies allow collaboration of clients that may possess data from distinct distributions, and have different resource constraints. For large scale training of LLMs, we propose novel meta-optimizers that speeds up convergence, and lower precision training strategies that enable training with lower memory and faster throughput. The contributions of this thesis can be summarized as follows:  \n• We introduce QuPeD, Quantized Personalization via Distillation, where we introduce a relaxed quantization objective coupled with knowledge distillation thats lets each client learn a model compressed to its own precision and architectural constraints while enabling collaboration. We develop an alternating proximal gradient algorithm that outperforms prior personalized FL baselines across diverse vision and language tasks. The convergence of our method reveals the coupling between heterogeneity and resource constraints.  \n• We generalize the hierarchical Bayes framework to personalized FL. We unify disparate personalization techniques—regularization, model interpolation, clustering—under one statistical lens, which also yields AdaPeD, an information geometry regularized algorithm. Extending the framework, we provide user level differential privacy guarantees without sacrificing personalized performance. Furthermore, we utilize a similar framework for unsupervised learning; viewing client models through a hierarchical prior, we design adaptive algorithms for personalized dimensionality reduction and diffusion based generation. Our analysis reveals the provable utility obtained from collaboration for generative models, and experiments confirm these gains. For the personalized diffusion generative models we further propose an architectural personalization, using a shared backbone and individual client identities to steer the shared backbone.  \n• For LLM training, we develop MADA, which parameterizes a spectrum of adaptive optimizer rules (Adam, AMSGrad, etc.), and employ hyper gradient descent to select the best optimizer interpolation a","cbCaidupPByHm1Ih","https://ap.wps.com/l/cbCaidupPByHm1Ih","pdf",19377693,1,364,"English","en",105,"# 1 Introduction\n## 1.1 Contributions and Outline\n# 2 Quantized Personalization via Distillation\n## 2.1 Introduction\n## 2.2 Problem Formulation\n## 2.3 Centralized Model Quantization Training\n## 2.4 Personalized Quantization for FL via Knowledge Distillation\n## 2.5 Experiments\n## 2.6 Discussion\n# 3 A Statistical View of Personalized Learning\n## 3.1 Introduction\n## 3.2 Personalized Estimation\n## 3.3 Personalized Learning","[{\"question\":\"What two main scaling challenges does the dissertation address?\",\"answer\":\"It addresses heterogeneous data generation at the edge for personalization and the growth of model parameters to billions that requires efficient training while maintaining numerical stability.\"},{\"question\":\"How does the thesis approach personalization in federated learning?\",\"answer\":\"It introduces QuPeD, using quantized personalization via distillation, and generalizes hierarchical Bayes frameworks to unify multiple personalization methods, including privacy guarantees.\"},{\"question\":\"What methods are proposed for efficient large language model training?\",\"answer\":\"It proposes MADA, which adaptively interpolates optimizer rules using hyper-gradient descent, and uses stochastic rounding to enable a full BF16 training recipe with improved throughput and convergence.\"}]","Personalized and Efficient Distributed Machine Learning | PDF",1785944995,917,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"personalized-and-efficient-distributed-machine-learning","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/personalized-and-efficient-distributed-machine-learning/128131/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-23","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What two main scaling challenges does the dissertation address?","Question",{"text":76,"@type":77},"It addresses heterogeneous data generation at the edge for personalization and the growth of model parameters to billions that requires efficient training while maintaining numerical stability.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does the thesis approach personalization in federated learning?",{"text":81,"@type":77},"It introduces QuPeD, using quantized personalization via distillation, and generalizes hierarchical Bayes frameworks to unify multiple personalization methods, including privacy guarantees.",{"name":83,"@type":74,"acceptedAnswer":84},"What methods are proposed for efficient large language model training?",{"text":85,"@type":77},"It proposes MADA, which adaptively interpolates optimizer rules using hyper-gradient descent, and uses stochastic rounding to enable a full BF16 training recipe with improved throughput and convergence.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]