[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83007-en":3,"doc-seo-83007-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83007,7971461740886,"Theodore","https://ap-avatar.wpscdn.com/davatar_3d24733baf745e90a7e4bdd5f77d97b2",8,"Research & Report","Differentially Private Natural Gradient Descent","Under a fixed privacy budget, the utility of differentially private (DP) training hinges on optimization efficiency. Standard first-order DP methods such as DP-SGD ignore loss curvature, causing zigzagging in ill-conditioned deep-learning landscapes and wasting privacy budget on uninformative steps. Natural Gradient Descent (NGD) preconditions gradients with curvature to align updates with the loss geometry, improving convergence and utility. Direct DP-NGD integration is challenged by expensive curvature estimation, incompatibility between isotropic DP noise and NGD’s anisotropic scaling, and instability from inverse-curvature amplification. DP-NGD addresses these issues by decoupling curvature estimation from private data, using a whitened-space mechanism, and dynamically clamping curvature. Experiments show state-of-the-art accuracy and up to 10× faster convergence under the same privacy budget.","Differentially Private Natural Gradient Descent  \nLi Pan  \nInstitute of Information Engineering, Chinese Academy of Sciences Beijing, China [lipan@iie.ac.cn](lipan@iie.ac.cn)  \nChen Kai  \nInstitute of Information Engineering, Chinese Academy of Sciences Beijing, China [chenkai@iie.ac.cn](chenkai@iie.ac.cn)  \nChang Shuai  \nInstitute of Information Engineering, Chinese Academy of Sciences Beijing, China [changshuai751x@iie.ac.cn](changshuai751x@iie.ac.cn)  \nZhang Sheng Zhi  \nDepartment of Computer Science, Metropolitan College, Boston University Boston, United States [shengzhi@bu.edu](shengzhi@bu.edu)  \nLv Pei Zhuo  \nNanyang Technological University Singapore, Singapore [lvpeizhuo@gmail.com](lvpeizhuo@gmail.com)  \nHe Jin Wen  \nInstitute of Information Engineering, Chinese Academy of Sciences Beijing, China [hejinwen@iie.ac.cn](hejinwen@iie.ac.cn)  \narXiv :2607 .05866v 1 [ cs .LG] 7 Jul 2026  \nABSTRACT  \nUnder a fixed privacy budget, the utility of differentially private (DP) training is ultimately determined by its optimization efficiency. Standard first-order DP optimizers such as DP-SGD rely solely on local gradients and ignore the underlying loss curvature. This geometric blindness causes severe zigzagging in ill-conditioned landscapes, squandering precious privacy budgets on inefficient iterations. Practitioners are thus trapped in a bind: either stop training prematurely or inject massive per-step noise, both of which critically compromise final model utility. Natural Gradient Descent (NGD) resolves this by preconditioning gradients with curvature, aligning updates with the loss geometry and extracting more efficient signal from every noisy step, offering a principled pathway to break the privacy-utility bottleneck.  \nDespite its theoretical appeal, directly integrating NGD with DP introduces fundamental challenges: curvature estimation itself consumes prohibitive privacy budgets, isotropic DP operations conflict with the anisotropic scaling of NGD, and the inverse curvature catastrophically amplify parameter updates in flat directions, causing training instability. We propose DP-NGD, a practical framework that systematically addresses these obstacles by decoupling curvature estimation from private data, reconciling isotropic DP constraints with anisotropic second-order optimization via a whitened-space mechanism, and dynamically clamping the curvature to stabilize training. Extensive experiments on standard benchmarks demonstrate that DP-NGD achieves state-of-the-art accuracy, breaking through the utility ceilings of first-order baselines while delivering up to a 10× convergence speedup under the same privacy budget.  \nPVLDB Reference Format:  \nLi Pan, Chen Kai, Chang Shuai, Zhang Sheng Zhi, Lv Pei Zhuo, and He Jin Wen. Differentially Private Natural Gradient Descent. PVLDB, 14(1): XXX-XXX, 2026 .  \ndoi:XX.XX/XXX.XX  \nThis work is licensed under the Creative Commons BY-NC-ND 4.0 International License. Visit [https://creativecommons.org/licenses/by-nc-nd/4.0/ to view a copy of](https://creativecommons.org/licenses/by-nc-nd/4.0/ to view a copy of)[ ](https://creativecommons.org/licenses/by-nc-nd/4.0/ to view a copy of)[this license. For any use beyond those covered by this license](this license. For any use beyond those covered by this license), [obtain permission by](obtain permission by)[emailing info@vldb.org. Copyright](emailing info@vldb.org. Copyright) is held by the owner/author(s). Publication rights licensed to the VLDB Endowment.  \nProceedings of the VLDB Endowment, Vol. 14, No. 1 ISSN 2150-8097 . doi:XX.XX/XXX.XX  \nPVLDB Artifact Availability:  \nThe source code, data, and/or other artifacts have been made available at [https://github.com/XXX](https://github.com/XXX).  \n1 INTRODUCTION  \nDifferential Privacy (DP) [14, 15] has emerged as the gold standard for privacy-preserving machine learning. Under a fixed privacy budget, the final utility of DP training is bottlenecked by its optimization efficiency. Standard fir","cbCaifU4m4BXOEK9","https://ap.wps.com/l/cbCaifU4m4BXOEK9","pdf",6532207,3,1,13,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"Why do standard DP optimizers like DP-SGD perform poorly under a fixed privacy budget?\",\"answer\":\"They rely only on local gradients and ignore loss curvature, leading to severe zigzagging and slow convergence. Because each iteration consumes privacy budget, inefficient progress wastes the budget and harms final utility.\"},{\"question\":\"How does Natural Gradient Descent improve DP training utility?\",\"answer\":\"NGD preconditions gradients with a curvature matrix (e.g., Fisher information), aligning updates with the loss geometry. This accelerates convergence in ill-conditioned landscapes, which increases utility under the same privacy guarantee.\"},{\"question\":\"What key obstacles arise when combining NGD with differential privacy?\",\"answer\":\"Curvature estimation consumes prohibitive privacy budget, isotropic DP operations conflict with NGD’s anisotropic scaling, and inverse curvature can catastrophically amplify updates in flat directions, causing instability.\"}]",1784184634,33,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"differentially-private-natural-gradient-descent","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/differentially-private-natural-gradient-descent/83007/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do standard DP optimizers like DP-SGD perform poorly under a fixed privacy budget?","Question",{"text":75,"@type":76},"They rely only on local gradients and ignore loss curvature, leading to severe zigzagging and slow convergence. Because each iteration consumes privacy budget, inefficient progress wastes the budget and harms final utility.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does Natural Gradient Descent improve DP training utility?",{"text":80,"@type":76},"NGD preconditions gradients with a curvature matrix (e.g., Fisher information), aligning updates with the loss geometry. This accelerates convergence in ill-conditioned landscapes, which increases utility under the same privacy guarantee.",{"name":82,"@type":73,"acceptedAnswer":83},"What key obstacles arise when combining NGD with differential privacy?",{"text":84,"@type":76},"Curvature estimation consumes prohibitive privacy budget, isotropic DP operations conflict with NGD’s anisotropic scaling, and inverse curvature can catastrophically amplify updates in flat directions, causing instability.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]