[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85000-en":3,"doc-seo-85000-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85000,13056703019404,"Miles","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining","Gradient communication is a primary scaling bottleneck in large language model pretraining. Lowering gradient precision (e.g., FP8, NVFP4) reduces communication volume, but Euclidean-space quantization can harm performance due to direction-dependent distortions from highly anisotropic gradients. GIFT introduces geometry-informed gradient scaling that quantizes in geometry-aware coordinates, transforming gradients into a near-isotropic space before FP8 communication. A simplified geometry-aware transform uses low-rank approximation and selective application to balance overhead. Experiments on Llama-300M/600M show faster pretraining and better downstream preservation than direct Euclidean FP8.","GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining  \nJieying Wang 1 , Shuyuan Fan2 , Mingkai Zheng2 , and Zhao Zhang2  \n1Department of Computer Science, Rutgers University, Piscataway, NJ, USA  \n2Department of Electrical and Computer Engineering, Rutgers University, Piscataway, NJ, USA  \narXiv :2607 .07494v 1 [ cs .DC] 8 Jul 2026  \nAbstract—Gradient communication is a primary scaling bottleneck in large language model (LLM) pretraining. Communicating gradients in low-precision formats, such as FP8 and NVFP4, can significantly reduce the communication volume. Existing methods quantize gradients via linear or nonlinear mappingsin Euclidean space, often degrading model performance because highly anisotropic gradients incur direction-dependent distortion. We present GIFT, a geometry-informed gradient scaling method that performs low-precision communication in geometry-aware coordinates. By transforming gradients into a near-isotropic space before quantization, GIFT makes low-precision representations substantially more faithful to their high-precision counterparts. GIFT only changes the coordinate system used for low-precision gradient communication and does not change the optimizer, training recipe, communication collective, or lowprecision format. We also develop a simplified geometry-aware transformation algorithm with low-rank approximation and selective application to balance the computation overhead and communication reduction. We examine the empirical convergence of GIFT using Llama-300M and Llama-600M models. Our results show that GIFT reduces the end-to-end pretraining time of Llama-600M by 7.6% on 64 NVIDIA GH200 Superchips, while improving the downstream task preservation profile over direct Euclidean FP8 communication under the same optimizer and communication path.  \nIndex Terms—large language models, distributed pretraining, gradient communication, communication compression, lowprecision communication, geometry-aware communication, KFAC  \nI. INTRODUCTION  \nGradient communication is a major scaling bottleneck in distributed LLM pretraining [1] . In data-parallel, 3D-parallel, or more complex parallel pretraining strategies, the optimizer needs to communicate gradients across graphics processing units (GPUs) . Standard communication libraries, such as NCCL and MPI, implement the ring allreduce algorithm [2], whose cost depends on the model size and job size (i.e., the number of GPUs) . Previous work reports that communication accounts for 40% of the overall pretraining time for GPT-8.3Bpretraining across 128 NVIDIA A100 GPUs [3] .  \nReducing communication volume via low-precision gradients is an effective way to lower communication overhead, asthe bits saved for each gradient representation directly translate into proportional reductions in bandwidth consumption. Existing low-precision training techniques, such as FP8-LM [4] and COAT [5], mainly apply scaling and quantization in Euclidean tensor coordinates, which can still lead to degraded model performance. They are designed specifically for the AdamW  \noptimizer [6] . FP8-LM uses FP8 format for model weights, activations, gradients, and first moment, while the second moment is stored in 16-bit format for division stability. COAT uses dynamic range strategies (i.e., SX = ~~ ~~amF~~3~~XP~~32~~)) and mixed-granularity activation quantization to quantize second moment and activation, while the gradients are still in 16-bit format for training stability. SDP4Bit [7] leverages the sharded data parallel model placement and implements 4-bit gradient communication via Fourier transform and hierarchical allto-all communication. However, neither FP8-LM nor COAT achieves comparable model performance to the BF16 baseline. FP8-LM shows that seven out often downstream tasks perform worse than the baseline. COAT has only one with a higher score than the baseline among the four downstream tasks. SDP4Bit only evaluates model performance using validation ","cbCaidNqGWamTvla","https://ap.wps.com/l/cbCaidNqGWamTvla","pdf",1321277,4,1,12,"English","en",105,"# Introduction\n## Gradient communication bottlenecks\n## Related low-precision gradient methods\n## Motivation for geometry-aware scaling\n# Method Overview","[{\"question\":\"Why is gradient communication a bottleneck in distributed LLM pretraining?\",\"answer\":\"In data-parallel or more complex parallel strategies, the optimizer must exchange gradients across GPUs using ring allreduce, whose cost grows with model and job size. Prior reports show communication can consume a large fraction of total pretraining time.\"},{\"question\":\"What limitation affects existing low-precision gradient quantization methods?\",\"answer\":\"Most methods quantize gradients in Euclidean tensor coordinates using linear or nonlinear mappings, which can degrade performance. Highly anisotropic gradients can be distorted in a direction-dependent way, causing the synchronized updates to drift from high-precision behavior.\"},{\"question\":\"How does GIFT improve fidelity and efficiency for FP8 gradient communication?\",\"answer\":\"GIFT changes only the coordinate system used for low-precision gradient communication. It transforms gradients into a near-isotropic, geometry-aware space before quantization, so FP8 scaling introduces more comparable errors across dimensions, improving downstream task preservation while reducing pretraining time.\"}]",1784200143,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"gift-geometry-informed-low-precision-gradient-communication-for-llm-pretraining","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/gift-geometry-informed-low-precision-gradient-communication-for-llm-pretraining/85000/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is gradient communication a bottleneck in distributed LLM pretraining?","Question",{"text":75,"@type":76},"In data-parallel or more complex parallel strategies, the optimizer must exchange gradients across GPUs using ring allreduce, whose cost grows with model and job size. Prior reports show communication can consume a large fraction of total pretraining time.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What limitation affects existing low-precision gradient quantization methods?",{"text":80,"@type":76},"Most methods quantize gradients in Euclidean tensor coordinates using linear or nonlinear mappings, which can degrade performance. Highly anisotropic gradients can be distorted in a direction-dependent way, causing the synchronized updates to drift from high-precision behavior.",{"name":82,"@type":73,"acceptedAnswer":83},"How does GIFT improve fidelity and efficiency for FP8 gradient communication?",{"text":84,"@type":76},"GIFT changes only the coordinate system used for low-precision gradient communication. It transforms gradients into a near-isotropic, geometry-aware space before quantization, so FP8 scaling introduces more comparable errors across dimensions, improving downstream task preservation while reducing pretraining time.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]