[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86440-en":3,"doc-seo-86440-105":29,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":11,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},86440,4398048950312,"Violet","https://ap-avatar.wpscdn.com/avatar/400002538284de19e3c?_k=1778320343897328908",8,"Research & Report","The Silent Freeze: Predicting When Low Precision Training Stops Learning","Reduced floating-point precision can silently halt model learning when gradient-descent updates become smaller than half a weight’s ULP, causing selected coordinates to round back onto the grid and freeze while gradients remain nonzero. The freeze is deterministic and per-coordinate, forecastable from a high-precision trajectory and the target mantissa length without low-precision data. Experiments show freezing in a small GPT with AdamW+cosine and in a 124M GPT-2 under 8-bit grid constraints; stochastic rounding mitigates it. The same rule transfers across regression, precision truncation emulation, small networks, and MNIST CNNs, identifying an arithmetic hazard distinct from noisy failure modes.","The Silent Freeze: Predicting When Low-Precision Training Stops Learning  \nZekai Shang 1, ∗  \n1 University of Illinois at Urbana-Champaign, Champaign, IL, USA  \n(Dated: July 14, 2026)  \narXiv :2607 .09800v 1 [ cs .LG] 9 Jul 2026  \nAbstract  \nTraining in reduced floating-point precision can silently halt learning: when a gradient-descent weight update falls below half the unit in the last place (ULP) of the weight, it rounds away and that coordinate freezes while its gradient is still nonzero. The freeze is deterministic, governed by a per-coordinate half-ULP condition, and predictable from a high-precision trajectory and the target mantissa length alone, without low-precision data. In a small GPT trained under the standard AdamW-plus-cosine recipe with bf16-equivalent stored weights, training proceeds normally and then permanently freezes just past mid-run, within four steps of the a-priori prediction. In a 124-million-parameter GPT-2 transformer whose weights are constrained to the 8-bit floating-point grid after every optimizer step, with no master weights, the dense weights freeze at initialization in both fp8 formats—predicted a priori from an fp32 reference—and validation loss plateaus while full precision keeps improving. Stochastic rounding removes the persistent freeze, and the same reference predicts that too. The condition transfers across frozen-feature regression, a mantissa-truncation emulator spanning 128 × in precision, small networks, and a CNN on MNIST: a computable axis of low-precision training, not diffuse noise.  \nReduced-precision arithmetic is now pervasive in training, and the bit width keeps falling: bfloat16 is routine and 8-bit floating point has reached frontier transformer training. Asthe bits have fallen, independent groups across scientific machine learning, low-precision optimizer design, and distributed post-training have collided with the same unnamed failure: training runs on, the loss is finite, the gradient is nonzero—and weights silently stop moving (Table I) . Each group engineered a workaround; none could say in advance when the failure would strike. Low-precision failure modes are usually treated as diffuse “noise.” This one is not noise at all but a deterministic, per-coordinate, and predictable event: the gradientunderflow freeze. It is, concretely, the failure mode that fp32 master weights exist to prevent and that pure low-mantissa weight updates expose. In gradient descent a weight is updated as w ← w − η∇ℓ; in finite precision the result is rounded back onto the floating-point grid. When the update η|∇ℓ| is smaller than half the unit in the last place (ULP) of w, the rounded result is w itself—the coordinate stops moving while its true gradient is still nonzero, and once enough coordinates cross this threshold the weight vector stops advancing even thoughthe loss gradient has not vanished. As ∥w∥ grows during training its ULP grows with it, so  \nTABLE I. Six independent reports of the same arithmetic hazard: weight updates too small to survive the destination grid. Each is a manifestation of half-ULP update swamping—in stored-weight training, mixed-precision pipelines, or cast-and-synchronize post-training—thoughthe surrounding mechanisms differ (Related work) . None predicts when the per-coordinate condition |∆wi| \u003C ~~1~~2ULPm(wi) engages, and only one [9] states the condition, as a gate for a fix rather than a forecast.  \n\n| Report | Observed symptom | Their response |\n| --- | --- | --- |\n| SciML PINNs, fp32 [4] | ∆w falls below fp32 ε; L-BFGS halts early | switch to fp64 |\n| SciML, mixed preci- | fp16 weight updates underflow to | fp32 master weights |\n| sion [6] | zero |  |\n| M+Adam [5] | bf16 Adam misaligned with the bf16 grid | additive mantissa + exponent updates |\n| bf16 fused Adam [7] | small updates cancel; weights stay stale | 16+16 storage (extra mantissa bits) |\n| Distributed RL, | ∼99% of per-step bf16 updates in- | exploited: sync only the |\n| bf16 [12] | visible af","cbCaiaSoEnDZoJRm","https://ap.wps.com/l/cbCaiaSoEnDZoJRm","pdf",504441,1,29,"English","en",105,"# Abstract\n## Gradient-descent half-ULP freezing mechanism\n## Predicting freeze onset from high precision trajectory\n## Evidence in GPT and GPT-2 with 8-bit constrained weights\n## Transfer across tasks and precision emulation\n## Relation to prior low-precision failure reports","[{\"question\":\"What experiments demonstrate this deterministic freeze behavior?\",\"answer\":\"A small GPT trained with AdamW-plus-cosine shows normal progress followed by permanent mid-run freezing within a few steps of the predicted time. A 124M GPT-2 with weights constrained to the 8-bit floating-point grid freezes dense weights at initialization in both fp8 formats, with validation loss plateauing.\"}]",1784211748,73,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":27},"the-silent-freeze-predicting-when-low-precision-training-stops-learning","",{"@graph":35,"@context":77},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/the-silent-freeze-predicting-when-low-precision-training-stops-learning/86440/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":11},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"What experiments demonstrate this deterministic freeze behavior?","Question",{"text":75,"@type":76},"A small GPT trained with AdamW-plus-cosine shows normal progress followed by permanent mid-run freezing within a few steps of the predicted time. A 124M GPT-2 with weights constrained to the 8-bit floating-point grid freezes dense weights at initialization in both fp8 formats, with validation loss plateauing.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,120,123,127],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":45,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":45,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":45,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":45,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":128,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":98,"slug":130},19,"General","general"]