[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119373-en":3,"doc-seo-119373-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},119373,1099514067438,"River Wang","https://ap-avatar.wpscdn.com/avatar/100002539ee87300030?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780474512215547542",8,"Research & Report","Gradient directions and relative inexactness in optimization and machine learning","Investigates how noise affects gradient estimates whose angle is acute with the true gradient in optimization and machine learning. Develops a relative noise model and provides theoretical calculations, theorems, and experimental results. Studies first-order methods under smoothness and relative error conditions, including convergence-rate preservation for classic gradient descent up to constants. Experiments use standard machine learning tasks such as linear and logistic regression plus vision and NLP, with modern deep learning models and empirically estimated noise parameters.","arXiv :2407 .00667v1 [math .OC] 30 Jun 2024  \nGradient directions and relative inexactness in optimization and machine learning  \nArtem Vasin a  \na Moscow Institute of Physics and Technology, Dolgoprudny, Russia;  \nARTICLE HISTORY  \nCompiled July 2, 2024  \nABSTRACT  \nIn this paper, we investigate the influence of noise giving an estimate of the gradi  \nent having a acute angle with the original. Noise amplitude has a relative model.  \nThe work offers both theoretical calculations and theorems, as well as experimental  \nresults. Classic machine learning problems were chosen as experiments-linear and  \nlogistic regression, computer vision and natural language processing.  \n1. Introduction  \nWe consider global optimization problem:  \nmin f(x) . (1)  \nx∈Rn  \nWe define f ∗ -as minimum value of f or solution for problem 1 and also for iterative methods with starting point x0 we can define R = assume that the objective f is L-smooth i.e., for all x, y ∈ Rn:  \n∥∇f(y) − ∇f(x)∥2 ⩽ L∥y − x∥2 , Or equivalent and more usefull:, f (y) ⩽ f (x) + ⟨∇f(x), y − x⟩ + L2 ∥x − y∥22  \nAlso we consider stochastic optimization:  \nf (x) → min ,  \nx∈Rn E hf (x)|xi = ∇f(x) .  \nx∗ : f (x∗ ) =∥x0 − x∗ ∥2 .  \nf ∗ ,  \nWe  \n(2)  \n(3)  \nThe classical formulation of the machine learning problem is the sum-structured problem:  \nf (θ) = 1m Xi fi(θ) → inRn . (4)  \nHere θ -model parameters and fi -loss function on sample i. Stochastic optimization in this problem is defined as follows:  \nI ∼ U{1 ... m} × · · · × U{1 ... m},  \n(θ) = 1B Xi∈I fi(θ) .  \nWhere B-batch size. We introduce two model errors in gradient:  \n∥f (x) − ∇f(x)∥2 ⩽ δ, (absolute error) or (5)  \n∥f (x) − ∇f(x)∥2 ⩽ α∥∇f(x)∥2 , (relative error) . (6)  \nFor 6 model we will use the following condition:  \nWe can inte(∀rxpr∈ Rnet the) c: o⟨nifti(xon) ,7∇fas(lxo)⟩we⩾rγbo∥fd(xfo)r∥2c∇fsine(xan) ∥2gle, bγetee(0n, 1gr] adient a(7nd)  \nit is estimation. Follow [2, 15] we can define (ν,ρ)– noise growth condition:  \nWe can note, that(∀xrela∈tiRvenm) :oνde∥∇l 6fw(x)∥ith2α ∥[; f1)(xim) ∥2pl⩽iesρ∇afnd(xg)r∥2owth model 8 wit(8h):  \nγ = p 1 − α2  ,ν = 1 − α,ρ = 1 + α . (9)  \nWe propose studies of the convergence of first-order methods with conditions 2, 7, 8 . We will prove, that classic gradient procedure 1 will preserves the order of convergence up to constants:  \nf (xN ) − f∗ = O 􀀒 νρ22γ2 LRN2 􀀓  \nIn Sections 5, 6 provided motivation and relevant to [10], [6] results, associated with p ~~µ~~L, where µ -constant of strong convexity 14 .  \nPaper contains a sufficient number of experiments with modern deep learning models with coefficients α,γ estimation. Sufficient conditions for the dataset are also given that guarantee the conditions 6, 7 .  \n2. Ideas behind the results  \nMost papers consider absolute model 5, but what δ should we choose for theoretical estimation of convergence. For example Algorithm 1 has convergence (for convex function with 2):  \nf (xN ) − f∗ = O 􀀒 LRN2 + δ􀀓 .  \nIf estimation for δ is large the theoretical convergence will be uninformative, but if we plot convergence plot we will see decreasing graph. As an example we can take dataset CIFAR-10 [11] for classification problem. Dataset consist of 50k training samples and 10k test samples of 32x32 images with 10 classes. We will use ResNet-18 [8] as classification model and PyTorch framework [12], because it provides batching. We will estimate α,δ and γ coefficient on each iteration by transforming epochs to iterations. Iterate over all batches (dataloader in PyTorch) we can summarize gradients per batches to gradient for whole train dataset, then choosing single batch we can evaluate required values.  \nFigure 1 . Convergence and coefficient evaluation per iterations for ResNet-18 CIFAR-10 .  \nWe can see, that γ -0.45, gives intuition to explore convergence with such conditions. More experiments provided at Section 8 . We should note, that for stochastic optimdetailizsatiato[ n16]3 convergence can be much better, using only δ∗ = E∥f (x∗ )∥2 , more","cbCaifuBeeHMQBS0","https://ap.wps.com/l/cbCaifuBeeHMQBS0","pdf",1544827,1,22,"English","en",105,"# Introduction\n## Global optimization and smoothness assumptions\n## Stochastic optimization and sum-structured learning\n## Absolute vs relative gradient noise models\n# Ideas behind the results\n## Choosing noise parameters for theoretical bounds\n## Experiments on CIFAR-10 with ResNet-18\n# Motivation for relative noise\n## Noise models via stochastic differential equations\n## Fokker-Planck stationary distributions\n# Gradient descent\n## Algorithm and convergence lemmas","[{\"question\":\"What noise setting does the paper analyze for gradient estimates?\",\"answer\":\"It studies noise in gradient estimates that maintains an acute angle with the original gradient, using a relative inexactness model (relative error proportional to the gradient magnitude). It also contrasts this with an absolute error model.\"},{\"question\":\"How does the paper model stochastic optimization in machine learning?\",\"answer\":\"It treats stochastic optimization using sum-structured objectives over samples and defines minibatch gradients by sampling an index set. This yields an expected objective related to the true gradient.\"},{\"question\":\"What do the results claim about classic gradient descent under the relative noise condition?\",\"answer\":\"The paper proves that classic gradient descent preserves the order of convergence up to multiplicative constants when the relative noise conditions are satisfied.\"}]","Gradient directions and relative inexactness in optimization and machine learning | PDF",1785723979,55,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"gradient-directions-and-relative-inexactness-in-optimization-and-machine-learning","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/gradient-directions-and-relative-inexactness-in-optimization-and-machine-learning/119373/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What noise setting does the paper analyze for gradient estimates?","Question",{"text":75,"@type":76},"It studies noise in gradient estimates that maintains an acute angle with the original gradient, using a relative inexactness model (relative error proportional to the gradient magnitude). It also contrasts this with an absolute error model.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the paper model stochastic optimization in machine learning?",{"text":80,"@type":76},"It treats stochastic optimization using sum-structured objectives over samples and defines minibatch gradients by sampling an index set. This yields an expected objective related to the true gradient.",{"name":82,"@type":73,"acceptedAnswer":83},"What do the results claim about classic gradient descent under the relative noise condition?",{"text":84,"@type":76},"The paper proves that classic gradient descent preserves the order of convergence up to multiplicative constants when the relative noise conditions are satisfied.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]