[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81984-en":3,"doc-seo-81984-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},81984,1099514068035,"Ezra","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","The Anatomy of Implicit Bias: Information Allocation in Neural Network Training","Implicit bias is commonly interpreted as an optimizer’s preference for certain final solutions and their geometry, yet such a view leaves unclear how bias is formed during training. The paper introduces a training-time information allocation perspective, where optimization writes a structured pattern of error signals across parameter paths, coordinate channels, and sample regions. It proposes allocation diagnostics and a collapse–persistence analysis to separate training progress from internal allocation. Experiments on image and facial tasks validate multiple sources and allocation-aware optimization benefits.","arXiv :2607 .07 156v2 [ cs .LG] 10 Jul 2026  \nThe Anatomy of Implicit Bias: Information Allocation in Neural Network Training  \nZhang Gongyue 1 , Wang Zhiyong 1 , Liu Donghan 1 , Ren Weihong 1 , Sheng Yixuan 1 , Liu Honghai 1*  \n1 State Key Laboratory of Robotics and Systems, Harbin Institute of  \nTechnology Shenzhen, Shenzhen, 518055, China.  \n*Corresponding author(s). E-mail(s): [honghai.liu@icloud.com](honghai.liu@icloud.com) ;  \nAbstract  \nImplicit bias is usually explained as the preference of an optimization process for certain final solutions and their geometry. This view helps explain where a model finally stops. It gives less direct explanation of how this bias is formed during training. This paper proposes a training-time information allocation view. Under this view, optimization forms a writing pattern for error signals across parameter paths, coordinate channels, and sample regions.  \nThis paper builds a set of observable allocation diagnostics. These diagnostics include gradient demand, actual update injection, coordinate gain induced by exponential moving averages, channel-level update ratios, and sample-wise loss distributions. To separate training progress from internal allocation, this paper introduces a collapse–persistence analysis. Under matched training loss, if external loss statistics collapse but internal allocation ratios remain separated, then the factor changes the internal allocation of the training signal. Controlled experiments show that the effect of the learning rate largely collapses after progress matching. Thus, the learning rate is closer to a progress-dominant source. In contrast, the preconditioning exponent p keeps clear channel-gain and update-ratio signatures after progress matching. This indicates that p directly changes coordinate allocation induced by EMA preconditioning. Batch size, optimizer memory, data statistics, model width, activation function, and scheduling also show different allocation signatures. These results suggest that training-time implicit bias has multiple sources.  \nThis paper further verifies these observations on standard neural network tasks. On CIFAR-100 with ResNet-18, changing p systematically affects the weight–bias gradient structure, sample loss quantiles, and the identity of hard samples. It also shows a median–tail trade-off. A larger p reduces the loss of most samples more strongly, but it can increase the high-loss tail. On facial expression recognition  \n1  \ntasks, dynamic information-allocation schedules improve convergence behavior and final performance. This shows that allocation-aware optimization can be useful in real tasks.  \nOverall, this paper extends the analysis of implicit bias from final-solution geometry to training-time signal allocation. The main claim is that implicit bias is not only reflected by the final solution. It is also reflected by which parameter paths, coordinate channels, and sample regions receive the error signal first and more strongly during training. Based on this view, this paper places different training factors into a unified information-allocation diagnostic framework. The framework gives a mechanism-level explanation of training-time implicit bias. It also provides a basis for future optimization methods that control training progress and signal allocation separately.  \n1 Introduction  \nBefore discussing neural network optimization, we start with a simple analogy from flight. A gliding seagull and an artificial aircraft can both adjust their posture to correct a flight path and approach a target. However, they rely on different degrees of freedom. A seagull has rich and continuous body control. It has lower requirements on external initial conditions such as altitude. An aircraft depends more strongly on external conditions when it optimizes its flight path. Thus, even for the same task of flying toward a target, different systems allocate path errors to different correction paths.  \nNeural networks have a simila","cbCain9lrt9YUax8","https://ap.wps.com/l/cbCain9lrt9YUax8","pdf",1663473,1,38,"English","en",105,"# Abstract\n# Introduction\n## Flight analogy and allocation mechanisms\n## Training factors and error-signal writing","[{\"question\":\"What new perspective does the paper propose for implicit bias?\",\"answer\":\"It reframes implicit bias as a training-time information allocation phenomenon, focusing on how optimization allocates error signals across parameter paths, coordinate channels, and sample regions rather than only final-solution geometry.\"},{\"question\":\"How does the paper separate training progress from internal allocation effects?\",\"answer\":\"It introduces a collapse–persistence analysis that checks whether external loss statistics collapse while internal allocation ratios remain separated under matched training loss.\"},{\"question\":\"Which training factors are shown to produce distinct allocation signatures?\",\"answer\":\"Besides learning rate and Adam-like preconditioning exponent p, the paper reports differing signatures from batch size, optimizer memory, data statistics, model width, activation function, and scheduling.\"}]",1784177418,96,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"the-anatomy-of-implicit-bias-information-allocation-in-neural-network-training","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/the-anatomy-of-implicit-bias-information-allocation-in-neural-network-training/81984/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-21","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What new perspective does the paper propose for implicit bias?","Question",{"text":75,"@type":76},"It reframes implicit bias as a training-time information allocation phenomenon, focusing on how optimization allocates error signals across parameter paths, coordinate channels, and sample regions rather than only final-solution geometry.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the paper separate training progress from internal allocation effects?",{"text":80,"@type":76},"It introduces a collapse–persistence analysis that checks whether external loss statistics collapse while internal allocation ratios remain separated under matched training loss.",{"name":82,"@type":73,"acceptedAnswer":83},"Which training factors are shown to produce distinct allocation signatures?",{"text":84,"@type":76},"Besides learning rate and Adam-like preconditioning exponent p, the paper reports differing signatures from batch size, optimizer memory, data statistics, model width, activation function, and scheduling.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]