[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82280-en":3,"doc-seo-82280-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82280,13056703019662,"Evangeline","https://ap-avatar.wpscdn.com/avatar/be000253a8e92610077?_k=1778726343310543188",8,"Research & Report","LionVote Per-Layer Learning Rate Adaptation for Lion","Per-layer diagnostics reveal that, at the prescribed learning rate, Lion’s effective scale is 2.6–2.8× too high for attention and MLP parameters and ~2× too high for normalisation layers on ViT-Tiny/CIFAR-100; this 32% cross-layer-type disparity cannot be reproduced by a single global rate. Measurement is provided by LionVote, a stateful per-layer learning rate mechanism that maintains a compound level and uses two diagnostics plus a validation-loss tiebreaker, bounded by structural cadence selection.","arXiv :2607 .09266v 1 [ cs .LG] 10 Jul 2026  \nLionVote: Per-Layer Learning Rate Adaptation for Lion  \nKris Atallah  \nNew York University  \nNew York, NY, USA  \n[kris.a@nyu.edu](kris.a@nyu.edu)  \nAbstract  \nPer-layer diagnostics reveal that, at the prescribed learning rate, Lion’s effective scale is 2.6–2.8 × too high for attention and MLP parameters and ∼2 × too high for normalisation layers on ViTTiny/CIFAR-100; this 32% cross-layer-type disparity cannot be reproduced by a single global rate.  \nThe measurement comes from LionVote, a per-layer learning rate mechanism in which each parameter tensor maintains a compound level, a persistent integer updated every c epochs by two diagnostics (gradient direction stability and momentum health) resolved by a validation loss tiebreaker. Voting thresholds derive from geometric identities, the EMA time constant, and a noise-floor estimate;  \ncadence is bounded structurally and selected by ablation. On ViT-Tiny/CIFAR-100, LionVote achieves 69 .7% top-1 accuracy vs. Lion’s 69.0%(p \u003C 0.02, Welch’s t-test) and AdamW’s 68.8% .  \nPer-layer adaptation value depends on both architectural heterogeneity and task; on uniform CNN architectures tuned SGD with cosine annealing remains dominant, and on ViT architectures gains are task-dependent.  \n1 Introduction  \nZhao et al. [Zhao et al., 2025] show that applying per-layer adaptive preconditioning only to the last layer and LayerNorm parameters recovers most of Adam’s advantage over SGD on autoregressive language models; this is evidence that layer types have different optimisation characteristics. Their analysis does not prescribe how much each type should diverge from a base rate, nor propose a mechanism to determine this during training.  \nExisting adaptive methods operate per-coordinate or per-layer but statelessly, and schedule-free approaches [Defazio et al., 2024] maintain a single global rate. No existing method combines stateful layer-level adaptation with data-driven rate adjustment (§2.4) .  \nLionVote is a per-layer learning rate mechanism for Lion [Chen et al., 2023] . Lion’s sign operation discards gradient magnitude, preventing directionally unstable layers from compensating through larger updates; this makes per-layer miscalibration more consequential for sign-based methods than for second-moment methods like Adam (§5.1) . Each parameter tensor maintains a compound level (a persistent integer, §3.1) that modulates the base learning rate exponentially. Every c epochs, two per-layer diagnostics (gradient direction stability and momentum health, with derived thresholds; Appendix A.1–A.2) vote on whether to increase, decrease, or maintain each layer’s rate. When the votes conflict, a validation loss tiebreaker resolves the decision. The mechanism adds one extra accumulation per parameter per batch; voting runs once every c epochs. With all votes disabled, compound levels decay to zero and LionVote reduces to standard Lion.  \nWe evaluate LionVote on WideResNet and ViT-Tiny across CIFAR-10 and CIFAR-100, with 8 seeds per configuration. The contributions are:  \n1. A per-layer adaptive mechanism for a sign-based optimizer whose voting thresholds derive from geometric identities and the EMA time constant; cadence and structural parameters (the exponent divisor d=2, maximum level L=4) are bounded rather than uniquely determined (Appendix A.1–A.6) . The threshold derivation methodology is reusable for principled design of per-layer mechanisms in other optimizers.  \n2. A quantified finding about Lion: its effective scale is 2.6–2.8 × too high for attention and MLP parameters and ∼2 × too high for normalisation parameters on ViT-Tiny/CIFAR-100, measured via compound level trajectories across 8 seeds. Normalisation layers receive 32% higher effective scale than attention layers (§5 . 1) . Because the compound multiplier scales both the sign update and decoupled weight decay, this measures joint LR+WD miscalibration (§5 . 1) .  \n3. Evidence that per-la","cbCaidBlCMEU5hgH","https://ap.wps.com/l/cbCaidBlCMEU5hgH","pdf",4165216,1,34,"English","en",105,"# Introduction\n# Related Work\n## Per-Parameter Adaptation\n## Per-Layer Adaptation","[{\"question\":\"What problem does LionVote address in Lion training?\",\"answer\":\"LionVote addresses cross-layer miscalibration in Lion: attention/MLP layers and normalisation layers receive effective scales that are systematically too high under the prescribed global learning rate.\"},{\"question\":\"How does LionVote decide whether to increase, decrease, or keep each layer’s learning rate?\",\"answer\":\"Each parameter tensor maintains a compound level, and every c epochs two diagnostics—gradient direction stability and momentum health—produce votes; a validation-loss tiebreaker resolves conflicts.\"},{\"question\":\"What evidence shows LionVote improves accuracy on ViT-Tiny/CIFAR-100?\",\"answer\":\"On ViT-Tiny/CIFAR-100, LionVote reaches 69.7% top-1 accuracy versus Lion’s 69.0% (p \\u003c 0.02, Welch’s t-test) and AdamW’s 68.8%.\"}]",1784179363,86,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"lionvote-per-layer-learning-rate-adaptation-for-lion","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/lionvote-per-layer-learning-rate-adaptation-for-lion/82280/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-21","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does LionVote address in Lion training?","Question",{"text":75,"@type":76},"LionVote addresses cross-layer miscalibration in Lion: attention/MLP layers and normalisation layers receive effective scales that are systematically too high under the prescribed global learning rate.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does LionVote decide whether to increase, decrease, or keep each layer’s learning rate?",{"text":80,"@type":76},"Each parameter tensor maintains a compound level, and every c epochs two diagnostics—gradient direction stability and momentum health—produce votes; a validation-loss tiebreaker resolves conflicts.",{"name":82,"@type":73,"acceptedAnswer":83},"What evidence shows LionVote improves accuracy on ViT-Tiny/CIFAR-100?",{"text":84,"@type":76},"On ViT-Tiny/CIFAR-100, LionVote reaches 69.7% top-1 accuracy versus Lion’s 69.0% (p \u003C 0.02, Welch’s t-test) and AdamW’s 68.8%.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]