[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85603-en":3,"doc-seo-85603-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85603,3848291630094,"Emma Wilson","https://eur-avatar.wpscdn.com/davatar_085a072bc5b1113ac321206ff7593b45",8,"Research & Report","Student Capacity Moderates Knowledge Distillation Effectiveness","Student capacity relationships modulate knowledge distillation effectiveness in ResNet-based image classification on CIFAR-10. A systematic study compares Logit-KD and Feature-KD across four teacher-student pairs (R50→R18, R34→R18, R50→R34, R101→R34) using a leakage-free evaluation protocol with validation-based hyperparameter selection, five-seed re-runs, and test-only final reporting. Distillation fidelity is assessed via teacher-student agreement and KL divergence. Results show Feature-KD consistently matches or exceeds Logit-KD, with capacity effects concentrated on R34 students, while architectural correctness dominates gains.","arXiv :2605 .3 1 19 1v2 [ cs .LG] 12 Jul 2026  \nStudent Capacity Moderates Knowledge Distillation  \nEffectiveness:  \nA Systematic Study Across ResNet Teacher-Student Pairs on  \nCIFAR-10  \nUmut Onur Yaşar  \n[umutonuryasar@gmail.com](umutonuryasar@gmail.com)  \nAbstract  \nWe investigate how teacher-student capacity relationships modulate knowledge distillation (KD) effectiveness in ResNet-based image classification on CIFAR-10 . Across four teacherstudent pairs—R50→R18, R34→R18, R50→R34, and R101→R34—we compare Logit-KD and Feature-KD under a strict evaluation protocol: hyperparameters and checkpoints are selected on a held-out validation split, selected configurations are re-run with five seeds, and the test set is used exclusively for final reporting. Beyond accuracy, we measure distillation fidelity directly via teacher-student agreement and KL divergence. We report four findings.  \nFirst, the student-capacity pattern survives the corrected protocol at reduced magnitude: the only statistically significant gains occur for R34 students under Feature-KD (+0 .19 and +0 .21 pp, p \u003C 0.05), in two pairs whose teachers differ two-fold in parameters but not inaccuracy — localizing the moderating variable on the student side — while no KD gain for R18 students is distinguishable from zero. Second, Feature-KD matches or outperforms Logit-KD in all four pairs, and its students land closer to the teacher’s output distribution (DKL at T = 1) than Logit-KD students despite never observing teacher logits. Third, top-1 teacher-student agreement is flat across all pairs, decoupling fidelity from accuracy gains.  \nFourth, architecture dominates KD: correcting the ResNet stem for 32×32 inputs is worth +5 .5 to +7 .2pp—more than 25 × the largest KD gain. We also retract an attribution made in v1 of this paper: a controlled re-run shows the reported gradient-clipping bug had no measurable effect, and v1’s larger gains are explained by test-set selection. All code and results are available at [github.com/umutonuryasar/kd-capacity-gap](github.com/umutonuryasar/kd-capacity-gap).  \n1 Introduction  \nDeep neural networks achieve state-of-the-art performance across a wide range of computer vision tasks, but their computational demands make deployment on resource-constrained hardware challenging. Knowledge Distillation (KD) offers a principled compression strategy: instead of training a small student model on hard one-hot labels alone, the student also learns from the soft output distribution of a larger, pre-trained teacher model. These soft targets encode inter-class similarity — the relative probabilities assigned to non-target classes — a form of structured information Hinton et al. [5] termed “dark knowledge.”  \nDespite its widespread use, the relative effectiveness of different distillation paradigms remains context-dependent and poorly understood. Two dominant approaches exist: Logit-KD, which aligns the student’s output distribution with the teacher’s via KL divergence, and Feature-KD, which aligns intermediate representations at one or more layers. Both approaches introduce hyperparameters—the distillation weight α, the softmax temperature T for Logit-KD, and layerselection and alignment metrics for Feature-KD—whose optimal values depend on the task and architecture.  \nPrior work has shown that excessive teacher-student capacity mismatch can impede knowledge transfer [3 , 7], but systematic evidence across multiple teacher-student pairs and KD paradigms under a leakage-free evaluation protocol remains limited. Moreover, recent work argues that distillation gains should be interpreted jointly with fidelity —how closely the student actually matches its teacher—rather than accuracy alone [9] .  \nIn this work, we make the following contributions:  \n1. A systematic comparison of Logit-KD and Feature-KD on CIFAR-10 across four teacherstudent pairs under a strict two-stage protocol: hyperparameters (α ∈ {0 .3 , 0.5 , 0. 7} , T ∈ {2, 3 , 4}) selected on a","cbCaisoavZaIuYcU","https://ap.wps.com/l/cbCaisoavZaIuYcU","pdf",690204,4,1,10,"English","en",105,"# Introduction\n## Knowledge Distillation\n## Feature-Based Distillation","[{\"question\":\"How are Logit-KD and Feature-KD compared in the study?\",\"answer\":\"The study evaluates Logit-KD by aligning the student and teacher output distributions via KL divergence, and evaluates Feature-KD by aligning intermediate representations across layers, using the same disciplined evaluation protocol across teacher-student pairs.\"},{\"question\":\"What evaluation protocol is used to ensure results are leakage-free?\",\"answer\":\"Hyperparameters and checkpoints are selected on a held-out validation split, the selected configurations are re-run with five different seeds, and the test set is reserved exclusively for final reporting.\"},{\"question\":\"What determines the largest performance gains: distillation method or architecture setup?\",\"answer\":\"Architectural correctness dominates: correcting the ResNet stem for 32×32 inputs yields roughly +5.5 to +7.2 percentage points, far exceeding the largest observed KD gains.\"}]",1784204864,25,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"student-capacity-moderates-knowledge-distillation-effectiveness","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/student-capacity-moderates-knowledge-distillation-effectiveness/85603/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"How are Logit-KD and Feature-KD compared in the study?","Question",{"text":75,"@type":76},"The study evaluates Logit-KD by aligning the student and teacher output distributions via KL divergence, and evaluates Feature-KD by aligning intermediate representations across layers, using the same disciplined evaluation protocol across teacher-student pairs.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What evaluation protocol is used to ensure results are leakage-free?",{"text":80,"@type":76},"Hyperparameters and checkpoints are selected on a held-out validation split, the selected configurations are re-run with five different seeds, and the test set is reserved exclusively for final reporting.",{"name":82,"@type":73,"acceptedAnswer":83},"What determines the largest performance gains: distillation method or architecture setup?",{"text":84,"@type":76},"Architectural correctness dominates: correcting the ResNet stem for 32×32 inputs yields roughly +5.5 to +7.2 percentage points, far exceeding the largest observed KD gains.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":22,"slug":133},"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]