[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-122325-en":3,"doc-seo-122325-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},122325,687197207057,"Sage","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Knowledge Distillation with Auxiliary Variable","Knowledge distillation (KD) transfers predictive knowledge from a teacher model to a student model by aligning predictive distributions, but common KD formulations mirror the teacher’s distribution modeling strategy and become sub-optimal when teacher and student capacities differ. This work introduces an auxiliary variable related to target variables to strengthen the student’s ability to model predictive distributions. A reformulated KD objective is derived, supported by theory for improved performance, and validated by experiments showing consistent gains over existing KD methods.","Knowledge Distillation with Auxiliary Variable  \nBo Peng 1 Zhen Fang 1 Guangquan Zhang 1 Jie Lu 1  \nAbstract  \nKnowledge distillation (KD) provides an efficient framework for transferring knowledge from a teacher model to a student model by aligning their predictive distributions. The existing KD methods adopt the same strategy as the teacher to formulate the student’s predictive distribution. However, employing the same distribution-modeling strategy typically causes sub-optimal knowledge transfer due to the discrepancy in model capacity between teacher and student models. Designing student-friendly teachers contributes to alleviating the capacity discrepancy, while it requires either complicated or student-specific training schemes.  \nTo cast off this dilemma, we propose to introduce an auxiliary variable to promote the ability of the student to model predictive distribution. The auxiliary variable is defined to be related to target variables, which will boost the model prediction.  \nSpecifically, we reformulate the predictive distribution with the auxiliary variable, deriving a novel objective function of KD. Theoretically, we provide insights to explain why the proposed objective function can outperform the existing KD methods. Experimentally, we demonstrate that the proposed objective function can considerably and consistently outperform existing KD methods.  \n1. Introduction  \nOver the past decades, deep learning has shown its significance by boosting the performance of various real-world tasks (Hassaballah & Awad, 2020 ; Hupkes et al., 2023) . The effectiveness of deep learning generally comes at the expense of huge computational complexity and massive storage requirements. This restricts the deployment of largescale models (teachers) in real-time applications where  \n1Faculty of Engineering & Information Technology, University of Technology Sydney, Sydney, Australia. Correspondence to: Zhen Fang \u003C[Zhen.Fang@uts.edu.au](Zhen.Fang@uts.edu.au) >.  \nProceedings of the 41 st International Conference on Machine Learning, Vienna, Austria. PMLR 235, 2024 . Copyright 2024 by the author(s) .  \nlightweight models (students) are preferable due to limited resources (Li et al., 2023) . In this context, knowledge distillation (KD) (Gou et al., 2021 ; Wang & Yoon, 2021) is introduced to transfer knowledge from a teacher to a student model. Conventionally, KD is approached by minimizing the Kullback-Leibler (KL) divergence between predictive distributions of the teacher and student (Hinton et al., 2015) . To implement this vision, an intuitive yet commonly accepted approach, initially introduced in Hinton et al. (2015), is that the student follows the pre-trained teacher to formulate predictive posterior probabilities with logit outputs. Consequently, knowledge can be distilled from the teacher to the student by matching their logit outputs.  \nThis logit-matching approach, however, is challenged by the counter-intuitive observations (Cho & Hariharan, 2019 ; Stanton et al., 2021) . Specifically, a larger teacher does not necessarily increase a student’s accuracy compared to a relatively smaller teacher. This is attributed to the capacity gap between the two models (Huang et al., 2022a ; Mirzadehet al., 2020) since the discrepancy between their predictions can be significantly large. Thus, directly aligning their predictive distributions would lead to sub-optimal knowledge transfer and even disturb the training of the student.  \nAdvanced methods introduce a novel direction to go beyond the logit-matching approach. These methods develop student-friendly teachers to shrink the capacity gap. For instance, TAKD (Mirzadeh et al., 2020) introduces multiple middle-sized teaching assistant models to guide the student; DGKD (Son et al., 2021) improves TAKD by densely gathering all the assistant models; SFTN (Park et al., 2021) provides the teacher with a snapshot of the student during training. Despite remarkable progress, these methods need to r","cbCaijtnKOGjTpYL","https://ap.wps.com/l/cbCaijtnKOGjTpYL","pdf",424514,1,15,"English","en",105,"# Introduction\n## Background: logit-matching KD\n## Challenge: capacity gap and sub-optimal transfer\n## Prior work: student-friendly teachers\n## Proposed idea: auxiliary variable for predictive distribution modeling","[{\"question\":\"What problem does the paper address in knowledge distillation?\",\"answer\":\"Existing KD methods align predictive distributions using the same modeling strategy as the teacher, which becomes sub-optimal when the student has weaker capacity and produces a large discrepancy. This can harm knowledge transfer and disturb student training.\"},{\"question\":\"How does the auxiliary variable improve KD in this work?\",\"answer\":\"The auxiliary variable is defined to be related to target variables and is used to reformulate predictive distributions for the student. This brings external knowledge that promotes the student’s ability to model predictive distributions.\"},{\"question\":\"What results are reported for the proposed objective function?\",\"answer\":\"The paper provides theoretical insights explaining why the new objective can outperform prior KD methods, and experimental results show it can considerably and consistently outperform existing approaches.\"}]","Knowledge Distillation with Auxiliary Variable | PDF",1785810014,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"knowledge-distillation-with-auxiliary-variable","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/knowledge-distillation-with-auxiliary-variable/122325/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address in knowledge distillation?","Question",{"text":75,"@type":76},"Existing KD methods align predictive distributions using the same modeling strategy as the teacher, which becomes sub-optimal when the student has weaker capacity and produces a large discrepancy. This can harm knowledge transfer and disturb student training.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the auxiliary variable improve KD in this work?",{"text":80,"@type":76},"The auxiliary variable is defined to be related to target variables and is used to reformulate predictive distributions for the student. This brings external knowledge that promotes the student’s ability to model predictive distributions.",{"name":82,"@type":73,"acceptedAnswer":83},"What results are reported for the proposed objective function?",{"text":84,"@type":76},"The paper provides theoretical insights explaining why the new objective can outperform prior KD methods, and experimental results show it can considerably and consistently outperform existing approaches.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]