[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86261-en":3,"doc-seo-86261-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":11,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},86261,687197207919,"Theodora","https://ap-avatar.wpscdn.com/avatar/a000253d6f5f7c60be?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779446848396160552",8,"Research & Report","Single-Teacher View Augmentation: Enhancing Knowledge Distillation with Student-Guided Perturbations","Knowledge distillation (KD) often uses a single, fixed-teacher perspective, which limits diversity in supervisory signals and may reduce generalization. Multi-teacher KD improves diversity but is expensive in computation and storage. Virtual-view generation from one teacher trades off efficiency versus controlled diversity. Shift-Augmented Knowledge Distillation (SAKD) uses evolving student features as dynamic conditions to generate parameter-free cyclic shifts, enabling single-stage training. Experiments on CIFAR-100 and ImageNet show SAKD surpasses random perturbations and matches two-stage accuracy with fewer parameters and no pre-training.","Single-Teacher View Augmentation: Enhancing Knowledge Distillation with Student-Guided  \nPerturbations  \nXuyi Yu, Yaohua Liu, Ziming Song, Yinghai Zhao, Huipeng Zhang, and Kuizhi Mei, Member, IEEE  \narXiv :2607 . 11557v1 [ cs .CV] 13 Jul 2026  \nAbstract—Knowledge distillation (KD) typically relies on the fixed perspective of a single teacher, limiting the diversity of supervisory signals. While multi-teacher distillation addresses this by aggregating knowledge from multiple models, it incurs prohibitive computational and storage costs. To balance efficiency and diversity, recent research has focused on generating virtual views from a single teacher. However, existing methods face a trade-off: random perturbation approaches offer efficiency but lack controlled diversity, while structured augmentation methods require multi-stage training and incur linear parameter growth. We observe that this trade-off stems from a common design choice: using the teacher’s strong but static features to generate views. Instead, we propose Shift-Augmented Knowledge Distillation (SAKD), a simple yet effective framework that leverages the student’s evolving features as a dynamic condition for perturbation generation. This shift in perspective enables singlestage training while producing adaptive, diverse views through a parameter-free cyclic shift. Extensive experiments on CIFAR-100 and ImageNet demonstrate that SAKD consistently outperforms random perturbation methods and achieves accuracy on par with two-stage approaches, while using significantly fewer parameters and eliminating pre-training requirements.  \nIndex Terms—Knowledge distillation, Model compression, Teacher augmentation, Structured perturbation  \nI. INTRODUCTION  \nThe deployment of deep neural networks in edge devices, mobile platforms, and other resource-constrained environments is often hindered by their substantial computational and memory requirements. Knowledge distillation (KD) [1] has emerged as a leading solution, transferring knowledge from a cumbersome but accurate teacher model to a lightweight student. While original KD aligned softened output probabilities, subsequent research has expanded to richer supervisory forms including feature-based [2], [3], [4], [5], relation-based [6], [7], [8], and logit-based approaches [9], [10], [11], [12],[13] . Despite these advances, a critical limitation remains: the student receives supervision from only one fixed perspective of the teacher, limiting the diversity of guidance and potentially constraining generalization.  \nXuyi Yu, Huipeng Zhang, and Kuizhi Mei are with the State Key Laboratory of Human-Machine Hybrid Augmented Intelligence, Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University, Xi’an 710049, China (e-mail: [yuxuyi@stu.xjtu.edu.cn](yuxuyi@stu.xjtu.edu.cn); [zhp429858387@stu.xjtu.edu.cn](zhp429858387@stu.xjtu.edu.cn); [meikuizhi@mail.xjtu.edu.cn](meikuizhi@mail.xjtu.edu.cn)).  \nYaohua Liu is with the Guangdong Institute of Intelligence Science and Technology, Zhuhai 519031, China (e-mail: [liuyaohua@gdiist.cn](liuyaohua@gdiist.cn)).  \nZiming Song is with the Institute of Collaborative Innovation, University of Macau, Taipa, Macau SAR, China (e-mail: zimingsong [um@163.com](um@163.com)).  \nYinghai Zhao is with the Beijing Huahang Institute of Radio Measurement, Beijing 102445, China (e-mail: [513043512@qq.com](513043512@qq.com)).  \nMulti-teacher distillation aggregates knowledge from multiple independently trained models to mitigate this issue [14],[15], [16], [17] . However, the cost of training and storing multiple large teachers is often prohibitive. A more practical direction generates multiple virtual teacher perspectives from a single pre-trained model. Existing approaches fall into two categories. Methods like TeKAP [18](Fig. 1(a)) inject random noise into teacher features or logits during training, creating diversity through controlled stochasticity in a single-stage, parameter-efficient man","cbCaicUIHoUkLbS7","https://ap.wps.com/l/cbCaicUIHoUkLbS7","pdf",914774,4,1,"English","en",105,"# Introduction\n## Knowledge distillation and its limitations\n## Existing single-teacher virtual-view methods","[{\"question\":\"What problem does single-teacher knowledge distillation face?\",\"answer\":\"The student is supervised using only one fixed teacher perspective, which limits diversity of guidance and can constrain generalization.\"},{\"question\":\"How does SAKD differ from random perturbation and structured augmentation methods?\",\"answer\":\"SAKD generates diversity using dynamic, evolving student features and a parameter-free cyclic shift, avoiding the randomness of random perturbations and the multi-stage parameter growth of structured reconstruction approaches.\"},{\"question\":\"What do the experiments on CIFAR-100 and ImageNet show for SAKD?\",\"answer\":\"SAKD consistently outperforms random perturbation methods and reaches accuracy comparable to two-stage approaches while using significantly fewer parameters and eliminating pre-training requirements.\"}]",1784209890,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"single-teacher-view-augmentation-enhancing-knowledge-distillation-with-student-guided-perturbations","",{"@graph":35,"@context":84},[36,52,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":20},"https://docshare.wps.com/document/single-teacher-view-augmentation-enhancing-knowledge-distillation-with-student-guided-perturbations/86261/",{"url":51,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":23,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":40,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"What problem does single-teacher knowledge distillation face?","Question",{"text":74,"@type":75},"The student is supervised using only one fixed teacher perspective, which limits diversity of guidance and can constrain generalization.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How does SAKD differ from random perturbation and structured augmentation methods?",{"text":79,"@type":75},"SAKD generates diversity using dynamic, evolving student features and a parameter-free cyclic shift, avoiding the randomness of random perturbations and the multi-stage parameter growth of structured reconstruction approaches.",{"name":81,"@type":72,"acceptedAnswer":82},"What do the experiments on CIFAR-100 and ImageNet show for SAKD?",{"text":83,"@type":75},"SAKD consistently outperforms random perturbation methods and reaches accuracy comparable to two-stage approaches while using significantly fewer parameters and eliminating pre-training requirements.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":28,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":28,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]