[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84950-en":3,"doc-seo-84950-105":29,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84950,7971461740886,"Theodore","https://ap-avatar.wpscdn.com/davatar_3d24733baf745e90a7e4bdd5f77d97b2",8,"Research & Report","Geometric Self Distillation for Reasoning Generalization","On-policy distillation provides large language models dense, token-level supervision by evaluating a teacher distribution on the student’s own trajectories. Privileged-context self-distillation supplies abundant hints or solution traces, yet supervision becomes harder to trust when the teacher’s privileged view diverges from what the student can justify. GEOSD treats accumulated drift as movement in predictive behavior, using a Hellinger-scaled pull and a proximal Fisher–Rao checkpoint penalty with natural-gradient updates. Results on reasoning benchmarks and three model families improve average out-of-distribution accuracy by 5.7–8.6 points.","GEOMETRIC SELF-DISTILLATION FOR REASONING GENERALIZATION  \narXiv :2607 .06855v 1 [ cs .LG] 7 Jul 2026  \nJosip Juki  \nILLC, University of Amsterdam [j.jukic@uva.nl](j.jukic@uva.nl)  \nIvan Titov  \nILLC, University of Amsterdam ILCC, University of Edinburgh [ititov@inf.ed.ac.uk](ititov@inf.ed.ac.uk)  \nABSTRACT  \nOn-policy distillation is a practical post-training recipe for large language models, supplying dense teacher supervision on the student’s own trajectories. In privilegedcontext self-distillation, teacher and student are the same model conditioned on the same prefix, but the teacher also sees a hint or the full solution trace. This makes supervision abundant but harder to trust: the teacher can be confident about continuations its privileged view makes obvious but the student cannot yet justify.  \nThe distillation pull is strongest where teacher and student disagree most, and over many updates it accumulates into drift that degrades out-of-distribution (OOD) reasoning. We introduce GEOSD, a geometric self-distillation objective that treats this drift as movement in the student’s predictive behavior and counters it in two complementary ways. A Hellinger loss scales each teacher preference by the overlap the student already shares with it, attenuating the pull on tokens the student cannot yet support. Since these pulls still compound over training, a proximal term penalizes how far the student’s predictions drift from a recent checkpoint, measured as a Fisher–Rao distance. Both are distances in the same geometry of next-token distributions, and a natural-gradient update takes its steps in that geometry rather than in parameter space. Across mathematical reasoning benchmarks and three model families, GEOSD preserves the in-distribution gains of self-distillation while improving average OOD accuracy by 5.7–8.6 points over the base model, with gains holding across model scales from 1.7B to 32B. Analyzing why standard matching fails out of distribution, we find it wins agreement with the teacher by draining mass from alternatives at high-entropy states, resulting in confident agreement on wrong answers, whereas GEO SD keeps those alternatives in reach.  \n1 INTRODUCTION  \nOn-policy distillation (OPD) has emerged as a practical post-training recipe for large language models (LLMs), providing dense teacher supervision on trajectories generated by the student itself (Agarwalet al., 2024; Song & Zheng, 2026) . Dense supervision is especially valuable for reasoning, where final-answer rewards are often too sparse to reveal which intermediate steps were useful (Lightmanet al., 2024) . Instead of assigning credit only after a completed solution, OPD evaluates the teacher’s distribution at the states the student visits, producing token-level targets throughout the trajectory. The recipe can scale cheaply, since it requires no separate, stronger teacher. In on-policy self-distillation (OPSD), the student acts as its own teacher, supervising the prefixes it generates. A common and effective variant grants this teacher privileged context, such as a hint or the full solution trace, while the student must reason from the original problem alone (Zhao et al., 2026; Ye et al., 2026) .  \nWhether dense supervision helps depends on the student’s ability to follow the teacher’s token-level preferences. Li et al. (2026) show that OPD progresses when student and teacher increasingly share mass on high-probability tokens, and stalls when that mass is largely disjoint. Privileged context is a systematic source of such mismatch: the extra information lets the teacher commit to continuations that are obvious given the solution but still uncertain from the student’s own context (H¨ubotter et al., 2026; Penaloza et al., 2026) . These mismatched states are also where imitation pulls hardest: standard divergences weight a teacher token by its own mass, so a token the teacher is sure of but the student barely supports produces a large gradient and a strong up","cbCaiqcAHqfanXYY","https://ap.wps.com/l/cbCaiqcAHqfanXYY","pdf",949638,1,24,"English","en",105,"# Abstract\n# Introduction\n## On-policy and privileged-context self-distillation\n## Mismatch, drift, and limitations of standard matching\n## GEOSD: geometry-aware distillation","[{\"question\":\"What performance improvements does GEOSD achieve over the base model?\",\"answer\":\"Across mathematical reasoning benchmarks and three model families, GEOSD preserves in-distribution gains of self-distillation while improving average out-of-distribution accuracy by 5.7–8.6 points compared with the base model. Gains hold across model scales from 1.7B to 32B.\"}]",1784199652,60,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":27},"geometric-self-distillation-for-reasoning-generalization","",{"@graph":35,"@context":77},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/geometric-self-distillation-for-reasoning-generalization/84950/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"What performance improvements does GEOSD achieve over the base model?","Question",{"text":75,"@type":76},"Across mathematical reasoning benchmarks and three model families, GEOSD preserves in-distribution gains of self-distillation while improving average out-of-distribution accuracy by 5.7–8.6 points compared with the base model. Gains hold across model scales from 1.7B to 32B.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":84},[85,89,93,97,101,106,111,114,119,122,126],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":45,"category_name":99,"show_sort_weight":28,"slug":100},5,"Comic","comic",{"id":102,"doc_module":4,"doc_module_name":45,"category_name":103,"show_sort_weight":104,"slug":105},6,"Technology",50,"technology",{"id":107,"doc_module":4,"doc_module_name":45,"category_name":108,"show_sort_weight":109,"slug":110},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":112,"slug":113},30,"research-report",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},9,"Religion & Spirituality",20,"religion-spirituality",{"id":117,"doc_module":4,"doc_module_name":45,"category_name":120,"show_sort_weight":117,"slug":121},"World Cup","world-cup",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":123,"slug":125},10,"Lifestyle","lifestyle",{"id":127,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":98,"slug":129},19,"General","general"]