[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83358-en":3,"doc-seo-83358-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83358,687197207919,"Theodora","https://ap-avatar.wpscdn.com/avatar/a000253d6f5f7c60be?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779446848396160552",8,"Research & Report","Different Teachers, Different Capabilities Sub-1B On-Device Distillation for Structured Text Enrichment","High-volume structured extraction imposes per-item latency costs on large language models, motivating distillation into small on-device models. The study evaluates distillation delivered by teacher quality per sub-task using news enrichment: each article maps to one JSON object with a short summary and five closed-vocabulary labels. An 8B reasoning teacher distills into a 0.6B student, running about 0.8s per article versus 39s, recovering 58% of the teacher gap and improving summary quality significantly while preserving faithfulness.","Different Teachers, Different Capabilities: Sub-1BOn-Device Distillation for Structured Text  \nEnrichment  \nVinay Kumar Chaganti  \narXiv :2607 .08268v 1 [ cs .AI] 9 Jul 2026  \nAbstract—High-volume structured extraction pays a large model’s latency on every item, so distilling the task into a small on-device model is attractive: comparable output at a fraction of the time and cost. We measure what that distillation actually delivers, per sub-task. Each news article is mapped to one JSON object with a short summary and five categorical labels. We distill an 8B reasoning teacher (deepseek-r1:8b) into a 0.6B student (Qwen3-0.6B; QLoRA, three seeds), and add two teacher controls: a same-size non-reasoning teacher and a larger managed pipeline. A blinded, reference-free, three-judge panel scores every arm against the full article, alongside two non-distillation baselines, few-shot prompting and constrained decoding. The student runs at about 0.8 s per article against the teacher’s 39 s, and recovers 58% of the base-to-teacher gap on summary quality, beating its primary baseline (constrained decoding) by +16.8 points and fewshot prompting by a secondary +4.9. A same-size non-reasoning teacher trains a student no better than the untuned base, so the summary gain follows from the teacher’s reasoning nature rather than its scale. Capabilities then split by teacher: the reasoning teacher transfers writing quality and the managed pipeline transfers label diversity, while a same-size instruction teacher’s students stay more grounded on the 22 short, thin-source articles in the 93-item test set (74 versus 55 faithful), where the reasoning-lineage student fabricates. That grounding difference is a consistent ordering rather than a significant aggregate effect, and the subgroup is small, so we report it as a direction. Because no single engine wins every field, the deliverable is a per-field routing map for on-device enrichment.  \nIndex Terms—Edge inference, knowledge distillation, large language models, LLM-as-a-judge evaluation, model compression, reasoning-model distillation, small language models, structured output generation, text summarization.  \nI. INTRODUCTION  \nHIGH-volume structured extraction is a common produc  \ntion pattern: one prompt, one schema-bound JSON object per item, repeated over thousands of items. It appears in ticket triage, log classification, document intake, moderation pre-filters, and catalog extraction. Run through a mid-sized or large model, it pays that model’s latency on every item. Distilling the task into a small on-device model promises comparable output far faster and at lower cost. This paper measures what that distillation delivers, per sub-task, against the cheaper alternatives a practitioner tries first.  \nThe instance we study is news enrichment. Each article is mapped to one JSON object with a short summary and five closed-vocabulary labels: sentiment, urgency, frame, tone,  \nThe author is an independent researcher (e-mail: [cvk.atreya@gmail.com](cvk.atreya@gmail.com)). Evaluation-harness code, figure-generation scripts, and portions of the manuscript were produced with AI-agent assistance; see the Acknowledgment.  \nand depth. An 8B reasoning teacher (deepseek-r1:8b) produces acceptable output but takes about 39 s per article, so a 500-item batch runs 5.4 hours on a consumer laptop. A 0.6B model is roughly 40 times faster and fits on-device; the open question is quality. Distillation raises a small model’s output toward the teacher’s, and we ask how much of that quality it recovers on each sub-task, and whether it preserves faithfulness, the one property a user-facing pipeline cannot compromise. Wereport per sub-task because the sub-tasks carry different failure costs: a single aggregate hides the axis where a regression is unacceptable.  \nEvery model under study runs locally. Fixed weights at temperature 0 give deterministic, unlimited reruns, so the measurement holds still and every reported number","cbCaicMbsggHeSIk","https://ap.wps.com/l/cbCaicMbsggHeSIk","pdf",799570,3,1,12,"English","en",105,"# Introduction\n## Problem: High-volume structured extraction\n## Study setup and evaluation design\n# Approach: Distillation and baselines\n## Teacher controls and student configuration\n## Reference-free judging and significance tests\n# Findings\n## Summary quality recovery vs baselines\n## Faithfulness and thin-source grounding effects\n## Per-field routing map for on-device enrichment","[{\"question\":\"What task and output format are used for news enrichment in this study?\",\"answer\":\"Each news article is mapped to one JSON object containing a short summary and five closed-vocabulary categorical labels (sentiment, urgency, frame, tone, and depth).\"},{\"question\":\"How is the distillation evaluated, and what is the scoring setup?\",\"answer\":\"A blinded, reference-free three-judge panel scores each experimental arm against the full articles. The evaluation includes non-distillation baselines and statistical significance testing via paired bootstrap methods.\"},{\"question\":\"How much faster is the distilled student model compared with the reasoning teacher, and what is the quality recovery result?\",\"answer\":\"The student runs at about 0.8 seconds per article versus about 39 seconds for the teacher, recovering about 58% of the base-to-teacher summary quality gap and beating the primary constrained-decoding baseline by +16.8 points.\"}]",1784186982,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"different-teachers-different-capabilities-sub-1b-on-device-distillation-for-structured-text-enrichment","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/different-teachers-different-capabilities-sub-1b-on-device-distillation-for-structured-text-enrichment/83358/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What task and output format are used for news enrichment in this study?","Question",{"text":75,"@type":76},"Each news article is mapped to one JSON object containing a short summary and five closed-vocabulary categorical labels (sentiment, urgency, frame, tone, and depth).","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is the distillation evaluated, and what is the scoring setup?",{"text":80,"@type":76},"A blinded, reference-free three-judge panel scores each experimental arm against the full articles. The evaluation includes non-distillation baselines and statistical significance testing via paired bootstrap methods.",{"name":82,"@type":73,"acceptedAnswer":83},"How much faster is the distilled student model compared with the reasoning teacher, and what is the quality recovery result?",{"text":84,"@type":76},"The student runs at about 0.8 seconds per article versus about 39 seconds for the teacher, recovering about 58% of the base-to-teacher summary quality gap and beating the primary constrained-decoding baseline by +16.8 points.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]