[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86507-en":3,"doc-seo-86507-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86507,1099514067415,"Rowan","https://ap-avatar.wpscdn.com/avatar/100002539d78ffe74a7?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779092875211072502",8,"Research & Report","Answer-Conditioned Chain-of-Thought Distillation for Few-Shot Industrial Vision with Small VLMs","Deploying AI-based visual inspection in manufacturing is difficult because requirements evolve, new defect types emerge, and large labeled datasets are rarely available. The proposed answer-conditioned chain-of-thought (CoT) distillation rapidly adapts small vision-language models (VLMs) to new industrial tasks using minimal labeled data. A frontier VLM generates label-justified visual explanations, and a 3B LoRA fine-tuning step trains on reasoning-augmented examples, improving performance even when teacher accuracy is low.","arXiv :2607 . 10666v1 [ cs .CV] 12 Jul 2026  \nAnswer-Conditioned Chain-of-Thought Distillation for Few-Shot Industrial Vision with Small VLMs  \nShubham Rao  \nEntropy AI Research Labs Private Limited  \n[director@entropyresearch.ai](director@entropyresearch.ai)  \n[https://entropyresearch.ai](https://entropyresearch.ai)  \nAbstract  \nDeploying AI-based visual inspection in manufacturing is hard because requirements change often, new defect types appear, and large labeled datasets are rarely available. We propose answer-conditioned chain-of-thought (CoT) distillation for rapidly adapting small vision-language models (VLMs) to new industrial tasks using minimal labeled data. A frontier VLM receives each training image along with its correct label and generates a justified visual explanation. A 3B-parameter model is then fine-tuned on these reasoning-augmented examples via LoRA. By conditioning on correct answers, we ensure all training reasoning is directed toward the correct conclusion, which is critical because frontier models score as low as 24.1% on our hardest task. We validate on four industrial classification tasks spanning three image modalities using only 18 to 30 labeled images per task. Across 4 seeds per task (32 training runs), our method outperforms direct fine-tuning on all 16 seed-task combinations, with mean improvements of +1.7 to +4.4 percentage points. A controlled equal-budget experiment confirms the improvement comes from reasoning quality, not additional training steps. An unconditioned baseline demonstrates that without answer-conditioning, wrong reasoning degrades performance by 17.8 percentage points. On weld radiograph classification, the fine-tuned 3B model outperforms GPT-4.1 by 10.0pp using just 24 training images.  \n1 Introduction  \nVisual quality inspection is central to manufacturing. Every production line needs to identify defects, classify materials, or verify compliance with standards. These requirements change often. Each change demands anew or retrained vision model.  \nTraditional approaches rely on CNNs trained on thousands of labeled images [2] . Collecting this data is expensive. In many settings, a factory has fewer than 30 labeled examples and needs a working classifier within days, not months.  \nLarge VLMs like GPT-4.1 can interpret images without task-specific training. But their accuracy on specialized industrial tasks is poor. On concrete aggregate grading (9 classes, DIN 1045 [4]), GPT-4.1 achieves 29.6% few-shot. Gemini 2.5 Pro achieves 24.1% . These models cannot be deployed on-premises due to cost and data privacy constraints.  \nSmall VLMs (1–7B parameters) [14] fit on edge hardware and can be fine-tuned with LoRA [6] . But standard fine-tuning on 18–30 images maps images to labels without transferring domain knowledge. The model memorizes rather than learns.  \nCoT distillation [5] trains smaller models on reasoning-augmented data from larger models, but existing work targets text-only settings. STaR [17] and Video-STaR [19] use answer-conditioned reasoning in text and video settings, but not for VLMs where the teacher itself performs poorly.  \nWe propose answer-conditioned CoT distillation for VLMs. A frontier model receives each training image with its correct label and explains why that label is correct. A small VLM is fine-tuned on these reasoning-augmented pairs via LoRA. The frontier model is not classifying. It already knows the answer. It is explaining what visual features justify that answer.  \nThis is critical because frontier model accuracy ranges from 24.1% to 91.1% across our tasks. If we asked it to classify and explain, up to 76% of training data would contain incorrect reasoning. We show experimentally that unconditioned CoT destroys performance (−17.8pp) when the teacher is mostly wrong.  \nOur contributions:  \n1. We apply answer-conditioned CoT distillation to VLMs for industrial few-shot classification, showing consistent improvement across 4 tasks, 3 modalities, a","cbCail6Ye7CpRSJq","https://ap.wps.com/l/cbCail6Ye7CpRSJq","pdf",8684558,4,1,11,"English","en",105,"# Introduction\n## Related Work\n# Method\n## Prompt Design","[{\"question\":\"为什么制造业视觉检测在实际落地中难以完成？\",\"answer\":\"需求经常变化、新缺陷类型会出现，而大规模标注数据通常难以获得。\"},{\"question\":\"答复条件（answer-conditioned）的 CoT 蒸馏是如何工作的？\",\"answer\":\"前沿 VLM 在看到训练图像时输出正确标签及其理由解释；随后用带有推理增强样本的数据对小型 VLM（例如 3B）进行 LoRA 微调，并通过“正确答案条件”把推理对齐到正确结论。\"},{\"question\":\"实验结果表明该方法带来的收益来自哪里？\",\"answer\":\"受控的等预算实验表明提升来自推理质量，而不是额外训练步骤；未进行 answer-conditioning 的基线在教师推理错误时会显著降性能。\"}]",1784212275,28,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"answer-conditioned-chain-of-thought-distillation-for-few-shot-industrial-vision-with-small-vlms","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/answer-conditioned-chain-of-thought-distillation-for-few-shot-industrial-vision-with-small-vlms/86507/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"为什么制造业视觉检测在实际落地中难以完成？","Question",{"text":75,"@type":76},"需求经常变化、新缺陷类型会出现，而大规模标注数据通常难以获得。","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"答复条件（answer-conditioned）的 CoT 蒸馏是如何工作的？",{"text":80,"@type":76},"前沿 VLM 在看到训练图像时输出正确标签及其理由解释；随后用带有推理增强样本的数据对小型 VLM（例如 3B）进行 LoRA 微调，并通过“正确答案条件”把推理对齐到正确结论。",{"name":82,"@type":73,"acceptedAnswer":83},"实验结果表明该方法带来的收益来自哪里？",{"text":84,"@type":76},"受控的等预算实验表明提升来自推理质量，而不是额外训练步骤；未进行 answer-conditioning 的基线在教师推理错误时会显著降性能。","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]