[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81820-en":3,"doc-seo-81820-105":31,"detail-sidebar-cat-0-en-105":89},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},81820,4398048950312,"Violet","https://ap-avatar.wpscdn.com/avatar/400002538284de19e3c?_k=1778320343897328908",8,"Research & Report","Revisiting Chain-of-Thought Reasoning under Limited Supervision: Semi-supervised Chain-of-Thought Learning","Chain-of-thought (CoT) reasoning activates latent reasoning abilities in large language models, yet existing CoT techniques largely discard generated reasoning traces after inference. Semi-supervised Chain-of-Thought Learning reuses these traces by turning unlabeled questions into pseudo supervision. Semi-CoT generates multiple pseudo-CoTs per unlabeled item, measures answer-level semantic entropy, and selects low-entropy chains as reliable pseudo-CoT demonstrations. Experiments on AQuA, SVAMP, GSM8K, and MultiArith show entropy gating yields 91.36%–100% pseudo-answer precision, with mixed transfer and limits.","arXiv :2607 .0 15 1 1v 1 [ cs .AI] 1 Jul 2026  \nRevisiting Chain-of-Thought Reasoning under Limited Supervision: Semi-supervised Chain-of-Thought  \nLearning  \nHongyang He 1 ,3†, Jiuming Liu2 , Victor Sanchez 1 1 University of Warwick, 2 University of Cambridge, 3 Manifolda.Ai  \n†Corresponding author  \nAbstract  \nChain-of-thought (CoT) reasoning has emerged as an effective approach for activating latent reasoning capabilities in large language models. However, most existing CoT methods use reasoning chains mainly as inference-time prompts, while the generated reasoning traces are rarely reused as semisupervised learning signals. In this report, we define Semi-supervised Chain-of-Thought Learning and propose Semi-CoT, a simple framework that uses unlabeled questions to construct pseudo reasoning supervision. Semi-CoT samples multiple pseudo-CoTs for each unlabeled question, estimates answer-level semantic entropy, and selects low-entropy reasoning chains as reliable pseudo-CoT demonstrations. This extends the self-training view of CoT from inference-time refinement to semi-supervised pseudo-supervision. Pilot experiments on AQuA, SVAMP, GSM8K, and MultiArith show that the entropy gate selects high-precision pseudo-CoTs, with pseudo-answer precision ranging from 91.36% to 100% . Semi-CoT also gives small gains on SVAMP and GSM8K, while AQuA shows negative transfer and MultiArith reaches a ceiling. These results suggest that unlabeled questions can provide reliable pseudo reasoning signals, but their effective use still requires stronger demonstrationselection or student training.  \n1 Introduction  \nChain-of-thought (CoT) reasoning has become an effective way to elicit the reasoning ability of large language models (LLMs) . By asking a model to generate intermediate reasoning steps before producing the final answer, CoT improves performance on arithmetic, symbolic, and commonsense reasoning tasks [19 , 38] . A series of follow-up studies further improve CoT through self-consistency, automatic demonstration construction, planning-based prompting, contrastive reasoning, and iterative re-reading [4, 36 , 37 , 41 , 44] . Recent surveys also show that CoT has become a central component of reasoning with foundation models [5, 32] . However, a common practice is still to use CoT only at inference time. The model generates reasoning paths for a test question, uses them to reach a final answer, and then discards the generated reasoning traces.  \nThis view leaves an important question underexplored: can model-generated reasoning chains be used as semi-supervised learning signals? In many reasoning tasks, high-quality CoT annotations are expensive. Annotating a final answer already requires human effort, and writing a complete reasoning chain requires even more cost. Moreover, different annotators may solve the same problem through different valid reasoning paths.  \nAt the same time, unlabeled questions are often much easier to collect. For example, math word problems, science questions, coding questions, and logic questions can exist in large quantities without human-written rationales [6, 9] . These unlabeled questions may still contain useful reasoning structures that can be discovered by an LLM.  \nOur motivation is closely related to self-training and semi-supervised learning. Self-training uses modelgenerated pseudo-labels to exploit unlabeled data and has a long history in machine learning [1, 21 , 30] . Modern semi-supervised learning further improves pseudo-labeling with entropy minimization, consistency regularization, confidence thresholding, curriculum pseudo-labeling, and teacher-student learning [10, 31 , 34 , 42 , 43] . These methods show that unlabeled data can be useful when pseudo-labels are sufficiently reliable. However, standard pseudo-labeling usually considers only the final label. For reasoning tasks, the supervision signal should include not only the final answer but also the reasoning process that leads to it.  \nRece","cbCaiszxXaHRLYaO","https://ap.wps.com/l/cbCaiszxXaHRLYaO","pdf",1216805,5,1,18,"English","en",105,"# Abstract\n# Introduction\n## Motivation: CoT as inference-only\n## Unlabeled questions and missing supervision\n## Self-training link and new problem setting\n## Reliability challenges and need for selection\n## Semantic entropy reliability mechanism","[{\"question\":\"What problem does Semi-supervised Chain-of-Thought Learning address?\",\"answer\":\"It addresses the gap where CoT methods use reasoning only at inference time, discarding traces instead of reusing them as learning signals. The goal is to leverage unlabeled questions to produce pseudo supervision that includes both an answer and a reasoning chain.\"},{\"question\":\"What do the pilot experiments suggest about unlabeled pseudo reasoning signals?\",\"answer\":\"Entropy gating selects high-precision pseudo-CoTs with pseudo-answer precision ranging from 91.36% to 100%. It yields small gains on SVAMP and GSM8K, shows negative transfer on AQuA, and reaches a ceiling on MultiArith.\"}]","Revisiting Chain-of-Thought Reasoning under Limited Supervision: Semi-supervised Chain-of-Thought Learning | PDF",1784176369,45,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":84,"head_meta":86,"extra_data":88,"updated_unix":29},"revisiting-chain-of-thought-reasoning-under-limited-supervision-semi-supervised-chain-of-thought-learning","",{"@graph":37,"@context":83},[38,55,70],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":54},"https://docshare.wps.com/document/revisiting-chain-of-thought-reasoning-under-limited-supervision-semi-supervised-chain-of-thought-learning/81820/",4,{"url":53,"name":13,"@type":56,"author":57,"headline":13,"publisher":59,"fileFormat":62,"inLanguage":24,"description":14,"dateModified":63,"datePublished":64,"encodingFormat":62,"isAccessibleForFree":65,"interactionStatistic":66},"DigitalDocument",{"name":9,"@type":58},"Person",{"url":42,"name":60,"@type":61},"DocShare","Organization","application/pdf","2026-07-29","2026-07-16",true,{"@type":67,"interactionType":68,"userInteractionCount":20},"InteractionCounter",{"@type":69},"ViewAction",{"@type":71,"mainEntity":72},"FAQPage",[73,79],{"name":74,"@type":75,"acceptedAnswer":76},"What problem does Semi-supervised Chain-of-Thought Learning address?","Question",{"text":77,"@type":78},"It addresses the gap where CoT methods use reasoning only at inference time, discarding traces instead of reusing them as learning signals. The goal is to leverage unlabeled questions to produce pseudo supervision that includes both an answer and a reasoning chain.","Answer",{"name":80,"@type":75,"acceptedAnswer":81},"What do the pilot experiments suggest about unlabeled pseudo reasoning signals?",{"text":82,"@type":78},"Entropy gating selects high-precision pseudo-CoTs with pseudo-answer precision ranging from 91.36% to 100%. It yields small gains on SVAMP and GSM8K, shows negative transfer on AQuA, and reaches a ceiling on MultiArith.","https://schema.org",{"og:url":53,"og:type":85,"og:title":13,"og:site_name":60,"og:description":14},"article",{"robots":87,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":90},[91,95,99,103,107,112,117,120,125,128,132],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":92,"show_sort_weight":93,"slug":94},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":96,"show_sort_weight":97,"slug":98},"Literature",80,"literature",{"id":54,"doc_module":4,"doc_module_name":47,"category_name":100,"show_sort_weight":101,"slug":102},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":104,"show_sort_weight":105,"slug":106},"Comic",60,"comic",{"id":108,"doc_module":4,"doc_module_name":47,"category_name":109,"show_sort_weight":110,"slug":111},6,"Technology",50,"technology",{"id":113,"doc_module":4,"doc_module_name":47,"category_name":114,"show_sort_weight":115,"slug":116},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":118,"slug":119},30,"research-report",{"id":121,"doc_module":4,"doc_module_name":47,"category_name":122,"show_sort_weight":123,"slug":124},9,"Religion & Spirituality",20,"religion-spirituality",{"id":123,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":123,"slug":127},"World Cup","world-cup",{"id":129,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":129,"slug":131},10,"Lifestyle","lifestyle",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":20,"slug":135},19,"General","general"]