[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81625-en":3,"doc-seo-81625-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81625,8796095360427,"Lucas Martin","https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d",8,"Research & Report","Hierarchical Chain-of-Thought Enhancing LLM Reasoning Performance and Efficiency","Chain-of-Thought (CoT) prompting improves large language model reasoning, yet traditional CoT often produces long, unstructured reasoning traces with redundancy and inefficiency on complex, multi-step problems. Hierarchical Chain-of-Thought (Hi-CoT) organizes reasoning into hierarchical substeps by alternating instructional planning with step-by-step execution, improving long-horizon coherence. Evaluations across diverse LLMs and math benchmarks show Hi-CoT raises average accuracy by 6.2% (up to 61.4%) and reduces reasoning trace length by 13.9% versus CoT. Strict adherence to the hierarchy maximizes accuracy and efficiency.","Hierarchical Chain-of-Thought: Enhancing LLM Reasoning Performance and Efficiency  \nXingshuai Huang 1 Derek Li 2 Bahareh Nikpour 1 Parsa Omidi 1  \narXiv :2604 .00 130v2 [ cs .CL] 10 Jul 2026  \nAbstract  \nChain-of-Thought (CoT) prompting has significantly improved the reasoning capabilities of large language models (LLMs) . However, conventional CoT often relies on unstructured, flat reasoning chains that suffer from redundancy and suboptimal performance. In this work, we introduce Hierarchical Chain-of-Thought (Hi-CoT), a structured reasoning paradigm specifically designed to address the challenges of complex, multi-step reasoning. Hi-CoT decomposes the reasoning process into hierarchical substeps by alternating between instructional planning and stepby-step execution. This decomposition enables LLMs to better manage long reasoning horizonsand maintain logical coherence. Extensive evaluations across diverse LLMs and mathematical  \nreasoning benchmarks show that Hi-CoT consistently improves average accuracy by 6.2%(up to 61.4% on certain models and tasks) while reducing reasoning trace length by 13.9% compared to CoT. We further show that accuracy and efficiency are maximized when models strictly adhere to the hierarchical structure. Our code is available at [https://github.com/XingshuaiHuang/Hi-CoT](https://github.com/XingshuaiHuang/Hi-CoT).  \n1. Introduction  \nLarge Language Models (LLMs) have shown strong capabilities on reasoning-intensive tasks when paired with appropriate prompting strategies. One influential approach is Chain-of-Thought (CoT) prompting (Wei et al., 2022), which encourages models to produce intermediate reasoning steps before a final answer. This simple modification has led to notable gains on multi-step reasoning tasks(Puri et al., 2021 ; Zhou et  al., 2025), such as mathematic tasks (Cobbe  \n1Huawei Technologies Canada, Canada 2Huawei Noah’s Ark Lab, Canada. Correspondence to: Xingshuai Huang \u003Cxing[shuai.huang@gmail.com](shuai.huang@gmail.com) >.  \nAccepted to the 4th Workshop on Planning in the Era of LLMs, LM4Plan@ICML 2026, Seoul, South Korea. Copyright 2026 by the author(s) .  \nAccuracy (%)  \n100  \n80  \n60  \n40  \n20  \n0  \nAMC MATH500  \nMinerva OlympiadBench  \n3500  \n3000  \n2500  \n2000  \n1500  \n1000  \nAvg tokens  \nStandard  Hi-CoT  \n CoT  Hi-CoT (correct format)  \n Plan-and-Solve  AMC (tokens)  \n Hi-CoT (format-relaxed)  \n MATH500 (tokens)  \n Minerva (tokens)  \n OlympiadBench (tokens)  \nFigure 1. Accuracy (%) and average token length for Qwen3-8B model (Yang et al., 2025a) across multiple prompting methods on mathematical reasoning benchmarks. For Hi-CoT (correct format), results are computed only from responses that strictly follow the hierarchical structure.  \net al., 2021 ; Liu et al., 2025a) .  \nDespite its success, standard CoT prompting has a strcutural shortcoming. The generated reasoning traces are typically linear and unstructured, which can introduce redundant steps (e.g., repeating explanations) and occasional deviations from the intended line of reasoning (Wang et al., 2023) . Redundancy arises because CoT imposes no compression presssure on the reasoning process. Without an explicit mechanism to filter low-value content, the model can repeat itself, hedge, or wander without penalty. Longer traces do not imply better reasoning (Liu et al., 2025b); they often reflect disorganized exploration rather than deliberate problem solving, while simultaneously incurring higher inference costs as computational overhead of transformer models grows with the number of generated tokens(Zeng et al., 2026 ; Omidiet al., 2025) . This lack of structure is particularly costly for problems that require coordinated planning across multiple steps, where unguided reasoning frequently fails to maintain logical coherence(He et al., 2024 ; Lewkowycz et al., 2022) .  \nOne solution is to impose a plan. Plan-and-Solve prompting (Wang et al., 2023) does exactly this, separating reasoning into a global planning phase fo","cbCaiiIJYsLedsFn","https://ap.wps.com/l/cbCaiiIJYsLedsFn","pdf",348418,4,1,12,"English","en",105,"# Abstract\n# Introduction\n## Challenges of standard CoT prompting\n## Plan-and-Solve and plan–execution drift\n## Proposed Hi-CoT paradigm","[{\"question\":\"What problem with standard CoT prompting does Hi-CoT address?\",\"answer\":\"Standard CoT often generates linear, unstructured reasoning traces that include redundant steps and can drift away from the intended reasoning path, increasing token cost. Hi-CoT introduces structure to reduce redundancy and maintain logical coherence.\"},{\"question\":\"How does Hi-CoT structure the reasoning process?\",\"answer\":\"Hi-CoT decomposes reasoning into hierarchical substeps by alternating between instructional planning and step-by-step execution. Each execution step is conditioned on the previous execution result to enable continuous plan refinement.\"},{\"question\":\"What improvements does Hi-CoT achieve compared with conventional CoT?\",\"answer\":\"Across diverse LLMs and mathematical reasoning benchmarks, Hi-CoT improves average accuracy by 6.2% and reduces reasoning trace length by 13.9% relative to CoT. Adhering strictly to the hierarchical structure further maximizes accuracy and efficiency.\"}]",1784174923,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"hierarchical-chain-of-thought-enhancing-llm-reasoning-performance-and-efficiency","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/hierarchical-chain-of-thought-enhancing-llm-reasoning-performance-and-efficiency/81625/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem with standard CoT prompting does Hi-CoT address?","Question",{"text":75,"@type":76},"Standard CoT often generates linear, unstructured reasoning traces that include redundant steps and can drift away from the intended reasoning path, increasing token cost. Hi-CoT introduces structure to reduce redundancy and maintain logical coherence.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does Hi-CoT structure the reasoning process?",{"text":80,"@type":76},"Hi-CoT decomposes reasoning into hierarchical substeps by alternating between instructional planning and step-by-step execution. Each execution step is conditioned on the previous execution result to enable continuous plan refinement.",{"name":82,"@type":73,"acceptedAnswer":83},"What improvements does Hi-CoT achieve compared with conventional CoT?",{"text":84,"@type":76},"Across diverse LLMs and mathematical reasoning benchmarks, Hi-CoT improves average accuracy by 6.2% and reduces reasoning trace length by 13.9% relative to CoT. Adhering strictly to the hierarchical structure further maximizes accuracy and efficiency.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]