[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86299-en":3,"doc-seo-86299-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86299,13056703020460,"Valentina","https://ap-avatar.wpscdn.com/avatar/be000253dac470eee5d?_k=1778207105932848923",8,"Research & Report","Think Through a Bottleneck: Hourglass Reasoning for Rigorous Induction","Self-refinement in large language models often fails to improve few-shot inductive reasoning, especially when models can directly patch outputs after feedback rather than consolidating the underlying rule. Hourglass reasoning introduces structurally enforced isolation between reasoning stages via a frozen meta-constructor that builds an encoder–decoder symbolic bottleneck. Only the compressed symbolic state crosses stages, so refinement remains anchored to the inferred rule. Experiments across visual abstraction, hardware synthesis, and textual rule induction show large accuracy and synthesis gains and ablations confirm the effect comes from stage isolation and induction quality, not prompt wording.","arXiv :2607 . 1 1696v 1 [ cs .AI] 13 Jul 2026  \nThink Through a Bottleneck: Hourglass Reasoning for  \nRigorous Induction  \nHuan Zhu∗  \nPeking University  \nABSTRACT  \nSelf-refinement often fails to strengthen few-shot inductive reasoning in large language models. Prompting a model to explicitly state its inferred rule does little on its own. What actually matters is a structurally enforced isolation between reasoning stages, so that information can only pass between them as a compressed symbolic state.  \nWe introduce Hourglass reasoning, which enforces strict context isolation between reasoning stages. The frozen LLM acts as a meta-constructor, building for each task a symbolic encoder–decoder: an Induction module compresses the support examples into a schema 􀁱 (encoder) and a transient scaffold 􀁉 ; a Deduction module derives rule 􀀩 (decoder) from these and discards 􀁉 ; an Implementer compiles (􀁱, 􀀩) into artifacts; an error-driven Refiner revises (􀁱, 􀀩) and regenerates artifacts from scratch. Only (􀁱, 􀀩) crosses stage boundaries, so all refinement stays anchored to the rule.  \nWe evaluate Hourglass across three benchmarks spanning visual abstraction, hardware synthesis, and textual rule induction, using GPT-5.5 and Gemini 3.1 Pro. On ARC-AGI-2, it raises best-of-5 accuracy by up to 14 points over an iterative-refinement baseline. On ChipBench, it nearly doubles Verilog synthesis accuracy with GPT-5.5, from 31% to 58% . BBEH-Linguini draws on puzzles from the International Linguistics Olympiad, a setting where prior work has shown that explicit verbalization can hurt performance. Hourglass mitigates this tendency, and on Gemini 3.1 Pro, it reverses the effect entirely.  \nAblations confirm that these gains come from the isolation between stages and the quality of the initial induction, not from prompt wording or the particular symbolic form used. It is how information flows through the reasoning process, rather than the language used to express it, that drives inductive reasoning in frozen LLMs.  \n1 Introduction  \nHumans can extract abstract rules from only a handful of examples. In artificial intelligence, this capability is rigorously tested by benchmarks like ARC-AGI-2 (Chollet, 2019), where each puzzle is governed by a single, precise transformation rule that must be inferred from few-shot demonstrations. Despite impressive  \n∗ Code and prompts: [https://github.com/ZhuHuan09/hourglass-reasoning](https://github.com/ZhuHuan09/hourglass-reasoning)  \nperformance on many natural-language tasks, current frontier LLMs still struggle with such rule-centric induction.  \nLarge language models have demonstrated remarkable capabilities across a broad spectrum, extending from natural language processing and code generation to hardware description language synthesis (Liu, Y., et al., 2025) and abstract spatial reasoning on benchmarks such as ARC-AGI-2 (Franzen et al., 2025) . However, these models remain prone to shortcut learning (Geirhos et al., 2019): instead of abstracting the latent rule, a model exploits superficial regularities in the support examples. When execution feedback is available, this often manifests as patchwork logic: hardcoded if-else branches keyed to specific coordinates, example indices, or local artifacts. Such patches force support examples to pass but fail on out-of-distribution queries (Moskvichev et al., 2023; Mitchell et al., 2023) . Moreover, a growing body of recent work demonstrates that the intrinsic self-correction capabilities of monolithic LLMs are brittle and frequently degrade performance in the absence of external structural guidance (Tsui, 2025; Sanz-Guerrero & Von Der Wense, 2025) .  \nThese pathologies are partly rooted in how information flows through dense, unrestricted context windows. When raw examples, current artifacts, error feedback, and repair instructions coexist without structured partition, the model tends to anchor on low-level perceptual details rather than generalizing to an abstract","cbCaik400v7VIsJh","https://ap.wps.com/l/cbCaik400v7VIsJh","pdf",490462,2,1,27,"English","en",105,"# Introduction\n## Bottleneck motivation and information flow\n## Hourglass reasoning overview\n## Evaluation across benchmarks\n## Ablation insights","[{\"question\":\"What problem does Hourglass reasoning target in few-shot inductive reasoning?\",\"answer\":\"It targets failures where self-refinement does not strengthen rule-centric induction in frozen LLMs. Models may instead exploit superficial regularities and patch outputs directly when feedback is available.\"},{\"question\":\"How does Hourglass reasoning enforce correct information flow across stages?\",\"answer\":\"It isolates reasoning stages so that only a compressed symbolic state (an encoder schema and a decoder rule) crosses stage boundaries. Intermediate traces, including transient scaffolds, are discarded to prevent instance-specific leakage.\"},{\"question\":\"What experimental benefits does Hourglass achieve and how are they validated?\",\"answer\":\"Across three benchmarks, Hourglass improves performance over a context-reset self-refinement baseline, including higher ARC accuracy and better Verilog synthesis. Ablations attribute gains to stage isolation and the quality of the initial induction rather than prompt wording or symbolic form.\"}]",1784210309,68,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"think-through-a-bottleneck-hourglass-reasoning-for-rigorous-induction","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/think-through-a-bottleneck-hourglass-reasoning-for-rigorous-induction/86299/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does Hourglass reasoning target in few-shot inductive reasoning?","Question",{"text":75,"@type":76},"It targets failures where self-refinement does not strengthen rule-centric induction in frozen LLMs. Models may instead exploit superficial regularities and patch outputs directly when feedback is available.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does Hourglass reasoning enforce correct information flow across stages?",{"text":80,"@type":76},"It isolates reasoning stages so that only a compressed symbolic state (an encoder schema and a decoder rule) crosses stage boundaries. Intermediate traces, including transient scaffolds, are discarded to prevent instance-specific leakage.",{"name":82,"@type":73,"acceptedAnswer":83},"What experimental benefits does Hourglass achieve and how are they validated?",{"text":84,"@type":76},"Across three benchmarks, Hourglass improves performance over a context-reset self-refinement baseline, including higher ARC accuracy and better Verilog synthesis. Ablations attribute gains to stage isolation and the quality of the initial induction rather than prompt wording or symbolic form.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]