[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84256-en":3,"doc-seo-84256-105":30,"detail-sidebar-cat-0-en-105":84},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84256,13056703019662,"Evangeline","https://ap-avatar.wpscdn.com/avatar/be000253a8e92610077?_k=1778726343310543188",8,"Research & Report","RL Post-Training Builds Compositional Reasoning Strategies","Whether RL post-training simply amplifies latent skills in a base model or composes primitive skills into higher-level strategies is tested in a fully observable rewrite-grammar environment. A Transformer is pretrained on primitive symbol-rewrite chains and post-trained on a trace-based contraction task with a binary reward. RL solves heldout problems rarely solved by pretraining, outperforming rejection fine-tuning after early gains plateau. Trace analysis shows phased procedural chunking: RL strengthens primitive reductions, then discovers reusable sequential and parallel composed procedures, while selection—not exploration volume—drives validity and consolidation. Pretraining determines whether reduction procedures exist for RL to compress and reuse.","RL Post-Training Builds Compositional Reasoning Strategies  \nAzwar Abdulsalam 1 Nishil Patel 1 Andrew Saxe 1  \narXiv :2607 .07646v 1 [ cs .AI] 8 Jul 2026  \nAbstract  \nDoes RL post-training merely amplify primitive skills already latent in a base model, or can it compose primitive skills into new higher-level strategies? We study this question in a fully observable rewrite-grammar environment where thepretraining distribution is known and every generated rewrite can be audited. A Transformer ispretrained on primitive symbol-rewrite chains and post-trained on a Trace-based reasoning task with only a binary final-answer reward. RL solves heldout problems that remain rarely solved by the pretrained model even under much larger sampling budgets, while rejection fine-tuning improves early but plateaus. Trace analysis shows that RL reorganizes primitive competence through a phased compositional mechanism: it first strengthens primitive reductions, then discovers valid composed procedures. These include sequential compositions, which collapse ordered chains of primitive contractions, and parallel compositions, which combine independent primitive contractions in a single step. The composed procedures are not isolated samples; they are reused and consolidated into a stable repertoire. Comparing RL with rejection fine-tuning shows that the key difference is not exploration volume but selectivity: RFT produces many shortcut-like rewrites, much of them invalid, whereas RL concentrates exploration into valid reusable structure. Pretraining ablations show that the emergence of compositional strategies is gated not by primitive exposure alone, but by whether pretraining organizes primitive competence into reduction procedures that RL can later compress. The base model provides weak procedural ingredients; RL builds them into reliable higher-level strategies.  \n1 Gatsby Computational Neuroscience Unit, UCL, London, United Kingdom. Correspondence to: Azwar Abdulsalam \u003C[azwar.azwar.25@ucl.ac.uk](azwar.azwar.25@ucl.ac.uk)> .  \nAccepted to the 2nd Workshop on Compositional Learning at ICML 2026, Seoul, South Korea. Copyright 2026 by the author(s) .  \n1. Introduction  \nDoes reinforcement learning (RL) post-training merely reweight behaviors already latent in a base model, or can it compose primitive skills into new higher-level strategies? Recent work has sharpened this into an active debate. Yue et al. (Yue et al., 2025) argue that RL with verifiable rewards often improves small-k success without expanding the largek capability frontier; ProRL (Liu et al., 2025a) and Yuan et al. (Yuan et al., 2025) present evidence that prolonged RL or controlled compositional tasks can uncover behaviors inaccessible to the base model under extensive sampling; and imitation-style baselines such as rejection fine-tuning are sometimes surprisingly competitive (Chu et al., 2025 ; Xiong et al., 2025) .  \nA central obstacle is that the relevant mechanisms are hard to observe in pretrained language models. When a post-trained model exhibits a new behavior, it is usually unclear whether that behavior was absent from the base model, merely lowprobability, or already present in pretraining data. Aggregate metrics such as pass@k leave three questions open: what strategies are being used, whether their emergence reflects broader exploration or selective filtering, and what pretraining structure makes them reachable in the first place.  \nWe study these questions in a fully observable rewritegrammar environment. The task abstracts a common structure in step-by-step reasoning: complex solutions can be built by composing local transformations into reusable procedures. A Transformer is pretrained from scratch on primitive rewrite chains and then post-trained on goal-directed contraction task. Because every generated rewrite can be audited against the grammar, we can decompose behavior into primitive rule use, valid composed strategies, and spurious invalid rewrites. The valid com","cbCaii3gQ3eKlK1Z","https://ap.wps.com/l/cbCaii3gQ3eKlK1Z","pdf",1530814,6,1,12,"English","en",105,"# Abstract\n# Introduction\n# Related Work","[{\"question\":\"Why does RL outperform rejection fine-tuning in this setup?\",\"answer\":\"The key difference is selectivity: rejection fine-tuning generates many shortcut-like, often invalid rewrites, while RL concentrates exploration into valid, reusable structure.\"}]",1784194412,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":79,"head_meta":81,"extra_data":83,"updated_unix":28},"rl-post-training-builds-compositional-reasoning-strategies","",{"@graph":36,"@context":78},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/rl-post-training-builds-compositional-reasoning-strategies/84256/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72],{"name":73,"@type":74,"acceptedAnswer":75},"Why does RL outperform rejection fine-tuning in this setup?","Question",{"text":76,"@type":77},"The key difference is selectivity: rejection fine-tuning generates many shortcut-like, often invalid rewrites, while RL concentrates exploration into valid, reusable structure.","Answer","https://schema.org",{"og:url":52,"og:type":80,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":82,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":85},[86,90,94,98,103,107,112,114,119,122,126],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":87,"show_sort_weight":88,"slug":89},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":91,"show_sort_weight":92,"slug":93},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Exam",70,"exam",{"id":99,"doc_module":4,"doc_module_name":46,"category_name":100,"show_sort_weight":101,"slug":102},5,"Comic",60,"comic",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":46,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":113},"research-report",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},9,"Religion & Spirituality",20,"religion-spirituality",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":120,"show_sort_weight":117,"slug":121},"World Cup","world-cup",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":123,"slug":125},10,"Lifestyle","lifestyle",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":99,"slug":129},19,"General","general"]