[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84773-en":3,"doc-seo-84773-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84773,4398048950312,"Violet","https://ap-avatar.wpscdn.com/avatar/400002538284de19e3c?_k=1778320343897328908",8,"Research & Report","Input Pathways Shape Few-Shot Not Zero-Shot Binding in Tiny Transformers: A Fully-Enumerable Study","How information enters a transformer—via symbolic in-context tokens, a clean per-factor oracle code, or an entangled perceptual vector—determines whether compositional binding is learned. The study uses tiny ∼6–10K parameter transformers on exhaustively enumerated factored worlds, evaluating every model over the entire input space with exact Bayes ceilings and repeated nonparametric robustness tests. Results include endpoint invariance for zero-shot, a two-factor account linking few-shot binding to pathway sharing and code readability, a training-time double dissociation, and a failure anatomy pinpointing readout loss versus mis-binding.","arXiv :2607 .04926v 1 [ cs .LG] 6 Jul 2026  \nInput Pathways Shape Few-Shot, Not Zero-Shot, Binding in Tiny Transformers: A Fully-Enumerable Study  \nYoshiyuki Ootani [info@ootanl. com](info@ootanl. com)  \nIndependent Researcher  \nAbstract  \nHow does the way information reaches a transformer—as symbolic in-context tokens, asa clean per-factor “oracle” code, or as an entangled perceptual vector—affect whether the model can bind that information compositionally? We study this in a deliberately extreme regime: transformers of ∼6–10K parameters on finite factored worlds that we enumerate exhaustively, so every behavioral measurement is evaluated on the entire input space (zero sampling variance), the informative input channels are matched in the sense that each uniquely determines the answer (verified by exact Bayes ceilings), and every headline comparison is run at 10–20 seeds with paired nonparametric tests (robustness arms use 5–20) .  \nWe report four findings. (1) Endpoint invariance: on held-out binding queries no informative route yields converged zero-shot composition—each ends at or below chance despite an exact Bayes ceiling of 1.0, so within our bounded sweep (a ∼25 × parameter increase, a 10× learning-rate range, weight decay, and longer training) all routes converge to lookup-like solutions despite information-sufficient inputs. (The ceiling certifies that the input determines the answer, not that the held-out mapping is identifiable from a training objective for which lookup suffices.) (2) A two-factor account of few-shot binding: when a small fraction of the held-out query type is leaked into training, within the tested cells sample efficiency is best predicted by (i) input-pathway parameter sharing (a shared projection transfers to unseen query types; per-factor embedding tables do not) and (ii) readability of the code (a poorly readable entangled code, whose raw linear decodability drops 0.95 → 0.58, has the worst pooled few-shot performance and saturates far below the readable alternatives;  \na dimension-matched control and a graded readability sweep—few-shot accuracy is monotone in decodability at fixed input dimension—isolate readability from dimension) . Two parameter-controlled pairwise dissociations (a three-cell partial factorial) separate the factors; the two core effects (sharing helps, low readability hurts) replicate across two held-out shape query types at 20 seeds, and are corroborated as a stress check in a modestly larger three-object world, while the weak-perceptual edge over the oracle is testbed-specific. Notably, the clean per-factor oracle is not the most sample-efficient readable route; shared readable pathways transfer better. (3) A double dissociation: early in training, distributed codes—but not index-like codes—pass through a transient phase of above-chance zero-shot transfer before collapsing into memorization (clearly for the symbolic and weak-perceptual routes; the strong-entangled route’s collapse is comparable but its peak CI overlaps chance);  \nthis trajectory effect tracks code format, whereas few-shot efficiency tracks pathway sharing.  \n(4) Failure anatomy: the symbolic route fails by losing the answer at the readout position;  \nindex routes fail by systematically mis-binding—the answer stays decodable in the residual, yet a direct input intervention shows the converged output is more sensitive to the wrong object slot than to the correct one—and the entangled route inherits, rather than improveson, the readability of its input. Of these, the central positive claim is the two-factor account of few-shot efficiency; endpoint invariance and failure anatomy are diagnostic constraints on it, and the transient a dynamical correlate. All code, manifests, and per-seed logs are available for exact reproduction.  \n1 Introduction  \nA transformer can receive the same fact in different ways.“The object on the right is a square” can arrive as words in the token stream, as a slot in a structured st","cbCaid4SE4rLuy2b","https://ap.wps.com/l/cbCaid4SE4rLuy2b","pdf",909623,1,16,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"What question does the study investigate about transformer binding?\",\"answer\":\"It examines how different information routes into a transformer (symbolic in-context tokens, a per-factor oracle code, or an entangled perceptual vector) affect whether the model can bind information compositionally.\"},{\"question\":\"Why does the paper claim to remove confounds from sampling noise and information mismatch?\",\"answer\":\"It evaluates models exhaustively over an enumerated entire query space (zero sampling variance), matches the informative input channels by construction, and computes exact Bayes-optimal ceilings per route while running robustness tests over multiple seeds.\"},{\"question\":\"What are the main findings about few-shot versus zero-shot binding?\",\"answer\":\"Held-out binding queries show endpoint invariance: informative routes fail to achieve converged zero-shot composition despite Bayes ceilings. For few-shot, performance is best predicted by pathway parameter sharing and by code readability, with poor readability entangled codes producing the worst outcomes.\"}]",1784198143,40,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"input-pathways-shape-few-shot-not-zero-shot-binding-in-tiny-transformers-a-fully-enumerable-study","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/input-pathways-shape-few-shot-not-zero-shot-binding-in-tiny-transformers-a-fully-enumerable-study/84773/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What question does the study investigate about transformer binding?","Question",{"text":75,"@type":76},"It examines how different information routes into a transformer (symbolic in-context tokens, a per-factor oracle code, or an entangled perceptual vector) affect whether the model can bind information compositionally.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why does the paper claim to remove confounds from sampling noise and information mismatch?",{"text":80,"@type":76},"It evaluates models exhaustively over an enumerated entire query space (zero sampling variance), matches the informative input channels by construction, and computes exact Bayes-optimal ceilings per route while running robustness tests over multiple seeds.",{"name":82,"@type":73,"acceptedAnswer":83},"What are the main findings about few-shot versus zero-shot binding?",{"text":84,"@type":76},"Held-out binding queries show endpoint invariance: informative routes fail to achieve converged zero-shot composition despite Bayes ceilings. For few-shot, performance is best predicted by pathway parameter sharing and by code readability, with poor readability entangled codes producing the worst outcomes.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":28,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]