[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85145-en":3,"doc-seo-85145-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85145,1374391974468,"Eden","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning","Robotic manipulation policies depend on how visual encoders represent scenes, since policies never observe raw inputs. A study compares dense global embeddings, dense patch grids, and object-centric slot representations on ManiSkill3 PickCube-v1 using a frozen encoder and held-out-seed evaluation. Object-centric SPOT (DINO ViT-B/16 + Slot Attention) yields 55.0±2.9% success, outperforming dense global features by 22.4%. Simply adding more tokens fails to improve dense representations. With explicit 2D spatial goals and native-resolution rendering, success reaches 68.7±4.2%, near a 3D-oracle upper bound. Failure taxonomy separates Near-Miss from No-Grasp, highlighting grounding benefits and occlusion bottlenecks transfer to StackCube-v1.","More Structure, Not More Capacity: Object-Centric Representations for Visuomotor  \nImitation Learning  \nYi Li 1 , Alexandre Chapin2 , Liming Chen2 , Jan Peters 1 , Alap Kshirsagar3  \n1TU Darmstadt 2École Centrale de Lyon 3IIT Delhi -Abu Dhabi  \narXiv :2607 .09825v1 [ cs .RO] 10 Jul 2026  \nAbstract—Robotic manipulation policies rely on pre-trained vision models that give either a global scene embedding or a dense patch grid. Both mix task-relevant and task-irrelevant features. Object-centric slot representations are a structured alternative: they group features into a few per-object slots. We test what this structure buys on ManiSkill3 PickCube-v1, with a frozen encoderand a held-out-seed evaluation. Holding the policy, goal token, rendering, and calibration fixed and changing only the encoder, a frozen object-centric SPOT representation (DINO ViT-B/16 + Slot Attention) reaches 55.0±2 .9% success, 22.4% above a dense DINO global-feature baseline (32 .6±1 .5%), with the same trainable policy and no encoder fine-tuning. More tokens alone do not help: a dense patch grid with 16 × the tokens performs no better than the global feature. Adding an explicit 2D spatial goal and native-resolution rendering raises the full system to 68.7±4 .2%, just below a privileged 3D-oracle upper bound (71 .7±4 . 1%). An automated kinematic failure taxonomy then separates spatialprecision (Near-Miss) failures from object-tracking (No-Grasp) failures: spatial grounding reduces Near-Miss while leaving NoGrasp unchanged. The same taxonomy transfers to the harder StackCube-v1 and points to occlusion as the main bottleneck.  \nI. INTRODUCTION  \nA visuomotor policy never sees the raw scene. It only sees what its visual encoder keeps. The choice of representation therefore decides what information reaches the policy. In robotics this question is most developed for mapping and localization, where geometric, probabilistic, dense, and implicit map representations each shape behavior differently [1] . It applies just as much to manipulation: how does the structure of the visual representation affect behavior when the policy must act on object and goal positions it never saw during training? We answer this with a controlled, task-level study on a simulated pick-and-place task.  \nSelf-supervised vision transformers such as DINO [2] transfer well to recognition. It is less clear how they behave when a policy must act on scene configurations not seen in training. Recent work across image classification [14], 3D perception [17], and gaze estimation [11] points to one shared pattern: a backbone trained under global or invariance-oriented supervision finds a low-cost cue that holds in-distribution and breaks under shift, and in each case the fix is structural rather than adding parameters.  \nThis paper studies a similar pattern in visuomotor imitation. Under matched conditions, we compare dense DINO features (both the global [CLS] token and dense patch  \ngrids) against an object-centric slot representation built on the same frozen backbone. We treat the slot representation as a small, competition-based bottleneck. Slot Attention [9], as implemented in SPOT [6], compresses the dense patches into a few slots and makes each slot bind to one coherent region. This is an architectural inductive bias toward objectlevel structure. Since the encoder stays frozen, this bias has to come from the module placed on top of its features, not from retraining the backbone. A strong pretrained backbone is not enough when the task needs signals that the pretraining objective suppressed.  \nIn this work, we make the following contributions:  \n1) A systematic, matched-condition comparison of dense global, dense patch, and object-centric representations for visuomotor imitation, under a single policy and a held-out-seed protocol that measures generalization to novel object and goal placements (a task-level evaluation of representation choice) .  \n2) The finding that representational structure,","cbCaieN2U6g0Wtyf","https://ap.wps.com/l/cbCaieN2U6g0Wtyf","pdf",654406,2,1,7,"English","en",105,"# Abstract\n# Introduction\n# Related Work","[{\"question\":\"Why does visual representation structure matter for visuomotor imitation learning?\",\"answer\":\"A visuomotor policy depends entirely on what the visual encoder preserves from the scene. Representation structure determines which task-relevant versus task-irrelevant information reaches the policy during action and generalization.\"},{\"question\":\"What performance advantage do object-centric slot representations provide?\",\"answer\":\"On ManiSkill3 PickCube-v1 with a frozen encoder and held-out-seed protocol, a frozen object-centric SPOT representation achieves 55.0±2.9% success, 22.4% higher than a dense DINO global-feature baseline (32.6±1.5%).\"},{\"question\":\"How does the failure taxonomy explain different error sources?\",\"answer\":\"An automated kinematic failure taxonomy distinguishes spatial-precision failures (Near-Miss) from object-tracking failures (No-Grasp). Spatial grounding reduces Near-Miss while leaving No-Grasp unchanged, and the same taxonomy on StackCube-v1 points to occlusion as the main bottleneck.\"}]",1784201370,18,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"more-structure-not-more-capacity-object-centric-representations-for-visuomotor-imitation-learning","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/more-structure-not-more-capacity-object-centric-representations-for-visuomotor-imitation-learning/85145/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why does visual representation structure matter for visuomotor imitation learning?","Question",{"text":75,"@type":76},"A visuomotor policy depends entirely on what the visual encoder preserves from the scene. Representation structure determines which task-relevant versus task-irrelevant information reaches the policy during action and generalization.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What performance advantage do object-centric slot representations provide?",{"text":80,"@type":76},"On ManiSkill3 PickCube-v1 with a frozen encoder and held-out-seed protocol, a frozen object-centric SPOT representation achieves 55.0±2.9% success, 22.4% higher than a dense DINO global-feature baseline (32.6±1.5%).",{"name":82,"@type":73,"acceptedAnswer":83},"How does the failure taxonomy explain different error sources?",{"text":84,"@type":76},"An automated kinematic failure taxonomy distinguishes spatial-precision failures (Near-Miss) from object-tracking failures (No-Grasp). Spatial grounding reduces Near-Miss while leaving No-Grasp unchanged, and the same taxonomy on StackCube-v1 points to occlusion as the main bottleneck.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]