[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-133381-en":3,"doc-seo-133381-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},133381,687197207639,"Asher","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","PEEK - Policy-agnostic Extraction of Essential Keypoints","Robotic manipulation policies often struggle to generalize because they must jointly learn where to attend, what object actions to take, and how to execute them. PEEK offloads high-level “where and what” reasoning to vision-language models while keeping policies focused on “how” low-level control. It fine-tunes VLMs to output a unified point-based intermediate representation: end-effector paths for action guidance and task-relevant masks for attention. An automatic annotation pipeline generates labeled data across 20+ robot datasets with 9 embodiments, enabling scalable training. Real-world evaluations show consistent zero-shot generalization gains, including 41.4× improvement for a simulation-only 3D policy and 2–3.5× gains for both large and small manipulation policies.","PEEK: Guiding and Minimal Image Representations for Zero-Shot Generalization of Robot Manipulation Policies  \nJesse Zhang⋆1 ,2 ,3 , Marius Memmel⋆1 ,2 , Kevin Kim3 , Dieter Fox 1 ,4 , Jesse Thomason3 , Fabio Ramos2 , Erdem Bıyık3 , Abhishek Gupta†1, Anqi Li†2  \narXiv :2509 . 18282v1 [ cs .RO] 22 Sep 2025  \nAbstract—Robotic manipulation policies often fail to generalize because they must simultaneously learn where to attend, what actions to take, and how to execute them. We argue that high-level reasoning about where and what can be offloaded to vision-language models (VLMs), leaving policies to specialize in how to act. We present PEEK (Policy-agnostic Extraction of Essential Keypoints), which fine-tunes VLMs to predict a unified point-based intermediate representation: (1) endeffector paths specifying what actions to take, and (2) taskrelevant masks indicating where to focus. These annotations are directly overlaid onto robot observations, making the representation policy-agnostic and transferable across architectures. To enable scalable training, we introduce an automatic annotation pipeline, generating labeled data across 20+ robot datasets spanning 9 embodiments. In real-world evaluations, PEEK consistently boosts zero-shot generalization, including a 41.4× real-world improvement for a 3D policy trained only in simulation, and 2–3.5× gains for both large VLAs and small manipulation policies. By letting VLMs absorb semantic and visual complexity, PEEK equips manipulation policies with the minimal cues they need—where, what, and how. Website at [https://peek-robot.github.io](https://peek-robot.github.io).  \nI. INTRODUCTION  \nImagine walking through a crowded store when your child suddenly cries out, “I want the Labubu!” Though you’ve never heard the word before, context clues guide your eyes to the fuzzy toy on the shelf, and you effortlessly weave through the crowd to grab it. What makes this possible is not raw perception ability, but the ability to interpret ambiguous instructions and distill them into just the right cues—where to focus, what actions to take, and how to perform these actions at the low level. Similarly, if given where to focus and what motions to take, a robot manipulation policy should be able to achieve the visual robustness and semantic generalization necessary for open-world deployment by focusing only on how to perform actions.  \nA common tactic for training manipulation policies is through imitation learning of human-collected robotics data [1]–[4], which attempts to learn the where, what, and how all at the same time. Yet their performance degradeson novel objects, clutter, or semantic variations [5], [6], since the policy alone bears the burden of handling task, semantic, and visual complexity. Such failures often entangle the axes of where, what, and how—for example, grasping a distractor simultaneously reflects misplaced attention, an incorrect object choice, and a wrong motion.  \n⋆ Co-first authors, †Equal Advising, 1University of Washington, 2NVIDIA, 3University of Southern California, 4Allen Institute for AI  \nFig. 1: PEEK enables policy generalization by modulating minimal representations of where to focus and what to do for robust policy learning.  \nOur key idea is to offload high-level reasoning to visionlanguage models (VLMs), which can excel at semantic and visual generalization [7], [8], leaving the policy to determine how low-level behavior should be executed. Instead of forcing the policy to directly parse raw images and instructions, a high-level VLM modulates the input representation to the low-level policy by providing: (1) a path that encodes what the policy should do, and (2) masks showing where to attend. By “absorbing” semantic and visual variation, the VLM provides the policy a simplified, annotated “peek” of the scene that gives the what and the where, while the policy only needs to learn how to perform the low-level actions. This intermediate representation helps policy exec","cbCaiaUeCi9Lpa08","https://ap.wps.com/l/cbCaiaUeCi9Lpa08","pdf",5393338,2,1,11,"English","en",105,"# Introduction\n## Core idea: VLM-modulated minimal representations\n## PEEK method and intermediate representation\n## Scalable automatic annotation pipeline\n## Real-world evaluation results","[{\"question\":\"What problem does PEEK address in robot manipulation policies?\",\"answer\":\"It addresses poor zero-shot generalization caused by entangling where-to-attend, what-to-do, and how-to-execute within a single policy.\"},{\"question\":\"How does PEEK use vision-language models with the low-level policy?\",\"answer\":\"A high-level VLM modulates the policy’s input by providing an action-guiding path and attention masks, letting the policy learn primarily how to act.\"},{\"question\":\"What intermediate representation does PEEK predict and how is it used?\",\"answer\":\"PEEK fine-tunes the VLM to predict point-based end-effector paths and task-relevant masking points; these are overlaid on robot observations to make the representation policy-agnostic and transferable.\"}]","PEEK - Policy-agnostic Extraction of Essential Keypoints | PDF",1787218812,28,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"peek-policy-agnostic-extraction-of-essential-keypoints","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,48,51],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":20},"https://docshare.wps.com/document/","Document",{"item":49,"name":12,"@type":44,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":44,"position":53},"https://docshare.wps.com/document/peek-policy-agnostic-extraction-of-essential-keypoints/133381/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-23","2026-08-20",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does PEEK address in robot manipulation policies?","Question",{"text":76,"@type":77},"It addresses poor zero-shot generalization caused by entangling where-to-attend, what-to-do, and how-to-execute within a single policy.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does PEEK use vision-language models with the low-level policy?",{"text":81,"@type":77},"A high-level VLM modulates the policy’s input by providing an action-guiding path and attention masks, letting the policy learn primarily how to act.",{"name":83,"@type":74,"acceptedAnswer":84},"What intermediate representation does PEEK predict and how is it used?",{"text":85,"@type":77},"PEEK fine-tunes the VLM to predict point-based end-effector paths and task-relevant masking points; these are overlaid on robot observations to make the representation policy-agnostic and transferable.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]