[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86172-en":3,"doc-seo-86172-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86172,962075114101,"Seraphina","https://ap-avatar.wpscdn.com/avatar/e000253a75eb197efd?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780044092746381165",8,"Research & Report","Pix2Act: Image-Space Manipulation Policies with Equivariant Augmentation","Pix2Act represents robot manipulation actions as continuous keypoint trajectories in each camera’s 2D image plane, enabling compact and interpretable learning of complex 3D control. The method avoids limitations of pixel discretization and out-of-frame uncertainty by shifting action representation from SE(3) paths to unbounded R2 trajectories, then recovering end-effector poses losslessly via triangulation. It further introduces per-camera rotation equivariant augmentation, training view-fused networks that respect independent image-space rotations, improving robustness and generalization across diverse simulated and real tasks.","arXiv :2607 . 11167v1 [ cs .RO] 13 Jul 2026  \nPix2Act: Image-Space Manipulation Policies with Equivariant Augmentation  \nHaojie Huang 1 ,∗ Linfeng Zhao2 Haotian Liu 1 Zhang Ye 1  \nSi-Yuan Huang3 Mingxi Jia4 Boce Hu 1 Fangzhou Lin5 Yu Qi 1  \nDian Wang2 Robin Walters 1 ,† Robert Platt 1 ,†  \n1Northeastern University 2 Stanford University 3University of Pennsylvania  \n4Brown University 5Texas A&M University ∗ Corresponding author †Equal advising [haojhuang.github.io/pix2act](haojhuang.github.io/pix2act) page [huang.haoj@northeastern.edu](huang.haoj@northeastern.edu)  \nAbstract: Representing manipulation actions as 2D trajectories in the camera plane provides a compact and interpretable basis for learning complex 3D manipulation policies. However, it also creates challenges from out-of-frame trajectories and limited precision. We propose Pix2Act, an imitation learning method that addresses these challenges by generating continuous image-space keypoint trajectories in each camera plane and losslessly recovering end-effector poses via triangulation.  \nThis reformulates high-dimensional 3D control as a simpler, more learnable 2D prediction problem. Crucially, it aligns observations and actions in the same coordinate space, enabling equivariant transformations to jointly rotate individual camera images together with their image-space actions. We analyze the symmetry properties of this augmentation and design a network architecture that can fuse multiple camera views while respecting their per-view rotations. As a result, Pix2Act implicitly enlarges the support of the data distribution and learns invariant action structures across transformations, yielding improved generalization and overall performance. Across diverse simulated and real-world manipulation tasks, Pix2Act outperforms state-of-the-art baselines and remains robust under camera perturbations.  \nKeywords: Manipulation Learning, Imitation Learning, Action Representation  \n1 Introduction  \nManipulation policy learning has seen significant advances, particularly in mimicking 3D trajectories for closed-loop control. Typically, prior work models a conditional mapping from 2D visual observations to a 3D action chunk, defined as a sequence of 3D translations and rotations. However, this formulation often treats the image space and the 3D Cartesian action space as independent domains, overlooking the inherent geometric relationship between them. This unstructured mapping lacks explicit spatial grounding and forces the network to implicitly infer complex 3D trajectories between the end-effector and objects, making the learning problem highly ambiguous and prone to overfitting. Projecting manipulation actions as 2D trajectories in the camera plane offers an alternative: complex 3D policies can be learned through compact and interpretable 2D representations. However, this is non-trivial due to precision limits and the difficulty of cross-view fusion. For example, prior work [1] denoises pixel keypoint coordinates independently in each view and triangulates them across two agent cameras to recover 3D positions. As a result, it suffers from precision loss due to pixel discretization, an inability to represent actions outside the image frame, large triangulation errors under widely separated viewpoints, and cross-view tracking inconsistencies.  \nIn this work, we provide the policy with explicit spatial grounding by shifting the action representation from trajectories in SE(3) to a set of continuous and unbounded paths in the image plane, i.e. in R2. This formulation eliminates the pixel discretization that limits prior 2D approaches, recovering full precision in the action representation. Figure 1 illustrates the key elements of our approach. First, we define a set of keypoints on the gripper (lower left of Figure 1a) . We encode a trajectory through SE(3) (the action chunk) with trajectories of these keypoints. Second, we project these keypoints  \nFigure 1: From left to right: (a) an il","cbCaihwhno4pcs5f","https://ap.wps.com/l/cbCaihwhno4pcs5f","pdf",50862534,4,1,20,"English","en",105,"# Introduction\n## Action representation in SE(3) vs image-space\n## Dual in-hand cameras and keypoint trajectories\n## Projection, diffusion-based trajectory generation, and triangulation\n## Per-view equivariant augmentation","[{\"question\":\"What problem does Pix2Act address in learning 3D manipulation policies from images?\",\"answer\":\"Prior approaches map 2D observations to 3D action chunks without explicit geometric grounding, making the mapping ambiguous and prone to precision loss and overfitting.\"},{\"question\":\"How does Pix2Act represent actions differently from existing methods?\",\"answer\":\"Pix2Act shifts from SE(3) trajectory prediction to continuous, unbounded paths of keypoints in the 2D image plane (R2), then reconstructs full 3D end-effector poses through triangulation.\"},{\"question\":\"What is equivariant augmentation in Pix2Act and why is it important?\",\"answer\":\"Because actions and observations share the same 2D image space, Pix2Act applies per-camera image rotations so the generated image-space action trajectories transform consistently. This n-equivariance (with n cameras) biases the model toward local visual features and improves robustness under camera perturbations.\"}]",1784209097,50,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"pix2act-image-space-manipulation-policies-with-equivariant-augmentation","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/pix2act-image-space-manipulation-policies-with-equivariant-augmentation/86172/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does Pix2Act address in learning 3D manipulation policies from images?","Question",{"text":75,"@type":76},"Prior approaches map 2D observations to 3D action chunks without explicit geometric grounding, making the mapping ambiguous and prone to precision loss and overfitting.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does Pix2Act represent actions differently from existing methods?",{"text":80,"@type":76},"Pix2Act shifts from SE(3) trajectory prediction to continuous, unbounded paths of keypoints in the 2D image plane (R2), then reconstructs full 3D end-effector poses through triangulation.",{"name":82,"@type":73,"acceptedAnswer":83},"What is equivariant augmentation in Pix2Act and why is it important?",{"text":84,"@type":76},"Because actions and observations share the same 2D image space, Pix2Act applies per-camera image rotations so the generated image-space action trajectories transform consistently. This n-equivariance (with n cameras) biases the model toward local visual features and improves robustness under camera perturbations.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":29,"slug":113},6,"Technology","technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":22,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":127,"show_sort_weight":22,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":46,"category_name":135,"show_sort_weight":106,"slug":136},19,"General","general"]