[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"detail-sidebar-cat-0-en-105":3,"doc-seo-140258-105":59,"doc-detail-140258-en":130},{"code":4,"msg":5,"data":6},0,"success",[7,13,18,23,28,33,38,43,48,51,55],{"id":8,"doc_module":4,"doc_module_name":9,"category_name":10,"show_sort_weight":11,"slug":12},1,"Document","Story & Novel",90,"story-novel",{"id":14,"doc_module":4,"doc_module_name":9,"category_name":15,"show_sort_weight":16,"slug":17},2,"Literature",80,"literature",{"id":19,"doc_module":4,"doc_module_name":9,"category_name":20,"show_sort_weight":21,"slug":22},4,"Exam",70,"exam",{"id":24,"doc_module":4,"doc_module_name":9,"category_name":25,"show_sort_weight":26,"slug":27},5,"Comic",60,"comic",{"id":29,"doc_module":4,"doc_module_name":9,"category_name":30,"show_sort_weight":31,"slug":32},6,"Technology",50,"technology",{"id":34,"doc_module":4,"doc_module_name":9,"category_name":35,"show_sort_weight":36,"slug":37},7,"Healthcare",40,"healthcare",{"id":39,"doc_module":4,"doc_module_name":9,"category_name":40,"show_sort_weight":41,"slug":42},8,"Research & Report",30,"research-report",{"id":44,"doc_module":4,"doc_module_name":9,"category_name":45,"show_sort_weight":46,"slug":47},9,"Religion & Spirituality",20,"religion-spirituality",{"id":46,"doc_module":4,"doc_module_name":9,"category_name":49,"show_sort_weight":46,"slug":50},"World Cup","world-cup",{"id":52,"doc_module":4,"doc_module_name":9,"category_name":53,"show_sort_weight":52,"slug":54},10,"Lifestyle","lifestyle",{"id":56,"doc_module":4,"doc_module_name":9,"category_name":57,"show_sort_weight":24,"slug":58},19,"General","general",{"code":4,"msg":60,"data":61},"ok",{"site_id":62,"language":63,"slug":64,"title":65,"keywords":66,"description":67,"schema_data":68,"social_meta":123,"head_meta":125,"extra_data":127,"updated_unix":129},105,"en","leveraging-equivariant-features-for-absolute-pose-regression-cvpr-2022-paper","Leveraging Equivariant Features for Absolute Pose Regression - CVPR 2022 paper","","Absolute pose regression remains less competitive than 3D geometry-based pose estimation methods, and prior end-to-end approaches often resemble image retrieval rather than leveraging intrinsic geometry. The work argues that standard CNN-learned statistical features lack sufficient geometric information. It proposes a translation- and rotation-equivariant convolutional neural network that maps camera motions directly into the feature space. This geometric equivariance enables implicit training-data augmentation over image-plane-preserving transformations. Extensive experiments validate a lightweight model that outperforms existing methods on standard datasets.",{"@graph":69,"@context":122},[70,84,105],{"@type":71,"itemListElement":72},"BreadcrumbList",[73,77,79,82],{"item":74,"name":75,"@type":76,"position":8},"https://docshare.wps.com","Home","ListItem",{"item":78,"name":9,"@type":76,"position":14},"https://docshare.wps.com/document/",{"item":80,"name":40,"@type":76,"position":81},"https://docshare.wps.com/document/research-report/",3,{"item":83,"name":65,"@type":76,"position":19},"https://docshare.wps.com/document/leveraging-equivariant-features-for-absolute-pose-regression-cvpr-2022-paper/140258/",{"url":83,"name":65,"@type":85,"image":86,"author":91,"headline":65,"publisher":94,"fileFormat":97,"inLanguage":63,"description":67,"dateModified":98,"datePublished":99,"encodingFormat":97,"isAccessibleForFree":100,"interactionStatistic":101},"DigitalDocument",{"url":87,"@type":88,"width":89,"height":90},"https://docshare.wps.com/thumbnails/leveraging-equivariant-features-for-absolute-pose-regression-cvpr-2022-paper/140258.png","ImageObject",300,407,{"name":92,"@type":93},"Melati","Person",{"url":74,"name":95,"@type":96},"DocShare","Organization","application/pdf","2026-09-17","2026-08-24",true,{"@type":102,"interactionType":103,"userInteractionCount":34},"InteractionCounter",{"@type":104},"ViewAction",{"@type":106,"mainEntity":107},"FAQPage",[108,114,118],{"name":109,"@type":110,"acceptedAnswer":111},"Why does absolute pose regression underperform compared with 3D geometry-based methods?","Question",{"text":112,"@type":113},"The paper attributes the gap to insufficient geometric information in features learned by classical CNNs, which makes absolute pose regression closer to image retrieval than to 3D structure reasoning.","Answer",{"name":115,"@type":110,"acceptedAnswer":116},"What is the key idea behind the proposed approach?",{"text":117,"@type":113},"A translation- and rotation-equivariant CNN is used so that representations directly encode camera planar motions in the feature space, leveraging equivariance properties for pose regression.",{"name":119,"@type":110,"acceptedAnswer":120},"How does equivariance affect training data?",{"text":121,"@type":113},"Equivariance implicitly augments training data by leveraging a group of image plane-preserving transformations, reducing the reliance on explicit data augmentation.","https://schema.org",{"og:url":83,"og:type":124,"og:title":65,"og:site_name":95,"og:description":67},"article",{"robots":126,"canonical":83},"index,follow",{"doc_id":128,"site_id":62},140258,1787571374,{"code":4,"msg":5,"data":131},{"doc_id":128,"user_id":132,"nickname":92,"user_avatar":133,"doc_module":4,"category_id":39,"category_name":40,"doc_title":65,"doc_description":67,"doc_content":134,"file_id":135,"file_url":136,"file_type":137,"file_size":138,"view_count":34,"is_deleted":4,"is_public":8,"is_downloadable":8,"audit_status":8,"page_count":139,"language":140,"language_code":63,"site_id":62,"html_lang":63,"table_of_contents":141,"faqs":142,"seo_title":143,"seo_description":67,"update_tm":129,"read_time":144},962085570644,"https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d","Leveraging Equivariant Features for Absolute Pose Regression  \nMohamed Adel Musallam Vincent Gaudillire Miguel Ortiz del Castillo  \n[mohamed.ali@uni.lu](mohamed.ali@uni.lu) [vincent.gaudilliere@uni.lu](vincent.gaudilliere@uni.lu) [miguel.ortizdelcastillo@uni.lu](miguel.ortizdelcastillo@uni.lu)  \nKassem Al Ismaeil Djamila Aouada  \n[kassem.alismaeil@gmail.com](kassem.alismaeil@gmail.com) [djamila.aouada@uni.lu](djamila.aouada@uni.lu)  \nInterdisciplinary Center for Security, Reliability and Trust (SnT) University of Luxembourg, Luxembourg  \nAbstract  \nWhile end-to-end approaches have achieved state-ofthe-art performance in many perception tasks, they are not yet able to compete with 3D geometry-based methods in pose estimation. Moreover, absolute pose regression has been shown to be more related to image retrieval. As a result, we hypothesize that the statistical features learned by classical Convolutional Neural Networks do not carry enough geometric information to reliably solve this inherently geometric task. In this paper, we demonstrate how a translation and rotation equivariant Convolutional Neural Network directly induces representations of camera motions into the feature space. We then show that this geometric property allows for implicitly augmenting the training data under a whole group of image plane-preserving transformations. Therefore, we argue that directly learning equivariant features is preferable than learning data-intensive intermediate representations. Comprehensive experimental validation demonstrates that our lightweight model outperforms existing ones on standard datasets.1  \n1. Introduction  \nIn computer vision, camera pose estimation, and its reference frame inverse, i.e., object pose estimation, have been extensively studied over the last decades [38, 42, 54] .  \nTraditionally, pose estimation has been addressed using 3D geometry. In practice, a set of 2D-3D feature correspondences is generated, then statistically leveraged to recover the camera pose [18, 39, 49, 64] . More recently, direct Absolute Pose Regression (APR) approaches have been introduced, drawing upon early successes of deep learning [1] .  \n1This work was funded by the Luxembourg National Research Fund (FNR), under the project reference BRIDGES2020/IS/14755859/MEETA/Aouada, and by LMO ([https://www.lmo.space](https://www.lmo.space)).  \nFigure 1 . Illustration of our approach - Our method adopts a translation and rotation-equivariant convolutional neural network to extract geometry-aware features that directly encode camera planar motions R, t. While camera moves, equivariance of the proposed feature extractor F leads to explicit image (ϕR(I,)t) and feature (ϕR(F,t) ) changes. This property is leveraged to propose amore efficient solution to the absolute pose regression problem.  \nThese methods consist in directly mapping an image to its pose using a suitably trained Convolutional Neural Network (CNN) . Therefore, end-to-end trainable methods have the advantage of providing fully differentiable results, enabling the optimization of all parameters in a comprehensive manner. Moreover, predictions are achieved at a steady speed and power consumption, whereas RANdom SAmple Consensus (RANSAC)-based methods [18] are less predictable  \nand likely to suffer from an efficiency drop when the inlier rate is low. However, state-of-the-art APR methods have been proven theoretically and shown experimentally to have a lower accuracy compared to 3D structure-based approaches [50] . Indeed, the former are more closely related to image retrieval than to 3D structure [50] .  \nThe questions we ask in this work are: Why do current APR methods fall short in accuracy ? How can they reach their full potential? Our hypothesis is that there is a lack of exploitation of the geometric properties of data. This happens typically at the level of the feature extraction layers commonly used in classical deep learning approaches. Specifically, we posit that in the case of","cbCaiu1R7WOCkrOR","https://ap.wps.com/l/cbCaiu1R7WOCkrOR","pdf",8593398,11,"English","# Abstract\n# Introduction\n## Motivation and background\n## Contributions\n# Method Overview","[{\"question\":\"Why does absolute pose regression underperform compared with 3D geometry-based methods?\",\"answer\":\"The paper attributes the gap to insufficient geometric information in features learned by classical CNNs, which makes absolute pose regression closer to image retrieval than to 3D structure reasoning.\"},{\"question\":\"What is the key idea behind the proposed approach?\",\"answer\":\"A translation- and rotation-equivariant CNN is used so that representations directly encode camera planar motions in the feature space, leveraging equivariance properties for pose regression.\"},{\"question\":\"How does equivariance affect training data?\",\"answer\":\"Equivariance implicitly augments training data by leveraging a group of image plane-preserving transformations, reducing the reliance on explicit data augmentation.\"}]","Leveraging Equivariant Features for Absolute Pose Regression - CVPR 2022 paper | PDF",28]