[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82891-en":3,"doc-seo-82891-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82891,8796095462418,"Noah","https://ap-avatar.wpscdn.com/avatar/80000253c1241d02b47?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778826106357471780",8,"Research & Report","Unsupervised Pixel-Level Semantic Left-Right Understanding of In-the-Wild Images","Pixel-level semantic left-right prediction in in-the-wild images remains difficult despite progress in reflective symmetry understanding for 3D data, due to missing 3D cues in single-view settings and the presence of occlusion, pose variation, partiality, and boundary complexity. This work presents an unsupervised framework that jointly leverages 3D shape data and image datasets to infer dense pixel-wise semantic left-right labels. A medium-scale 3D dataset with human- and quadruped-like shapes, combined with diverse images, enables high-quality predictions even for entirely unseen object categories, outperforming prior state-of-the-art methods on both rendered and real-world datasets.","arXiv :2607 .05006v 1 [ cs .CV] 6 Jul 2026  \nUnsupervised Pixel-Level Semantic Left-Right Understanding of In-the-Wild Images  \nWeikang Wang 1 ,2 , Tobias Weißberg 1 ,2 , and Florian Bernard 1 ,2  \n1 University of Bonn, Germany  \n2 Lamarr Institute, Germany  \nPartiality Unbalanced left-right Occlusion Complex boundaries  \nDifferent scales Unseen categories Intra-class consistency Inter-class consistency  \nFig. 1: We propose the first unsupervised pixel-level semantic left-right prediction framework for in-the-wild images, which is robust across diverse challenging settings.  \nAbstract. While various works address reflective symmetry understanding in 3D data and images, pixel-level semantic left-right prediction of in-the-wild images remains challenging, due to certain difficulties including the lack of 3D information, occlusion, object pose variation, partiality, etc. In this work, we propose an unsupervised learning framework to tackle this challenge. Leveraging recent advances in vertex-wise semantic left-right understanding of 3D data, our unsupervised learning method jointly utilises 3D shape and image datasets to infer pixel-wise semantic left-right predictions in single-view images. In particular, we show that a medium-scale 3D shape dataset comprising mainly of human-and quadruped animal-like shapes, combined with diverse in-the-wild image data, are sufficient to achieve high-quality semantic left-right prediction in images, even for entirely unseen 3D object categories, such as cars or trains. Overall, our approach achieves superior performance in dense pixel-wise semantic left-right predictions on both rendered and in-thewild image datasets when compared to existing state-of-the-art methods.  \nKeywords: Left-Right Symmetry · Semantic Image Understanding  \n2 Weikang Wang et al.  \n1 Introduction  \nReflective symmetry understanding, especially left-right symmetry understanding, has been a long standing topic that is studied in different areas of visual computing. The high relevance of left-right symmetry stems from the fact that it is a ubiquitous property observed in various object categories, including humans, animals and man-made objects such as cars, bicycles, or aeroplanes.  \nFrom a geometric perspective, reflective symmetries (including left-right symmetry) are conventionally classified as extrinsic or intrinsic ones [15, 19] . For 3D data, such as point clouds and meshes, numerous methods capable of detecting extrinsic [1, 9, 13, 31] or intrinsic reflective symmetries [10, 14, 19, 22, 23, 30] have demonstrated high-quality results.  \nIn contrast, left-right understanding in 2D images is significantly more challenging. Left-right symmetry observed in images arises from the underlying 3D structure of the object they depict, and common difficulties including the lack of 3D information in single-view images, occlusion, pose variation and partiality, make geometric modelling of symmetries in 2D images ill-posed in many settings.  \nExisting methods attempt to address this challenge, but they possess certain limitations. One category of methods focuses on general symmetry axis detection [26,27,33,34], yet they are restricted to objects exhibiting (near) perfect extrinsic symmetry and struggle to handle cases like occlusion or partiality. Another line of works [8,29,35] utilises semantic correspondence as a proxy to refine left-right aware image features; however, keypoints are required as supervision signals.  \nRecent works [30,32] introduce a new semantic perspective to tackle left-right understanding problem. Emerging studies [3, 6, 16, 35] show vision foundation models encode rich left-right semantics. Leveraging this, [30, 32] formalise leftright understanding of 3D shapes as a vertex-wise semantic prediction task. By utilising shape descriptors decorated with features from vision foundation models [4], these methods extract reliable predictions without requiring annotated data, effectively introducing a semantic par","cbCaip7hqLiDwwnE","https://ap.wps.com/l/cbCaip7hqLiDwwnE","pdf",33639258,3,1,17,"English","en",105,"# Introduction\n# Related Works","[{\"question\":\"Why is pixel-level semantic left-right understanding difficult in in-the-wild 2D images?\",\"answer\":\"Single-view images lack explicit 3D information, and factors such as occlusion, pose variation, partiality, and complex boundaries make geometric modeling ill-posed in many settings.\"},{\"question\":\"What does the proposed unsupervised framework do for left-right prediction?\",\"answer\":\"It jointly uses a 3D shape dataset and in-the-wild image data to infer dense, pixel-wise semantic left-right predictions without requiring annotated left-right labels.\"},{\"question\":\"How can the method generalize to unseen 3D object categories?\",\"answer\":\"A medium-scale 3D shape dataset covering human- and quadruped-like shapes, when combined with diverse image datasets, is sufficient to produce high-quality semantic left-right predictions even for unseen categories such as cars or trains.\"}]",1784183732,43,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"unsupervised-pixel-level-semantic-left-right-understanding-of-in-the-wild-images","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/unsupervised-pixel-level-semantic-left-right-understanding-of-in-the-wild-images/82891/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is pixel-level semantic left-right understanding difficult in in-the-wild 2D images?","Question",{"text":75,"@type":76},"Single-view images lack explicit 3D information, and factors such as occlusion, pose variation, partiality, and complex boundaries make geometric modeling ill-posed in many settings.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What does the proposed unsupervised framework do for left-right prediction?",{"text":80,"@type":76},"It jointly uses a 3D shape dataset and in-the-wild image data to infer dense, pixel-wise semantic left-right predictions without requiring annotated left-right labels.",{"name":82,"@type":73,"acceptedAnswer":83},"How can the method generalize to unseen 3D object categories?",{"text":84,"@type":76},"A medium-scale 3D shape dataset covering human- and quadruped-like shapes, when combined with diverse image datasets, is sufficient to produce high-quality semantic left-right predictions even for unseen categories such as cars or trains.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]