[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85289-en":3,"doc-seo-85289-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},85289,687197207919,"Theodora","https://ap-avatar.wpscdn.com/avatar/a000253d6f5f7c60be?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779446848396160552",8,"Research & Report","Desc++ Efficient Descriptor Enhancement for Data Association in Existing Visual SLAM Systems","Reliable visual data association is fundamental to visual SLAM (V-SLAM), directly shaping camera pose estimation quality and map consistency. Existing real-time systems rely on handcrafted descriptors that degrade under illumination and viewpoint changes, while learning-based front-ends often require replacing the pipeline and add overhead. The proposed Desc++ enhances descriptors within the original format, jointly encoding descriptor representations and keypoint geometry and aggregating spatial context via a hybrid linear-time design. Experiments across multiple tasks and four V-SLAM systems show improved accuracy and trajectory stability with a practical efficiency–accuracy balance, enabling drop-in integration.","Desc++: Efficient Descriptor Enhancement for Data Association in Existing Visual SLAM Systems  \nTing-Wei Ou 1 , Huang-Ting Lin2 , and Kuu-Young Young2  \narXiv :2607 . 11099v1 [ cs .RO] 13 Jul 2026  \nAbstract—Reliable visual data association is fundamental to visual SLAM (V-SLAM), as it directly determines the quality of the camera pose estimation and map consistency. However, the handcrafted descriptors used by most mature real-time systems degrade under illumination and viewpoint changes, while learning-based front-ends that address this weakness typically require replacing the extraction-and-matching pipeline and introduce substantial computational overhead. Descriptor enhancement offers a compromise by refining existing descriptors within their original format, yet current methods rely on simplified attention mechanisms whose limited contextual modeling constrains the achievable matching quality. To resolve this trade-off between contextual expressiveness and efficiency, we propose Desc++, a lightweight enhancement module that jointly encodes descriptor representations and keypoint geometry and aggregates spatial context through a hybrid architecture that combines order-agnostic global attention with geometry-aware sequential modeling in linear time. The enhanced descriptors retain their original dimensionality and matching interface, enabling integration into deployed V-SLAM systems without modifying the pipeline. Experiments across descriptor matching, correspondence analysis, and system-level benchmarks with four different V-SLAM systems demonstrate that Desc++ improves matching accuracy over the state-ofthe-art enhancement method, translates these gains into more accurate and stable trajectory estimation, and achieves a favorable balance between accuracy and efficiency for practical integration into existing real-time V-SLAM pipelines. The source code and pretrained model weights are available at: [https://github.com/ouotingwei/DescPP.git](https://github.com/ouotingwei/DescPP.git).  \nI. INTRODUCTION  \nAutonomous systems operating in unknown, GPS-denied environments require continuous state estimation from onboard sensing, a capability for which visual SLAM (VSLAM) has become a core solution in robotics [45] . The performance of a V-SLAM system largely depends on reliable visual data association, which establishes feature correspondences across images for camera tracking, pose optimization, and map construction. In sparse feature-based pipelines, these correspondences are established by matching local feature descriptors extracted around detected keypoints. Consequently, descriptor discriminability directly affects matching accuracy and overall SLAM robustness. To meet the computational requirements of real-time applications, most mature V-SLAM systems [2]–[4], [41] rely on handcrafted descriptors, which are computationally efficient but remain sensitive to illumination changes and viewpoint variations, often leading to degraded data association and accumulated trajectory drift.  \nTo improve visual data association, recent studies have explored learning-based local feature extraction [8]–[10] and feature matching [12]–[14] . Although these approaches  \nsubstantially improve matching robustness and accuracy, integrating them into existing V-SLAM systems generally requires replacing the original front-end pipeline, limiting their adoption in deployed real-time systems. Descriptor enhancement [7] provides an alternative design strategy by improving the discriminative capability of existing feature descriptors while preserving the original detector, descriptor format, and matching interface. Such a plug-and-play design enables deployed V-SLAM systems to benefit from learned representations with minimal system modification. Contextual relationships among neighboring keypoints provide complementary geometric and structural cues that are beneficial for descriptor discrimination. However, the state-of-the-art enhancement method ","cbCaij4C2OMWnMSp","https://ap.wps.com/l/cbCaij4C2OMWnMSp","pdf",6027478,1,12,"English","en",105,"# I. Introduction\n## Visual data association in V-SLAM\n## Limitations of handcrafted descriptors\n## Learning-based front-ends and integration cost\n## Descriptor enhancement as a plug-and-play alternative\n## Proposed Desc++ framework and contributions","[{\"question\":\"What problem does Desc++ address in visual SLAM?\",\"answer\":\"Desc++ targets unreliable visual data association caused by descriptor degradation under illumination and viewpoint changes, which leads to poorer pose estimation and accumulated drift.\"},{\"question\":\"How does Desc++ improve descriptors without changing the existing V-SLAM pipeline?\",\"answer\":\"It enhances the original descriptors in-place by retaining their dimensionality and matching interface, so the same downstream tracking and matching logic can be used.\"},{\"question\":\"Why is Desc++ described as efficient while still capturing context?\",\"answer\":\"Desc++ aggregates spatial context using a hybrid architecture: order-agnostic global attention without attention-heavy cost and geometry-aware sequential modeling along a serialized keypoint sequence in linear time.\"}]",1784202277,30,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"desc-efficient-descriptor-enhancement-for-data-association-in-existing-visual-slam-systems","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/desc-efficient-descriptor-enhancement-for-data-association-in-existing-visual-slam-systems/85289/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does Desc++ address in visual SLAM?","Question",{"text":75,"@type":76},"Desc++ targets unreliable visual data association caused by descriptor degradation under illumination and viewpoint changes, which leads to poorer pose estimation and accumulated drift.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does Desc++ improve descriptors without changing the existing V-SLAM pipeline?",{"text":80,"@type":76},"It enhances the original descriptors in-place by retaining their dimensionality and matching interface, so the same downstream tracking and matching logic can be used.",{"name":82,"@type":73,"acceptedAnswer":83},"Why is Desc++ described as efficient while still capturing context?",{"text":84,"@type":76},"Desc++ aggregates spatial context using a hybrid architecture: order-agnostic global attention without attention-heavy cost and geometry-aware sequential modeling along a serialized keypoint sequence in linear time.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":28,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]