[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82281-en":3,"doc-seo-82281-105":29,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82281,13056703019662,"Evangeline","https://ap-avatar.wpscdn.com/avatar/be000253a8e92610077?_k=1778726343310543188",8,"Research & Report","REMIND Re-Identification with Memory for Indoor Navigation","Mobile robots navigating indoors must re-identify objects after long temporal gaps, large viewpoint shifts, and major illumination changes, yet existing tracking and re-identification pipelines are optimized for short-term video association or lack persistent identity memory and global consistency. REMIND provides an online monocular RGB tracker without camera pose or depth, combining frozen DINOv3 features with a dual-bank multi-prototype appearance memory, context reasoning, and ambiguity-aware assignment. Experiments on dedicated indoor revisits and ScanNet++ show leading IDF1, while the full system and dataset are publicly released.","REMIND: RE-Identification with Memory for  \nINDoor Navigation  \nPablo Diaz-Pereda, Alejandro Rodriguez-Ramos, David Perez-Saura, and Pascual Campoy  \narXiv :2607 .09267v1 [ cs .CV] 10 Jul 2026  \nAbstract—Mobile robots operating indoors must re-identify previously observed objects after long temporal gaps, significant viewpoint changes, and severe illumination variations. This remains a challenging problem: multi-object tracking methods are optimized for short-term association of pedestrians and vehicles at video rates, person and vehicle re-identification approaches lack persistent memory mechanisms, and state-of-the-art video object segmentation techniques rely on reactive distractor filtering rather than enforcing global identity consistency.  \nTo address these limitations, we present REMIND, an online tracker designed for long-term multi-object re-identification of generic indoor objects from monocular RGB imagery, requiring neither camera pose nor depth. Motivated by evidence from visual cognition that humans rely on accumulated appearance familiarity and spatial context rather than explicit self-localization, REMIND combines frozen DINOv3 features with a dual-bank multi-prototype appearance memory, part-and background-level descriptors, a neighbour-context reasoning module exploiting spatial co-occurrence, and joint Hungarian assignment with ambiguity-aware safeguards. On a purpose-built indoor dataset featuring controlled revisits and dense same-class clutter, REMIND reaches 90.35% IDF1, nearly 20 points above a stateof-the-art video object segmentation baseline and more than 36 above a strong tracking-by-detection baseline. On ScanNet++, it attains the highest IDF1 in every setting but one, end-to-end detection over all scenes, where the tracking-by-detection baseline is marginally ahead while REMIND still associates and recovers identities more accurately; it also completes every scene, whereas the video object segmentation baseline exhausts GPU memory on 66.9% under YOLO detections. The complete system, evaluation framework, and dataset are publicly released.  \nI. INTRODUCTION  \nHumans recognise previously seen objects with remarkable ease, even after long absences and under substantially different viewing conditions. Returning to a room after days, a person re-identifies familiar items despite changes in viewpoint, illumination, and partial occlusion, without any conscious estimation of their own position in space. Research in visual cognition has shown that this ability relies heavily on accumulated appearance familiarity and on the spatial relationships between co-occurring objects [1], rather than on explicit geometric reconstruction of the scene. The robustness of this process, and its independence from self-localisation, points to a computational principle directly relevant to autonomous systems that must maintain object identity over time.  \nMobile robots navigating indoor environments face precisely this challenge. A robot leaves a room, traverses a  \nAll authors are with the Computer Vision and Aerial Robotics Group at Centre for Automation and Robotics C.A.R. (UPM-CSIC), Universidad Politcnica de Madrid (CVAR-UPM), Calle Jose Gutierrez Abascal 2, 28006 Madrid, Spain.  \nFig. 1: Long-term object re-identification scenario for indoor robot navigation. (1) First visit: the robot enters Room A, perceives the scene, and assigns persistent identities ID 1 ,..., ID5 to the visible objects. (2) Leaving and exploring: the robot traverses other rooms for several minutes, during which the previously observed objects are out of view. (3) Re-entry: upon returning to Room A, every previously catalogued object must be re-identified despite substantial changes in viewpoint, illumination, and partial occlusion, and without relying on camera pose or scene geometry.  \ncorridor, and returns minutes later; it must re-identify the same objects under substantial viewpoint and illumination changes (Fig. 1) . Unlike frame-to-frame tr","cbCaiu0DyHuRAcSf","https://ap.wps.com/l/cbCaiu0DyHuRAcSf","pdf",1712673,1,12,"English","en",105,"# Introduction\n## Long-term re-identification challenge\n## Cognitive motivation and system requirements\n## Limitations of existing paradigms","[{\"question\":\"What evidence supports REMIND’s design choice of using appearance and spatial context?\",\"answer\":\"The document motivates the approach using visual cognition findings that humans rely on accumulated appearance familiarity and spatial relationships between co-occurring objects rather than explicit self-localization, aligning with REMIND’s monocular appearance-and-context pipeline.\"}]",1784179366,30,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":27},"remind-re-identification-with-memory-for-indoor-navigation","",{"@graph":35,"@context":77},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/remind-re-identification-with-memory-for-indoor-navigation/82281/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"What evidence supports REMIND’s design choice of using appearance and spatial context?","Question",{"text":75,"@type":76},"The document motivates the approach using visual cognition findings that humans rely on accumulated appearance familiarity and spatial relationships between co-occurring objects rather than explicit self-localization, aligning with REMIND’s monocular appearance-and-context pipeline.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,114,119,122,126],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":45,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":45,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":45,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":28,"slug":113},"research-report",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},9,"Religion & Spirituality",20,"religion-spirituality",{"id":117,"doc_module":4,"doc_module_name":45,"category_name":120,"show_sort_weight":117,"slug":121},"World Cup","world-cup",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":123,"slug":125},10,"Lifestyle","lifestyle",{"id":127,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":98,"slug":129},19,"General","general"]