[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84565-en":3,"doc-seo-84565-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84565,34359740700684,"Finn","https://ap-avatar.wpscdn.com/avatar/1f400023980c374ae676?_k=1777273430885731487",8,"Research & Report","LIST3R Long Sequence Instance Aware 3D Reconstruction","LIST3R introduces an instance-aware framework for long-sequence 3D reconstruction, inspired by how humans organize spatial memory around stable, recognizable objects. The method partitions a long video into overlapping subsequences and builds structured local instance libraries for each partial reconstruction. Persistent instance anchors with semantic and geometric evidence are matched across subsequences to recover revisited regions and enforce object-aware alignment. Evolving geometric evidence updates the libraries and progressively consolidates them into a unified global 3D instance library, improving trajectory accuracy and reconstruction quality on long-horizon benchmarks.","arXiv :2607 .00375v 1 [ cs .CV] 1 Jul 2026  \nLIST3R: Long-sequence Instance-aware 3D Reconstruction  \nJing Gao1, Wei Wang1,,† Feiran Wang2, Yan Yan2  \n1Beijing Jiaotong University 2University of Illinois Chicago  \nMissed Revisit  \nRecoverd  \nRevisit  \nSmooth Alignment  \nGT  \nVGGT-Long  \nOurs  \nRigid Alignment  \nVGGT-Long Camera Pose Ours  \nFigure 1: LIST3R leverages instance guidance to recover more effective revisits and smoother crosssubsequence alignment, producing more accurate and stable camera trajectories than the baseline.  \nAbstract  \nWe present LIST3R, an instance-aware framework for long-sequence 3D reconstruction inspired by the way humans organize spatial memory around stable and recognizable objects. LIST3R organizes long-sequence reconstruction around instance anchors, using them to reconnect fragmented subsequences and consolidate local observations into a coherent global 3D scene. Given a long video, our approach partitions it into overlapping subsequences and builds a structured local instance library for each partial reconstruction, maintaining persistent trackable anchors with semantic and geometric evidence. These anchors are matched across subsequences to recover revisited regions and provide object-aware constraints for fragment alignment, producing a consistent global reconstruction. During this process, the evolving geometric evidence updates the local instance libraries and progressively organizes them into a unified global 3D instance library. Experiments on long-sequence benchmarks show that our method produces more accurate trajectories and higher-quality 3D reconstructions, highlighting the effectiveness of persistent instance anchors for organizing long-horizon 3D reconstruction. Our code is available on the project page: [https://yixn965.github.io/LIST3R/](https://yixn965.github.io/LIST3R/) .  \n1 Introduction  \nHumans organize long visual experiences around a small number of stable and recognizable objects. When navigating through an environment, these persistent cues help structure spatial memory, infer  \nlocation, and reconnect observations made at different times [13, 10, 3, 21] . Even under viewpoint †Corresponding Author.  \nchanges, occlusion, or partial visibility, they remain effective for recognizing revisited places and maintaining a coherent understanding of the surrounding space [9, 7, 8] . Such an ability is particularly important for long-horizon visual understanding, where the challenge is not only to interpret the current observation, but also to consistently relate it to observations from the distant past.  \nBuilding on these insights, long-sequence 3D reconstruction calls for a framework that unifies three key capabilities: 1) maintaining persistent anchors that remain usable across time, viewpoint changes, and fragmented observations; 2) re-identifying revisited places over long temporal horizons; and 3) integrating partial reconstructions into a globally consistent representation. These capabilities are essential for extending reconstruction beyond strong local geometry, enabling the system to preserve reliable long-range associations and consistent scene structure throughout the full sequence.  \nRecent work on long-sequence 3D reconstruction mainly follows two directions. One line of work adopts streaming formulations [23, 5], where frames are processed sequentially while a compact latent memory is updated over time. While such methods enable long-horizon processing, earlier scene information may be gradually forgotten as the sequence grows, making it difficult to preserve reliable long-range associations over extended videos. The other line partitions a long video into overlapping subsequences, reconstructs each subsequence independently with a strong feed-forward foundation model, and then merges the partial results into a unified scene [6, 26] . By retaining the full reconstruction power of the base model [22, 25, 11] within each local window, this paradigm often yields hi","cbCaiqrsdV7Xhl3B","https://ap.wps.com/l/cbCaiqrsdV7Xhl3B","pdf",18478553,1,23,"English","en",105,"# Introduction\n## Motivation from spatial memory\n## Key capabilities for long-horizon reconstruction\n## Related work and limitations\n## Proposed approach: LIST3R","[{\"question\":\"What problem does LIST3R address in long-sequence 3D reconstruction?\",\"answer\":\"LIST3R targets the difficulty of maintaining reliable long-range associations across time, viewpoint changes, and fragmented observations when reconstructing a long video into a coherent 3D scene.\"},{\"question\":\"How does LIST3R use instance anchors during reconstruction?\",\"answer\":\"LIST3R partitions the video into overlapping subsequences and constructs local instance libraries from trackable instance observations. Instance anchors are matched across subsequences to reconnect revisited regions and provide object-aware constraints for alignment.\"},{\"question\":\"What improves global consistency in LIST3R?\",\"answer\":\"LIST3R jointly merges revisited and neighboring subsequences using instance-guided constraints, then refines alignment with confidence-weighted optimization. As geometry evidence evolves, local instances are consolidated into a unified global 3D instance library.\"}]",1784196842,58,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"list3r-long-sequence-instance-aware-3d-reconstruction","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/list3r-long-sequence-instance-aware-3d-reconstruction/84565/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does LIST3R address in long-sequence 3D reconstruction?","Question",{"text":75,"@type":76},"LIST3R targets the difficulty of maintaining reliable long-range associations across time, viewpoint changes, and fragmented observations when reconstructing a long video into a coherent 3D scene.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does LIST3R use instance anchors during reconstruction?",{"text":80,"@type":76},"LIST3R partitions the video into overlapping subsequences and constructs local instance libraries from trackable instance observations. Instance anchors are matched across subsequences to reconnect revisited regions and provide object-aware constraints for alignment.",{"name":82,"@type":73,"acceptedAnswer":83},"What improves global consistency in LIST3R?",{"text":84,"@type":76},"LIST3R jointly merges revisited and neighboring subsequences using instance-guided constraints, then refines alignment with confidence-weighted optimization. As geometry evidence evolves, local instances are consolidated into a unified global 3D instance library.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]