[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82031-en":3,"doc-seo-82031-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},82031,7971461740886,"Theodore","https://ap-avatar.wpscdn.com/davatar_3d24733baf745e90a7e4bdd5f77d97b2",8,"Research & Report","Whareformer: Learning to Track What is Where in Long Egocentric Videos","The paper addresses the “Out of Sight, Not out of Mind” (OSNOM) egocentric-video task, which requires tracking moved objects while preserving both instance locations and identities when objects leave the field of view or become heavily occluded. It introduces Whareformer, a transformer-based method with an updatable memory of tracks and a feed-forward track assignment module. The model jointly reasons over evolving appearance (“what”) and updated 3D location (“where”), using a dedicated New Track token for novel objects. Trained on 56 videos, it achieves state-of-the-art results on 260 long test videos across EPIC-KITCHENS-100, IT3DEgo, and HD-EPIC.","arXiv :2607 .08537v 1 [ cs .CV] 9 Jul 2026  \nWhareformer: Learning to Track What is Wherein Long Egocentric Videos  \nJacob Chalk 1 Saptarshi Sinha 1 Dima Damen 1 Yannis Kalantidis2 Diane Larlus2  \n1 University of Bristol, UK  \n2 NAVER LABS Europe, France  \n[https://jacobchalk.github.io/Whareformer/](https://jacobchalk.github.io/Whareformer/)  \nAbstract. The recently established ‘Out of Sight, Not out of Mind’(OSNOM) task for egocentric videos focuses on tracking objects that are moved by the camera wearer, online, maintaining knowledge of instance locations throughout the video even when they leave the field of view or become heavily occluded. In this paper, we propose the first learningbased solution to the OSNOM task: Whareformer, a transformer-based  \nmodel with two components: an updatable memory of established tracks and a track assignment module that associates observations with existing tracks in a feed-forward manner. Whareformer jointly reasons over evolving object appearance (what) and updated 3D location (where), and employs a dedicated New Track token to reason about novel objects.  \nThanks to its design choices of using relative distances and evolving track representations, Whareformer is trained on a small set of 56 videos but achieves SOTA performance on 260 long test videos from three datasets:  \nEPIC-KITCHENS-100 (unseen videos), IT3DEgo, and HD-EPIC, with significant absolute improvements over prior work.  \n1 Introduction  \nHumans effortlessly maintain a mental map of their surroundings, simultaneously tracking multiple objects of interest, such as the keys on the counter, the phone on the coffee table, the mug by the sink. This ability to maintain an understanding of “what is where” for multiple objects, even when they are temporarily out of sight, is a core component of how humans interact with the world. For an autonomous agent to perform complex tasks in a human environment, replicating this capacity is a step towards human-like spatial intelligence. While this spatial understanding could be learned implicitly, representing it as an explicit 3D map provides a stable, viewpoint-invariant basis for locating and interacting with objects within this environment.  \nTo formalise this concept for egocentric videos, Plizzari et al. introduced the Out of Sight, Not Out of Mind (OSNOM) task [26] . It measures a model’s ability to (1) recall the 3D locations of objects over time, even when out of view  \n2 J. Chalk et al.  \nFig. 1: Overview of Whareformer. In the OSNOM task, the goal is to assign the current observation (highlighted in yellow in the current frame) of a 3D egocentric scene with one of the known objects (which respectively correspond to a kettle in red, a knife in green, or a tin of chopped tomatoes in blue) . We propose Whareformer, a model that jointly reasons about the appearance and the location of objects to decide whether to associate the observation with an existing track or initialise a new one.  \ndue to egomotion, object movement, or occlusion and (2) maintain the identity of objects as they change their appearance due to interactions and manipulations. The task can be simplified by explicit 3D reasoning, particularly when an object is occluded or when egomotion observes the same object from different viewpoints. An effective method would (1) maintain a coherent understanding of what is where within a consistent 3D coordinate frame, mirroring the spatial cognitive ability of humans and (2) learn how object appearances change during interactions, enabling accurate tracking for objects that are closely positioned.  \nOSNOM focuses on unstructured, human-centric egocentric scenarios characterised by rapid camera motion, frequent occlusions, and the need to maintain object permanence over extended timescales (typically up to half an hour) . Prior work [7, 16, 26, 40] tackles these challenges using heuristic association strategies based on fixed, manually defined thresholds. This restricts adaptabil","cbCaiamwotpnECpn","https://ap.wps.com/l/cbCaiamwotpnECpn","pdf",31154697,4,1,23,"English","en",105,"# Introduction\n# Related Work","[{\"question\":\"What problem does OSNOM in egocentric videos aim to solve?\",\"answer\":\"OSNOM requires tracking objects moved by the camera wearer while maintaining knowledge of instance 3D locations and identities over time, even when objects are out of view or heavily occluded.\"},{\"question\":\"How does Whareformer model “what” and “where” for long-term tracking?\",\"answer\":\"Whareformer uses a What-and-Where Transformer that jointly reasons over evolving object appearance (“what”) and updated 3D location (“where”) for consistent tracking across frames.\"},{\"question\":\"How does Whareformer decide whether to associate an observation with an existing track or start a new one?\",\"answer\":\"Whareformer includes a track assignment module and an explicit New Track token, treating new track creation as a primary learned decision rather than relying on fixed manual cost thresholds.\"}]","Whareformer: Learning to Track What is Where in Long Egocentric Videos | PDF",1784177693,58,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"whareformer-learning-to-track-what-is-where-in-long-egocentric-videos","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":20},"https://docshare.wps.com/document/whareformer-learning-to-track-what-is-where-in-long-egocentric-videos/82031/",{"url":53,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-30","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does OSNOM in egocentric videos aim to solve?","Question",{"text":76,"@type":77},"OSNOM requires tracking objects moved by the camera wearer while maintaining knowledge of instance 3D locations and identities over time, even when objects are out of view or heavily occluded.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does Whareformer model “what” and “where” for long-term tracking?",{"text":81,"@type":77},"Whareformer uses a What-and-Where Transformer that jointly reasons over evolving object appearance (“what”) and updated 3D location (“where”) for consistent tracking across frames.",{"name":83,"@type":74,"acceptedAnswer":84},"How does Whareformer decide whether to associate an observation with an existing track or start a new one?",{"text":85,"@type":77},"Whareformer includes a track assignment module and an explicit New Track token, treating new track creation as a primary learned decision rather than relying on fixed manual cost thresholds.","https://schema.org",{"og:url":53,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]