[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82233-en":3,"doc-seo-82233-105":29,"detail-sidebar-cat-0-en-105":95},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82233,962075114765,"Quinn","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","TSR-Ego: Temporally Guided Stereo Refinement Framework for Egocentric 3D Human Pose Estimation","Egocentric 3D human pose estimation from head-mounted stereo fisheye cameras is hindered by distortion, self-occlusion, and truncation of body joints outside the field of view. Many stereo methods use temporal cues only as auxiliary pose context, leaving weak, occluded, or ambiguous current-frame evidence underexploited. TSR-Ego injects motion evidence at the feature level using a causal temporal feature mixer and a single-stage causal stereo decoder with temporally guided cross-attention, improving robustness. Experiments validate state-of-the-art performance on UnrealEgo2 and UnrealEgo-RW.","TSR-Ego: Temporally Guided Stereo Refinement Framework for Egocentric 3D Human Pose Estimation  \nMd Mushfiqur Azam  \n[mdmushfiqur.azam@utsa.edu](mdmushfiqur.azam@utsa.edu)[ ](mdmushfiqur.azam@utsa.edu)John Quarles  \n[john.quarles@utsa.edu](john.quarles@utsa.edu)  \nKevin Desai  \n[kevin.desai@utsa.edu](kevin.desai@utsa.edu)  \nThe University of Texas at San Antonio  \n[ cs .CV] 10 Jul 2026  \nAbstract  \nEgocentric 3D human pose estimation from head-mounted stereo cameras is challenging due to fisheye distortion, severe self-occlusion, and frequent truncation of body joints outside the camera field of view. Recent stereo egocentric methods have improved performance through heatmap lifting, stereo correspondence, and transformer-based refinement, but they often rely heavily on frame-local evidence or use temporal information only as auxiliary pose-level context. This limits robustness when current-frame stereo cues are weak, occluded, or ambiguous. We propose TSR-Ego, a temporally guided stereo framework that couples short-term motion evidence with projection-guided feature sampling. The model first enriches dense stereo feature maps using a causal depthwiseseparable temporal convolution, allowing past visual evidence to influence the feature space before deformable cross-attention. A single-stage causal stereo decoder then refines learned 3D joint queries through temporal self-attention, joint self-attention, and  \n© 2026 . The copyright of this document resides with its authors. It may be distributed unchanged freely in print or electronic forms.  \nsevere for lower-body joints, which are often weakly observed by head-mounted cameras and must be inferred from incomplete visual evidence. Stereo fisheye cameras provide an attractive sensing setup because they introduce cross-view geometric cues while remaining compatible with compact wearable devices. However, effectively exploiting these cues remains difficult when body parts are truncated, stereo evidence is weak or one-sided, and visible joint evidence changes rapidly over time.  \nRecent datasets and methods have substantially advanced stereo egocentric 3D pose estimation. Large-scale benchmarks such as UnrealEgo, UnrealEgo2, and UnrealEgo-RW provide synthetic and real-world head-mounted fisheye stereo data for studying full-body pose estimation under severe occlusion and limited field of view [1, 2] . Building on these benchmarks, early approaches often rely on intermediate 2D heatmaps and lift them to 3D, which provides strong image-space localization but remains underconstrained when joints are occluded or outside the camera view. Recent stereo egocentric methods have improved 3D pose estimation by lifting image-space heatmaps into 3D poses, exploiting stereo correspondence and egocentric geometric cues, and refining joint representations with transformer-based attention [2, 11, 12, 31] . Despite this progress, existing stereo egocentric methods still leave an important gap between spatial stereo reasoning and temporal guidance. Heatmap-based methods depend on intermediate 2D detections, geometry-aware methods rely on reliable current-frame stereo cues, and transformer-based refinement methods often build their 3D hypotheses primarily from the current stereo observation. Video-based approaches introduce temporal context, but temporal information is commonly used to augment joint representations or improve pose consistency rather than to condition the stereo features sampled during refinement. This separation is limiting in egocentric views, where joints may be occluded, truncated by the headset field of view, or visible in only one fisheye camera. In such cases, current-frame spatial evidence alone may produce an unreliable pose hypothesis, eventhough recent frames provide useful cues about joint motion and visibility.  \nWe address this limitation with TSR-Ego, a temporally guided stereo refinement framework for egocentric 3D human pose estimation. Our key idea is to inject temp","cbCaikwQ8umvmiNp","https://ap.wps.com/l/cbCaikwQ8umvmiNp","pdf",2657397,1,18,"English","en",105,"# Abstract\n# Contributions\n# Related Work","[{\"question\":\"What problem does TSR-Ego address in egocentric stereo 3D pose estimation?\",\"answer\":\"It targets the gap where prior methods rely too heavily on frame-local stereo evidence or use temporal cues only after pose-level reasoning, which breaks down when joints are occluded or truncated.\"},{\"question\":\"How does TSR-Ego incorporate temporal information into stereo refinement?\",\"answer\":\"It enriches dense per-view stereo feature maps using a causal depthwise-separable temporal convolution, then refines joint queries with a single-stage causal decoder using temporal self-attention and projection-guided fisheye deformable stereo cross-attention.\"},{\"question\":\"Why is feature-level temporal conditioning beneficial in head-mounted fisheye setups?\",\"answer\":\"Because occluded or one-sided joint visibility can still be inferred from recent motion, and temporally enriched features provide better conditioning for deformable stereo sampling than frame-only representations.\"},{\"question\":\"On which datasets does TSR-Ego report improved results?\",\"answer\":\"State-of-the-art results are reported on UnrealEgo2 and UnrealEgo-RW.\"}]",1784179017,45,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":90,"head_meta":92,"extra_data":94,"updated_unix":27},"tsr-ego-temporally-guided-stereo-refinement-framework-for-egocentric-3d-human-pose-estimation","",{"@graph":35,"@context":89},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/tsr-ego-temporally-guided-stereo-refinement-framework-for-egocentric-3d-human-pose-estimation/82233/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81,85],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does TSR-Ego address in egocentric stereo 3D pose estimation?","Question",{"text":75,"@type":76},"It targets the gap where prior methods rely too heavily on frame-local stereo evidence or use temporal cues only after pose-level reasoning, which breaks down when joints are occluded or truncated.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does TSR-Ego incorporate temporal information into stereo refinement?",{"text":80,"@type":76},"It enriches dense per-view stereo feature maps using a causal depthwise-separable temporal convolution, then refines joint queries with a single-stage causal decoder using temporal self-attention and projection-guided fisheye deformable stereo cross-attention.",{"name":82,"@type":73,"acceptedAnswer":83},"Why is feature-level temporal conditioning beneficial in head-mounted fisheye setups?",{"text":84,"@type":76},"Because occluded or one-sided joint visibility can still be inferred from recent motion, and temporally enriched features provide better conditioning for deformable stereo sampling than frame-only representations.",{"name":86,"@type":73,"acceptedAnswer":87},"On which datasets does TSR-Ego report improved results?",{"text":88,"@type":76},"State-of-the-art results are reported on UnrealEgo2 and UnrealEgo-RW.","https://schema.org",{"og:url":51,"og:type":91,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":93,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":96},[97,101,105,109,114,119,124,127,132,135,139],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},"Exam",70,"exam",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},5,"Comic",60,"comic",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},6,"Technology",50,"technology",{"id":120,"doc_module":4,"doc_module_name":45,"category_name":121,"show_sort_weight":122,"slug":123},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":125,"slug":126},30,"research-report",{"id":128,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":130,"slug":131},9,"Religion & Spirituality",20,"religion-spirituality",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":130,"slug":134},"World Cup","world-cup",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":136,"slug":138},10,"Lifestyle","lifestyle",{"id":140,"doc_module":4,"doc_module_name":45,"category_name":141,"show_sort_weight":110,"slug":142},19,"General","general"]