[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82186-en":3,"doc-seo-82186-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82186,2336464648746,"Skyler","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","DETRAM End-to-End Detection Tracking and Recovery of Human Meshes","DETRAM presents an end-to-end Transformer framework for multi-person human mesh recovery (HMR) and tracking in video, targeting the difficulties of crowded scenes with frequent occlusions, truncations, and individuals entering or leaving. Traditional multi-stage pipelines accumulate errors, increase latency, and rely on external detectors, limiting tracked entity counts. DETRAM performs detection, reconstruction, and identity-consistent tracking jointly in a single feed-forward pass using persistent learnable query embeddings with automatic or user-prompted tracking. Experiments report state-of-the-art tracking and competitive reconstruction on major benchmarks.","DETRAM: End-to-end DEtection, Tracking and Recovery of HumAn Meshes  \nChunggi Lee 1 , Seonwook Park2 , Wanhua Li3 , Umar Iqbal2,*, and Hanspeter  \nPfister 1,*  \n1 Harvard University 2 NVIDIA 3 Nanyang Technological University  \narXiv :2607 .09089v1 [ cs .CV] 10 Jul 2026  \nFig. 1: Comparison between traditional multi-stage pipelines and our unified one-stage DETRAM framework. Conventional approaches decompose the problem into sequential detection, tracking, and human mesh recovery (HMR) modules, causing error accumulation, increased latency, and difficulty maintaining consistent identities over time. In contrast, DETRAM performs all three tasks jointly in a single feed-forward pass. This enables efficient, end-to-end multi-person reconstruction and tracking without requiring either an external detector or a separate learned tracking module. DETRAM discovers and tracks new people as they enter the scene, while also supporting optional user-prompted tracking, allowing individuals to be added or specified when desired.  \nAbstract. In the task of human mesh recovery (HMR), multi-person scenes are particularly difficult to handle due to the many entities that appear and occlusions between them over time. In particular for video inputs, there is a need to track each entity reliably and consistently.  \nExisting methods rely on pretrained human detection modules, increasing their runtime and limiting the number of tracked entities. We present DETRAM, a unified framework for multi-person HMR and tracking that simultaneously detects, reconstructs, and tracks humans across time, both automatically and via user prompts. DETRAM uses a single transformer decoder with an identity-consistent set of learnable query embeddings that persist across frames: detection queries discover new people, tracking queries maintain pose and shape for existing individuals, and prompt queries follow user-specified identities. Our approach achieves state-of-the-art tracking results on PoseTrack21, 3DPW, BEDLAM, and MuPoTS-3D, and competitive reconstruction accuracy on BEDLAM and 3DPW, while uniquely supporting prompt-based tracking of individuals in multi-person scenes. To our knowledge, this is the first method to unify promptability and multi-person HMR with tracking in an endto-end trainable framework, enabling user-directed human analysis in videos.  \n*  \nCorresponding authors.  \n2 C. Lee et al.  \n1 Introduction  \nVideos of public places, sporting events, and urban environments typically contain many humans and their movements over time. Tracking each human and recovering their 3D pose and shape in such crowded videos enables applications in sports analytics, mixed reality, safety monitoring, and human behavior understanding. However, these scenes are highly challenging: people frequently occlude one another, body parts are truncated by crops or the camera view, and individuals can enter or leave the scene at any time. These factors often lead to tracking failures or inconsistent 3D reconstructions, especially in long sequences.  \nExisting methods for multi-person human mesh recovery (HMR) and tracking are mostly built as multi-stage pipelines. A detector first outputs bounding boxes, which are subsequently cropped and fed into pose or mesh recovery networks. Tracking IDs are then assigned in a separate association stage based on appearance or pose features [14, 42 , 43] . While effective in some settings, this design has several weaknesses. Bounding box errors or misaligned crops can remove important spatial context, leading to missing limbs or ambiguous associationsin crowded scenes. Moreover, errors propagate across stages, each module may need to be trained independently, and the overall runtime increases.  \nRecent transformer-based approaches address some of these issues by operating directly on the full image and regressing human pose and shape without cropping, within detection-transformer architectures for multi-person HMR [2,46 ,47] . By reasoning o","cbCaigqX1LWj4hsY","https://ap.wps.com/l/cbCaigqX1LWj4hsY","pdf",4398144,1,25,"English","en",105,"# Introduction\n## Background and Motivation\n## Limitations of Existing Multi-stage Pipelines\n## Transformer-based Approaches and Remaining Gaps\n## DETRAM: End-to-end Detection, Tracking, and Recovery","[{\"question\":\"Why is multi-person human mesh recovery and tracking difficult in real videos?\",\"answer\":\"Crowded scenes cause frequent occlusions, body truncations from crops or camera views, and people entering or leaving the frame over time. These factors lead to tracking failures and inconsistent 3D reconstructions, especially in long sequences.\"},{\"question\":\"What limitation of multi-stage pipelines does DETRAM address?\",\"answer\":\"Multi-stage systems rely on an external detector and separate association steps, where bounding-box or crop errors can remove key spatial context. Errors also propagate across modules, requiring independent training and increasing runtime latency.\"},{\"question\":\"How does DETRAM perform end-to-end detection, tracking, and recovery?\",\"answer\":\"DETRAM uses persistent learnable identity-consistent query embeddings in a transformer decoder. Detection queries discover new people, tracking queries carry pose/shape/identity across frames, and prompt queries enable user-specified identity control without a separate learned tracking module.\"}]",1784178677,63,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"detram-end-to-end-detection-tracking-and-recovery-of-human-meshes","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/detram-end-to-end-detection-tracking-and-recovery-of-human-meshes/82186/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is multi-person human mesh recovery and tracking difficult in real videos?","Question",{"text":75,"@type":76},"Crowded scenes cause frequent occlusions, body truncations from crops or camera views, and people entering or leaving the frame over time. These factors lead to tracking failures and inconsistent 3D reconstructions, especially in long sequences.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What limitation of multi-stage pipelines does DETRAM address?",{"text":80,"@type":76},"Multi-stage systems rely on an external detector and separate association steps, where bounding-box or crop errors can remove key spatial context. Errors also propagate across modules, requiring independent training and increasing runtime latency.",{"name":82,"@type":73,"acceptedAnswer":83},"How does DETRAM perform end-to-end detection, tracking, and recovery?",{"text":84,"@type":76},"DETRAM uses persistent learnable identity-consistent query embeddings in a transformer decoder. Detection queries discover new people, tracking queries carry pose/shape/identity across frames, and prompt queries enable user-specified identity control without a separate learned tracking module.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]