[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83451-en":3,"doc-seo-83451-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83451,1099513958607,"Jiven","https://ap-avatar.wpscdn.com/avatar/100002390cf8733938c?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778829742770036399",8,"Research & Report","Progressive Pose-Guided 4D Animal Reconstruction from Monocular Video","Progressive test-time optimization enables high-fidelity 4D animal reconstruction from a single monocular video despite large inter-species variation, complex articulations, and missing reliable templates. The method builds on 3D Gaussian Splatting using a coarse shape prior plus principled disentanglement between articulated pose and nonrigid deformation. It employs symmetry-aware temporal encoding with bilateral cues to absorb camera drift, and a part-conditioned deformation mechanism driven by learnable part anchors and a learnable skinning field. Experiments show strong generalization across species with improved geometric accuracy, temporal consistency, and visual fidelity under severe prior mismatch.","arXiv :2607 .00157v1 [ cs .CV] 30 Jun 2026  \nProgressive Pose-Guided 4D Animal Reconstruction from Monocular Video  \nSiyuan Li 1 , Weiying Chen 1, Yilin Wang 1, Xinxin Zuo2, Xingyu Li 1, and Li Cheng 1B  \n1 University of Alberta, Edmonton, AB, Canada {sli20, [lcheng5}@ualberta.ca](lcheng5}@ualberta.ca)  \n2 Concordia University, Montreal, QC, Canada  \nFig. 1: Given monocular videos of animals (top), our method produces high-fidelity 4D models enabling free-viewpoint rendering across time and viewing angles (bottom) . Center: canonical 3D Gaussians colored by skinning weights.  \nAbstract. Reconstructing 4D animals from monocular videos is challenging due to large inter-species variation, complex articulations, and the lack of reliable templates. Existing approaches typically rely on either strict category-specific priors that restrict generalization, or unconstrained generative models that sacrifice input fidelity. To bridge this gap, we present a progressive test-time optimization framework built on 3D Gaussian Splatting for high-fidelity 4D animal reconstruction from a single video. Our key insight is that a coarse shape prior suffices when coupled with a progressive strategy that disentangles articulated pose from nonrigid deformation. Specifically, we employ a symmetry-aware temporal encoding that exploits bilateral cues while absorbing camera estimation drift and a part-conditioned deformation mechanism guided by learnable part anchors and a learnable skinning field. Extensive experiments demonstrate that our approach generalizes robustly across diverse species, achieving superior geometric accuracy, temporal consistency, and visual fidelity compared to existing baselines, even under severe prior mismatch.  \nProject page: [https://syl-322.github.io/ReWild4D/](https://syl-322.github.io/ReWild4D/)  \n2 S. Li et al.  \n1 Introduction  \nAnimals in the natural world display a stunning diversity of shapes and behaviors. Accurately reconstructing their 3D shape and motion from visual data is crucial for various applications ranging from wildlife monitoring, animal conservation and ethology research, to immersive media content creation. Despite the wide accessibility of monocular video, the task of creating realistic 4D animal models from monocular video presents a significant challenge in computer vision. This is primarily due to the inherent complexity of animal morphology and behaviors, as well as the fact that their appearance and motion are only partly observable from a monocular video.  \nThe task of 3D animal reconstruction presents unique challenges compared to 3D human reconstruction. Human models, such as SMPL [20, 23], benefit from well-studied anatomical structures and abundant 3D motion capture datasets. In contrast, diverse animal species exhibit extreme shape and motion variations, yet very little animal motion capture benchmarks are available. The pioneer SMAL [50] is a parametric SMPL-like animal model learned from a limited collection of toy figurines; it captures a rather limited category of species and lacks realistic details of shape and motion. The dilemma of both scarcely labeled, partially observable animal data, and extraordinarily diverse shape & motion variations across animal species, forces a trade-off where, category-specific methods [2, 27, 28, 31, 33] achieve higher reconstruction quality by training on annotated dataset but struggle with generalization, while category-agnostic approaches [1, 6, 16] improve coverage at the cost of reconstruction fidelity.  \nExtending to 4D animal reconstruction from monocular videos reveals a fundamental tension between representation flexibility and computational efficiency. Mesh-based methods [29, 38, 39] are efficient but topologically constrained, while neural implicit methods [40, 41] offer flexibility at prohibitive computational costs. 3D Gaussian Splatting (3DGS) [10] emerges as a powerful alternative: it combines the rendering efficiency of explicit geometry wi","cbCaiuWWGtqAnO4w","https://ap.wps.com/l/cbCaiuWWGtqAnO4w","pdf",7729689,4,1,19,"English","en",105,"# Introduction\n## Challenges in 3D/4D Animal Reconstruction\n## Related Work and Existing Trade-offs\n## Proposed Progressive Pose-Guided Framework\n## Key Contributions","[{\"question\":\"What problem does the paper address?\",\"answer\":\"It addresses reconstructing high-fidelity 4D animal models (shape and motion over time) from a single monocular video, despite variability across species and complex articulations.\"},{\"question\":\"Why do existing methods struggle, according to the document?\",\"answer\":\"Prior approaches either rely on strict category-specific priors that limit generalization, or use unconstrained generative models that reduce input fidelity. The paper attributes the trade-off to treating shape priors as rigid constraints rather than flexible initializations.\"},{\"question\":\"What is the core idea of the proposed method?\",\"answer\":\"A progressive test-time optimization framework based on 3D Gaussian Splatting uses a coarse shape initialization and continuously evolves it by disentangling articulated pose from nonrigid deformation.\"}]",1784188020,48,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"progressive-pose-guided-4d-animal-reconstruction-from-monocular-video","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/progressive-pose-guided-4d-animal-reconstruction-from-monocular-video/83451/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address?","Question",{"text":75,"@type":76},"It addresses reconstructing high-fidelity 4D animal models (shape and motion over time) from a single monocular video, despite variability across species and complex articulations.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why do existing methods struggle, according to the document?",{"text":80,"@type":76},"Prior approaches either rely on strict category-specific priors that limit generalization, or use unconstrained generative models that reduce input fidelity. The paper attributes the trade-off to treating shape priors as rigid constraints rather than flexible initializations.",{"name":82,"@type":73,"acceptedAnswer":83},"What is the core idea of the proposed method?",{"text":84,"@type":76},"A progressive test-time optimization framework based on 3D Gaussian Splatting uses a coarse shape initialization and continuously evolves it by disentangling articulated pose from nonrigid deformation.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},"General","general"]