[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82232-en":3,"doc-seo-82232-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82232,1374391974468,"Eden","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","4D Human-Scene Reconstruction from Low-Overlap Captures","High-fidelity 4D human capture typically depends on dense, highly overlapping camera arrays, but real-world studios often provide only a few uncalibrated, low-overlap cameras with frequent occlusions and multiple interacting people. StudioRecon addresses these limitations by decoupling background and human representations, densifying background supervision via camera-controlled novel-view synthesis from a video diffusion model, robustly initializing deformable Gaussian humans through cross-view identity association and multiview keypoint fitting, and refining results with recursive enhancement using motion-adaptive consistency injection.","arXiv :2607 .09125v1 [ cs .CV] 10 Jul 2026  \n4D Human-Scene Reconstruction from Low-Overlap Captures  \nMINHYUK HWANG∗ , Seoul National University, Republic of Korea SANGMIN KIM∗ , Seoul National University, Republic of Korea SEUNGUK DO, Seoul National University, Republic of Korea DANEUL KIM, Seoul National University, Republic of Korea JAESIK PARK†, Seoul National University, Republic of Korea  \nLow-overlap Initial Reconstruction  \n4 Input Videos  \nOur Enhanced Result  \nFig. 1. Given only as few as four sparse, low-overlap input videos (left), StudioRecon first reconstructs decoupled Gaussians for background and humans (right) . The reconstructed Gaussians enable rendering from novel viewpoints, and our recursive enhancement module further refines the rendered output (bottom) .  \n∗ Both authors contributed equally.†Corresponding author.  \nAuthors’ Contact Information: Minhyuk Hwang, Seoul National University, Seoul, Republic of Korea, [mhhlego@snu.ac.kr](mhhlego@snu.ac.kr); Sangmin Kim, Seoul National University, Seoul, Republic of Korea, [sm.kim@snu.ac.kr](sm.kim@snu.ac.kr); Seunguk Do, Seoul National University, Seoul, Republic of Korea, [seunguk.do@snu.ac.kr](seunguk.do@snu.ac.kr); Daneul Kim, Seoul National University, Seoul, Republic of Korea, [carpedkm@snu.ac.kr](carpedkm@snu.ac.kr); Jaesik Park, Seoul National University, Seoul, Republic of Korea, [jaesik.park@snu.ac.kr](jaesik.park@snu.ac.kr).  \nThis work is licensed under a Creative Commons Attribution-NonCommercialNoDerivatives 4 .0 International License.  \nSIGGRAPH Conference Papers ’26, Los Angeles, CA, USA © 2026 Copyright held by the owner/author(s) .  \nACM ISBN 979-8-4007-2554-8/26/07  \n[https://doi.org/10.1145/3799902.3811165](https://doi.org/10.1145/3799902.3811165)  \nExisting volumetric capture of dynamic human performance achieves high fidelity with dense camera arrays. However, in real-world scenarios, only a handful of low-overlap cameras are available, which degrades the output quality and leaves large areas unobserved. Recent 4D reconstruction methods have focused on low-overlap settings, yet they still produce noticeable artifacts in under-observed regions. Video diffusion models have emerged as another option, but they show geometrically inconsistent results for humans. To address these limitations, we propose StudioRecon, a pipeline that reconstructs 4D human scenes from sparse, low-overlap cameras by decoupling background and humans. We densify background supervision by synthesizing hundreds of camera-controlled novel views with a video diffusion model. We also robustly initialize deformable Gaussian humans with cross-view identity association and triangulated multiview keypoint fitting. Finally, our recursive enhancement module with motion-adaptive consistency injection harmonizes the composed output, thereby further avoiding remaining artifacts. We achieve state-of-the-art  \nSIGGRAPH Conference Papers ’26, July 19–23, 2026, Los Angeles, CA, USA.  \nnovel view synthesis across four real-world datasets and demonstrate applications such as novel trajectory rendering and human replacement. Project page: [https://sisyphm.github.io/studiorecon-page/](https://sisyphm.github.io/studiorecon-page/) .  \nCCS Concepts: • Computing methodologies → Reconstruction; Rendering.  \nACM Reference Format:  \nMinhyuk Hwang, Sangmin Kim, Seunguk Do, Daneul Kim, and Jaesik Park.  \n2026. 4D Human-Scene Reconstruction from Low-Overlap Captures. In Special Interest Group on Computer Graphics and Interactive Techniques Conference Conference Papers (SIGGRAPH Conference Papers ’26), July 19–23, 2026, Los Angeles, CA, USA. ACM, New York, NY, USA, 17 pages. [https:](https:)//[doi.org/10.1145/3799902.3811165](doi.org/10.1145/3799902.3811165)  \n1 Introduction  \nHigh-fidelity 4D human capture has become essential for entertainment, sports broadcasting, and virtual production. Professional volumetric systems achieve compelling results but require dozens to hundreds of cameras in contro","cbCaim8hOTSLPGXi","https://ap.wps.com/l/cbCaim8hOTSLPGXi","pdf",26922457,1,17,"English","en",105,"# Introduction\n## Problem: low-overlap in-the-wild studio capture\n## Related work and limitations\n## Proposed approach: StudioRecon","[{\"question\":\"What problem does StudioRecon target in real-world 4D human capture?\",\"answer\":\"It targets the quality degradation and unobserved regions caused by using only a handful of sparse, low-overlap cameras in uncalibrated, in-the-wild studio settings with multiple people and frequent occlusions.\"},{\"question\":\"How does StudioRecon handle background and human reconstruction differently?\",\"answer\":\"It decouples background and humans, using video diffusion to synthesize plausible, evidence-consistent background from novel viewpoints, while initializing deformable Gaussian humans using cross-view identity association and triangulated multiview keypoint fitting.\"},{\"question\":\"How does StudioRecon reduce remaining artifacts after the initial reconstruction?\",\"answer\":\"It applies a recursive enhancement module with motion-adaptive consistency injection to harmonize the composed output and further avoid geometric artifacts.\"}]",1784179013,43,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"4d-human-scene-reconstruction-from-low-overlap-captures","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/4d-human-scene-reconstruction-from-low-overlap-captures/82232/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does StudioRecon target in real-world 4D human capture?","Question",{"text":75,"@type":76},"It targets the quality degradation and unobserved regions caused by using only a handful of sparse, low-overlap cameras in uncalibrated, in-the-wild studio settings with multiple people and frequent occlusions.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does StudioRecon handle background and human reconstruction differently?",{"text":80,"@type":76},"It decouples background and humans, using video diffusion to synthesize plausible, evidence-consistent background from novel viewpoints, while initializing deformable Gaussian humans using cross-view identity association and triangulated multiview keypoint fitting.",{"name":82,"@type":73,"acceptedAnswer":83},"How does StudioRecon reduce remaining artifacts after the initial reconstruction?",{"text":84,"@type":76},"It applies a recursive enhancement module with motion-adaptive consistency injection to harmonize the composed output and further avoid geometric artifacts.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]