[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82063-en":3,"doc-seo-82063-105":28,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":11,"language":21,"language_code":22,"site_id":23,"html_lang":22,"table_of_contents":24,"faqs":25,"seo_title":13,"seo_description":14,"update_tm":26,"read_time":27},82063,13056703019404,"Miles","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","StereoSplat+ Feed-Forward Stereo Gaussian Splatting with Diffusion-Assisted Progressive Inference","Recent advances in 3D Gaussian Splatting (3DGS) provide render-ready scene representations for novel-view synthesis, but typical pipelines require multi-view observations or non-causal access to future frames. StereoSplat+ addresses causal reconstruction from a single stereo pair, where occlusions, limited field of view, and missing geometry make geometry coverage difficult. The method uses StereoSplat, an input-invariant feed-forward 3D Gaussian estimator with cost-volume and triplane 3D-volume branches, plus continuous pose encoding for view-count generalization. A diffusion-enhanced one-shot progressive loop renders novel views, refines them via a one-step diffusion enhancer, and reinjects them as pseudo inputs. Experiments on KITTI-360 show improved novel-view quality and geometry accuracy, especially under occlusion and strong view extrapolation.","StereoSplat+: Feed-Forward Stereo Gaussian Splatting with Diffusion-Assisted Progressive Inference  \nZihua Liu 1 and Masatoshi Okutomi.1  \narXiv :2607 .08808v 1 [ cs .CV] 9 Jul 2026  \nAbstract— Recent advances in 3D Gaussian Splatting (3DGS) have enabled high-quality, render-ready scene representations for novel-view synthesis. However, most existing 3DGS pipelines rely on multi-view observations (or non-causal access to future frames) to achieve sufficient coverage, which is often unavailable in on-device robotics and AR settings where sensing is restricted to a single stereo rig. Recovering a high-quality 3DGS scene from one stereo observation, therefore, remains challenging due to occlusions, limited field of view, and missing geometry. We present StereoSplat+, a diffusion-enhanced feed-forward framework that enables causal reconstruction from a single stereo pair. Our method builds on two key components. First, we propose StereoSplat, an input-invariant feed-forward 3D Gaussian estimator that takes a variable number of posed stereo pairs as input and predicts high-quality 3D Gaussians. StereoSplat fuses complementary geometry cues via a costvolume branch and a triplane-based 3D volume branch, and leverages continuous pose encoding to generalize across view counts and camera configurations. Second, since multiple posed stereo pairs are typically unavailable at inference time, we introduce a diffusion-enhanced one-shot progressive inference scheme called StereoSplat+: starting from one stereo pair, we render novel stereo views from the predicted 3DGS, refine them with a one-step diffusion enhancer, and feed them back as additional inputs to update the 3DGS. Experiments on the KITTI-360 dataset show that StereoSplat+ improves novel-view rendering quality and geometry accuracy, especially in occluded regions and under strong view extrapolation, outperforming recent feed-forward 3DGS baselines.  \nI. INTRODUCTION  \nFeed-forward 3D Gaussian Splatting (3DGS) [1], [2], [3],[4],[5] enables real-time and generalizable scene reconstruction by directly predicting a set of view-renderable Gaussians from multi-view inputs. However, in many practical roboticsand on-device AR settings, the input is often restricted toa single stereo pair. With such limited coverage and field of view, distant surfaces and regions occluded in the stereo observations are only weakly constrained, causing existing feed-forward 3DGS pipelines to struggle with reliable geometry coverage and photorealistic novel-view synthesis. Incorporating more views can mitigate these issues, but it increases latency and typically introduces non-causal requirements (e.g., access to future frames or accurate multi-view poses), which are incompatible with real-time deployment. To tackle this issue, we introduce StereoSplat+, an inputinvariant, feed-forward 3DGS framework for a single stereo pair input with diffusion-enhanced one-shot progressive inference. As shown in Figure 1, compared with conventional feed-forward 3DGS estimator,our key idea is to couple  \n1 Department of systems and Control Engineering, Institute of Science Tokyo, Japan. {zliu,[mxo](mxo}@ok.sc.e.titech.ac.jp)[}](mxo}@ok.sc.e.titech.ac.jp)[@ok.sc.e.titech.ac.jp](mxo}@ok.sc.e.titech.ac.jp)  \nFig. 1. The pipeline of our proposed StereoSplat+ . StereoSplat+ performs progressive inference: starting from one stereo pair, it estimates an initial 3DGS, renders novel stereo views, refines them with a one-step diffusion prior, and re-injects the enhanced views as pseudo inputs to update the 3DGaussians.  \na view-count-agnostic 3DGS predictor with a single-step diffusion enhancer for progressive inference: Starting from one stereo pair, we predict a provisional 3D Gaussian set, render novel stereo views, enhance the rendered images with a one-step diffusion model, and re-inject the enhanced views as pseudo views to provide additional geometric cues for new Gaussian estimation. The final result is the confidencebased fus","cbCaiiGiEfiHRHor","https://ap.wps.com/l/cbCaiiGiEfiHRHor","pdf",4540957,1,"English","en",105,"# Introduction\n## Method Overview\n## Dual-Branch Predictor\n## Progressive One-Shot Diffusion Inference\n## Experiments and Results","[{\"question\":\"What problem does StereoSplat+ address in 3D Gaussian Splatting?\",\"answer\":\"It targets causal, on-device reconstruction from only a single stereo pair, where limited coverage and occlusions hinder conventional feed-forward 3DGS pipelines that rely on multi-view inputs.\"},{\"question\":\"How does StereoSplat+ generate 3D Gaussian representations from stereo inputs?\",\"answer\":\"It uses StereoSplat, an input-invariant feed-forward estimator that fuses a cost-volume branch with a lightweight 3D volume (triplane-based) branch, and uses continuous pose encoding to generalize across view counts and camera configurations.\"},{\"question\":\"What role does diffusion-assisted progressive inference play during inference?\",\"answer\":\"Starting from one stereo pair, it renders novel stereo views using the predicted 3DGS, refines the rendered images with a one-step diffusion model, and reinjects the enhanced views as pseudo inputs to update the 3D Gaussian scene in exactly one render–enhance–reinject round.\"}]",1784177956,20,{"code":4,"msg":29,"data":30},"ok",{"site_id":23,"language":22,"slug":31,"title":13,"keywords":32,"description":14,"schema_data":33,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":26},"stereosplat-feed-forward-stereo-gaussian-splatting-with-diffusion-assisted-progressive-inference","",{"@graph":34,"@context":84},[35,52,67],{"@type":36,"itemListElement":37},"BreadcrumbList",[38,42,46,49],{"item":39,"name":40,"@type":41,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":43,"name":44,"@type":41,"position":45},"https://docshare.wps.com/document/","Document",2,{"item":47,"name":12,"@type":41,"position":48},"https://docshare.wps.com/document/research-report/",3,{"item":50,"name":13,"@type":41,"position":51},"https://docshare.wps.com/document/stereosplat-feed-forward-stereo-gaussian-splatting-with-diffusion-assisted-progressive-inference/82063/",4,{"url":50,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":22,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":39,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"What problem does StereoSplat+ address in 3D Gaussian Splatting?","Question",{"text":74,"@type":75},"It targets causal, on-device reconstruction from only a single stereo pair, where limited coverage and occlusions hinder conventional feed-forward 3DGS pipelines that rely on multi-view inputs.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How does StereoSplat+ generate 3D Gaussian representations from stereo inputs?",{"text":79,"@type":75},"It uses StereoSplat, an input-invariant feed-forward estimator that fuses a cost-volume branch with a lightweight 3D volume (triplane-based) branch, and uses continuous pose encoding to generalize across view counts and camera configurations.",{"name":81,"@type":72,"acceptedAnswer":82},"What role does diffusion-assisted progressive inference play during inference?",{"text":83,"@type":75},"Starting from one stereo pair, it renders novel stereo views using the predicted 3DGS, refines the rendered images with a one-step diffusion model, and reinjects the enhanced views as pseudo inputs to update the 3D Gaussian scene in exactly one render–enhance–reinject round.","https://schema.org",{"og:url":50,"og:type":86,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":88,"canonical":50},"index,follow",{"doc_id":7,"site_id":23},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,126,129,133],{"id":20,"doc_module":4,"doc_module_name":44,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":45,"doc_module":4,"doc_module_name":44,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":51,"doc_module":4,"doc_module_name":44,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":44,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":44,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":44,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":44,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":44,"category_name":124,"show_sort_weight":27,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":27,"doc_module":4,"doc_module_name":44,"category_name":127,"show_sort_weight":27,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":44,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":44,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]