[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86251-en":3,"doc-seo-86251-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86251,687197207919,"Theodora","https://ap-avatar.wpscdn.com/avatar/a000253d6f5f7c60be?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779446848396160552",8,"Research & Report","HyperGS: Fast and Generalizable Gaussian Video Representation","Gaussian Splatting enables strong video representation, yet prior approaches optimize Gaussians per video, causing slow encoding and weak transfer across different videos. HyperGS introduces a feedforward, optimization-free framework that predicts explicit Gaussian splatting parameters directly from any input video in a single forward pass. A factorized spatiotemporal Transformer extracts video tokens, while a query-based Transformer outputs 8-parameter Gaussians per frame. Training uses an adaptive rank-based geometric regularizer to prevent needle-like degeneration. Results show +2.9–3.1 dB PSNR gains, and 104–105× faster encoding with zero-shot higher-resolution generalization.","arXiv :2607 . 11500v1 [ cs .CV] 13 Jul 2026  \nHyperGS: Fast and Generalizable Gaussian Video Representation  \nFatimah Zohra 1 , Chen Zhao 1,2 , Shuming Liu 1 , Yahya Al Malallah 1 , Bernard Ghanem 1  \n1 King Abdullah University of Science and Technology (KAUST)  \n2 Harbin Institute of Technology, Shenzhen  \nAbstract  \nGaussian Splatting has emerged as an effective representation for video, but existing methods rely on per-video optimization. This leads to slow encoding and limits generalization across videos. To amortize this optimization, we propose HyperGS, a feedforward, optimizationfree approach that directly predicts Gaussian representations from any video in a single forward pass, speeding up encoding and decoding by orders of magnitude while generalizing to out-of-distribution videos at higher resolutions. In HyperGS, we design a factorized spatiotemporal Transformer to extract tokens from video, and a learnable query-based Transformer to obtain 8-parameter Gaussian representations for each video frame. We find that naively predicting Gaussians across diverse videos induces a needle-like degeneration that collapses training, and address this with a rank-based geometric regularizer whose strength adapts dynamically to stabilize optimization. HyperGS achieves encoding at 104–105 × the speed of per-video Gaussian optimization at matched reconstruction quality while generalizing zero-shot to 720p video, enabling higher-resolution rendering without re-encoding.  \nHyperGS improves PSNR by +2.9–3.1 dB over the prior video encoders on K400, SSv2, and UCF101 at a smaller video representation size. By predicting explicit 2D Gaussians in a single forward pass, HyperGS combines the fast, flexible rendering of Gaussian Splatting with the speed and generalization of feedforward prediction, advancing Gaussians as a practical direction for fast and generalizable video representation.  \n1 Introduction  \nRecent advances in 3D reconstruction have increasingly been introduced into the video representation field, changing how video data is encoded and reconstructed. Among these, 3D Gaussian Splatting (3DGS) Kerblet al. (2023) has emerged as a particularly effective representation, modeling a scene as a collection of explicit, differentiable primitives that can be rendered directly through rasterization rather than through an implicit, coordinate-based function. The adaptation of 3DGS to dynamic video in particular has brought improvements in rendering speed and geometric stability. Yet state-of-the-art 3DGS video models Bond et al.(2025); Liu et al. (2025); Sun et al. (2024) remain bound by a fundamental limitation: per-video optimization. Because each video’s Gaussians are fit independently through iterative, gradient-based optimization—often tens of thousands of steps and on the order of hours per video—none of the structure learned while encoding one video transfers to the next. This makes per-video Gaussian methods impractical for real-time processing.  \nA feedforward, optimization-free alternative instead amortizes this encoding by training a shared model to predict a video’s representation in a single forward pass. Prior work along this line predicts the weights of an implicit neural decoder, tying the representation to a specific neural architecture. We show that predicting explicit primitives instead decouples the representation from any implicit decoder, extending the efficient rendering and flexibility of Gaussian Splatting to a feedforward regime.  \nOur choice of Gaussians as the predicted representation follows directly from this decoupling. Gaussian parameters reduce the regression target into geometrically meaningful subproblems—position, scale, orientation, color—providing a natural inductive bias for the prediction task. Rendering via a sum of Gaussians  \nis also a direct learning objective, since the image is composed linearly from the predicted primitives, while an implicit decoder maps predicted weights to the reconstruct","cbCaio4MaacClWJx","https://ap.wps.com/l/cbCaio4MaacClWJx","pdf",7360966,6,1,19,"English","en",105,"# Introduction\n# Related Works","[{\"question\":\"What problem does HyperGS address compared with per-video Gaussian optimization?\",\"answer\":\"HyperGS addresses the high latency of per-video optimization, where each video’s Gaussians are fitted independently with iterative gradient-based steps. This prevents learned structure from transferring across videos and limits real-time processing.\"},{\"question\":\"How does HyperGS produce Gaussian representations for a video?\",\"answer\":\"HyperGS uses a factorized spatiotemporal Transformer to extract tokens from the video, then a learnable query-based Transformer to predict 8-parameter Gaussian splatting representations for each frame in a single forward pass.\"},{\"question\":\"What prevents training collapse when predicting Gaussians feedforward?\",\"answer\":\"Naively predicting Gaussians across diverse videos can drive them toward needle-like anisotropy, collapsing training mid-way. HyperGS stabilizes optimization using an adaptive rank-based geometric regularizer with dynamically adapting strength based on effective rank.\"}]",1784209823,48,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"hypergs-fast-and-generalizable-gaussian-video-representation","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/hypergs-fast-and-generalizable-gaussian-video-representation/86251/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does HyperGS address compared with per-video Gaussian optimization?","Question",{"text":76,"@type":77},"HyperGS addresses the high latency of per-video optimization, where each video’s Gaussians are fitted independently with iterative gradient-based steps. This prevents learned structure from transferring across videos and limits real-time processing.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does HyperGS produce Gaussian representations for a video?",{"text":81,"@type":77},"HyperGS uses a factorized spatiotemporal Transformer to extract tokens from the video, then a learnable query-based Transformer to predict 8-parameter Gaussian splatting representations for each frame in a single forward pass.",{"name":83,"@type":74,"acceptedAnswer":84},"What prevents training collapse when predicting Gaussians feedforward?",{"text":85,"@type":77},"Naively predicting Gaussians across diverse videos can drive them toward needle-like anisotropy, collapsing training mid-way. HyperGS stabilizes optimization using an adaptive rank-based geometric regularizer with dynamically adapting strength based on effective rank.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":107,"slug":137},"General","general"]