[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85860-en":3,"doc-seo-85860-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85860,1099514068035,"Ezra","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","CoSAG Compact Semantic Anchor Gaussians via Training-Free Rate-Distortion Coding","Open-vocabulary 3D scene understanding often uses CLIP-like 2D vision-language features embedded into 3D Gaussian Splatting, yielding a text-queryable semantic field. Yet assigning high-dimensional features to millions of Gaussians inflates storage and deployment to gigabytes. CoSAG separates construction from storage: it formulates storage as a rate–distortion coding problem over per-Gaussian binding to a small anchor table, constructed without per-scene training. A spatially predictive entropy coder ships no decoder, achieving sub-megabyte storage with strong accuracy improvements.","CoSAG: Compact Semantic Anchor Gaussians via Training-Free  \nRate–Distortion Coding  \nYuang Jia 1 , Jinlong Wang 1 , Junhong Lin 1 , Ruiting Dai2 , Wei Gao 1 ∗  \n1 SECE, Peking University 2University of Electronic Science and Technology of China  \n[yuangjia8@gmail.com](yuangjia8@gmail.com) [gaowei262@pku.edu.cn](gaowei262@pku.edu.cn)  \n[https://github.com/YuangJia/CoSAG](https://github.com/YuangJia/CoSAG)  \narXiv :2607 . 10237v1 [ cs .CV] 11 Jul 2026  \nAbstract  \nOpen-vocabulary 3D scene understanding is commonly achieved by embedding 2D vision-language features such as CLIP into a 3D Gaussian Splatting scene, turning it into a text-queryable semantic field. However, attaching a highdimensional feature to each of millions of Gaussians inflatesa single scene to gigabytes, which makes storage and deployment the real bottleneck of these fields. Existing compact methods each learn and ship a per-scene codec, an autoencoder, a quantized codebook, or a distilled feature field, entangling field construction with field storage and never compressing the per-Gaussian assignment that holds the bulk of the cost. We argue that construction and storage should bedecoupled, and that storage is a rate–distortion problem over the per-Gaussian binding to a small anchor table, a structure no prior open-vocabulary method compresses. We present CoSAG, which constructs the field without any per-scene training through a closed-form transmittance-weighted lift, spatially grounded semantic anchors, and multi-view denoising, and stores it with a spatially predictive entropy coder that ships no decoder. Because the anchors are spatially grounded, the binding is predictable and therefore highly compressible. The transmittance-weighted lift and multi-view denoising yield a clean, view-consistent assignment, so the entropy coder spends almost no rate on correcting noise and instead codes only the residual against its spatial prediction. CoSAG reaches sub-megabyte storage while matching or exceeding the state of the art across the 2D-rendered, 3D-selection, and dense-LSeg protocols, reducing field size by 37 to 76 × relative to LangSplatV2 at higher accuracy.  \n1 Introduction  \nOpen-vocabulary 3D scene understanding, which segments and localizes arbitrary text-queried objects in a reconstructed scene, underpins intelligent robotics (Peng et al. 2023; Jatavallabhula et al. 2023), autonomous driving (Deng et al. 2025), and augmented reality (Azuma 1997) . Built on 3D Gaussian Splatting (3DGS) (Kerbl et al. 2023), the prevailing approach distills the features of an image-text model such as CLIP (Radford et al. 2021) into the Gaussian primitives, so that the reconstructed scene answers free-form text queries (Qin et al. 2024; Zhou et al. 2024; Wu et al. 2024; Jun-Seong et al. 2025; Zuo et al. 2025; Li et al. 2026a; Jiao  \n∗Corresponding author.  \n(a) Auto Encoder-Decoder Methods (b) Codebook Indexing Methods  \nRender  \nOriginal GS Semantic GS Feature  \n(c) Additional Semantic Field Methods  \nGS-to-Anchor Binding Low  \nEntropy  \n(d) Training-Free Anchor Gaussian Field (Ours)  \nFigure 1: Compressing an open-vocabulary semantic field on the same pretrained 3D Gaussians. (a) Autoencoders suffer bottleneck loss and slow decoding. (b) Codebook indexing incurs extra training and quantization loss. (c) Additional-field methods train a separate semantic field, leaving two disjoint Gaussian fields whose appearance and semantics must be rendered separately. Our CoSAG (d) is training-free: it binds Gaussians to spatially grounded anchors and entropy-codes the resulting binding, reaching submegabyte storage with no decoder and higher accuracy.  \net al. 2025; Li et al. 2025; Zhou et al. 2026) . Such fields deliver strong open-vocabulary capability, but endowing every Gaussian with semantics is costly: a CLIP feature is highdimensional and a scene holds millions of Gaussians, so a single semantic field reaches the gigabyte scale. It dwarfs the geometry it annotates and becomes","cbCaiaIffSsNZGJF","https://ap.wps.com/l/cbCaiaIffSsNZGJF","pdf",11062447,3,1,14,"English","en",105,"# Abstract\n# 1 Introduction\n## Motivation: open-vocabulary 3D semantic fields\n## Limits of existing compact methods\n## Core idea: decouple construction and storage","[{\"question\":\"Why do open-vocabulary semantic fields become too large for storage and deployment?\",\"answer\":\"Because each of millions of 3D Gaussians gets a high-dimensional vision-language feature, making the semantic field reach gigabyte scale and dominating overall cost.\"},{\"question\":\"What is the key contribution of CoSAG compared with prior compact methods?\",\"answer\":\"CoSAG decouples field construction from storage by using a training-free construction strategy and treating storage as rate–distortion coding over binding to a small spatial anchor table.\"},{\"question\":\"How does CoSAG enable effective compression with a training-free approach?\",\"answer\":\"It constructs the semantic field without per-scene training using a closed-form transmittance-weighted lift and spatially grounded semantic anchors, then entropy-codes the predictable per-Gaussian binding residuals.\"}]",1784206754,35,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"cosag-compact-semantic-anchor-gaussians-via-training-free-rate-distortion-coding","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/cosag-compact-semantic-anchor-gaussians-via-training-free-rate-distortion-coding/85860/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do open-vocabulary semantic fields become too large for storage and deployment?","Question",{"text":75,"@type":76},"Because each of millions of 3D Gaussians gets a high-dimensional vision-language feature, making the semantic field reach gigabyte scale and dominating overall cost.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the key contribution of CoSAG compared with prior compact methods?",{"text":80,"@type":76},"CoSAG decouples field construction from storage by using a training-free construction strategy and treating storage as rate–distortion coding over binding to a small spatial anchor table.",{"name":82,"@type":73,"acceptedAnswer":83},"How does CoSAG enable effective compression with a training-free approach?",{"text":84,"@type":76},"It constructs the semantic field without per-scene training using a closed-form transmittance-weighted lift and spatially grounded semantic anchors, then entropy-codes the predictable per-Gaussian binding residuals.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]