[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84942-en":3,"doc-seo-84942-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84942,687197207639,"Asher","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","ELSA3D Elastic Semantic Anchoring for Unified 3D Understanding and Generation","Unified 3D foundation models aim to generate 3D assets and reason about them using language in a single backbone, yet text–3D interaction is often treated implicitly. ELSA3D introduces elastic semantic anchoring by pairing language and geometry at matched abstraction scales. Geometry is encoded with a scale-aware octree tokenizer, while Anchor Tokens are sparse cross-modal units that select semantic cues, route them to the most relevant 3D scale, retrieve scale-specific evidence, and write fused signals back into the unified representation. A lightweight router makes computation elastic. ELSA3D reaches state-of-the-art results in image-to-3D, text-to-3D, and 3D captioning, while cutting FLOPs and latency versus a non-elastic baseline.","arXiv :2607 .06565v 1 [ cs .CV] 7 Jul 2026  \nELSA3D: Elastic Semantic Anchoring for Unified 3D Understanding and Generation  \nTianjiao Yu, Xinzhuo Li, Yifan Shen, Onkar Susladkar, Yuanzhe Liu, Xiaona Zhou, Ismini Lourentzou  \nUniversity of Illinois Urbana-Champaign  \n{ty41, [lourent2}@illinois.edu](lourent2}@illinois.edu)  \n\n| Abstract. Unified 3D foundation models aspire to generate 3D assets and reason about them in language within a single backbone, but their text–3D interaction remains largely implicit. Existing methods concatenate text and 3D tokens into a flat sequence and rely on self-attention, collapsing coarse structural cues and fine geometric details into one undifferentiated representation. We introduce ELSA3D, a unified 3D model that addresses this with elastic semantic anchoring, structuring language and geometric reasoning jointly along matched abstraction scales. ELSA3D represents geometry with a scale-aware octree tokenizer and introduces Anchor Tokens, sparse cross-modal units that select semantic cues, route them to the most relevant 3D scale, retrieve scale-specific geometric evidence, and write the fused signal back into the unified representation, keeping interaction sparse yet precise. A lightweight per-block router makes both computation and reasoning elastic, choosing which text tokens instantiate anchors at which geometric scale so that cross-modal capacity concentrates where alignment is most needed. ELSA3D achieves state-of-the-art performance across image-to-3D generation, text-to-3D generation, and 3D captioning, outperforming the strongest unified baseline while roughly halving FLOPs and inference latency relative to the non-elastic version of the same model. |  |  |\n| --- | --- | --- |\n| [https://plan-lab.github.io/elsa3D](https://plan-lab.github.io/elsa3D) |  |  |\n\n1. Introduction  \nUnified 3D models [87, 91, 96] aim to bridge 3D understanding and generation within one backbone, where a single model can reconstruct a 3D object from an image, generate one from language, describe its structure in text, and support downstream reasoning over geometry. This unification is appealing because generation and understanding can support each other within a shared representation, but it also imposes stronger requirements on multimodal reasoning: a unified 3D model should balance global structure and local detail, translate open-ended language into concrete geometric decisions, and allocate compute dynamically when semantic-geometric alignment requires finer reasoning.  \nCurrent systems fall short on each of these requirements because text-3D interaction remains largely implicit. Previous works [91, 96] concatenate text and 3D tokens into a monolithic sequence and rely on self-attention to discover cross-modal correspondences. Recent 3D advances [13, 14, 21, 103] make the geometric representation multiscale, but the reasoning architecture has not evolved accordingly. What is missing, therefore, isnot merely a stronger hierarchical 3D representation, but a unified design that structures language reasoning and geometric reasoning jointly, enabling structured interaction between semantic cues and geometric content.  \nTo address this gap, we introduce ELSA3D, a unified 3D model built around elastic semantic anchoring. ELSA3D first represents 3D shapes with an octree VQ-VAE in which every content token carries an explicit deterministic scale tag, exposing multiple geometric resolutions to the model. Then, the model organizes language into a semantic trace spanning Global, Structure, and Appearance cues, decomposing text descriptions  \ninto finer semantic granularity. To connect both semantic and geometric abstractions, ELSA3D introduces Anchor Tokens. Each anchor is a transient cross-modal unit instantiated from a selected semantic token, routed to the most relevant 3D scale, fused with scale-specific geometric evidence, and written back into the unified sequence. Anchors keep cross-modal interaction sparse and ","cbCairTM0zXtnEkU","https://ap.wps.com/l/cbCairTM0zXtnEkU","pdf",30380025,1,30,"English","en",105,"# Introduction\n## Unified 3D foundation models\n## Elastic semantic anchoring and Anchor Tokens\n## Elastic routing mechanism\n## Experimental results and contributions","[{\"question\":\"What problem does ELSA3D address in existing unified 3D models?\",\"answer\":\"It targets the largely implicit text–3D interaction in current approaches, where concatenating text and 3D tokens into a single sequence with self-attention tends to blur structural cues and fine geometric details.\"},{\"question\":\"How does ELSA3D perform text–geometry alignment?\",\"answer\":\"It introduces elastic semantic anchoring: Anchor Tokens are instantiated from selected semantic cues, routed to the most relevant 3D scale, fused with scale-specific geometric evidence, and written back into the unified representation.\"},{\"question\":\"What enables ELSA3D to be computationally efficient while keeping precise grounding?\",\"answer\":\"A lightweight elastic router selects which text tokens create anchors, which 3D scale to query, whether each transformer block runs, and how much MLP width to use—concentrating cross-modal reasoning only where alignment matters most.\"}]",1784199608,76,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"elsa3d-elastic-semantic-anchoring-for-unified-3d-understanding-and-generation","",{"@graph":35,"@context":84},[36,53,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/elsa3d-elastic-semantic-anchoring-for-unified-3d-understanding-and-generation/84942/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":4},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"What problem does ELSA3D address in existing unified 3D models?","Question",{"text":74,"@type":75},"It targets the largely implicit text–3D interaction in current approaches, where concatenating text and 3D tokens into a single sequence with self-attention tends to blur structural cues and fine geometric details.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How does ELSA3D perform text–geometry alignment?",{"text":79,"@type":75},"It introduces elastic semantic anchoring: Anchor Tokens are instantiated from selected semantic cues, routed to the most relevant 3D scale, fused with scale-specific geometric evidence, and written back into the unified representation.",{"name":81,"@type":72,"acceptedAnswer":82},"What enables ELSA3D to be computationally efficient while keeping precise grounding?",{"text":83,"@type":75},"A lightweight elastic router selects which text tokens create anchors, which 3D scale to query, whether each transformer block runs, and how much MLP width to use—concentrating cross-modal reasoning only where alignment matters most.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,121,126,129,133],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":21,"slug":120},"research-report",{"id":122,"doc_module":4,"doc_module_name":45,"category_name":123,"show_sort_weight":124,"slug":125},9,"Religion & Spirituality",20,"religion-spirituality",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":124,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]