[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83762-en":3,"doc-seo-83762-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83762,137441390410,"Hazel","https://ap-avatar.wpscdn.com/avatar/2000252f4ab5702993?_k=1776741390130283984",8,"Research & Report","Scalable Semantic Steering of Embedding Projections","Low-dimensional projections enable interactive exploration of high-dimensional embedding spaces, yet their layouts often fail to reflect semantic relationships defined by analysts. LLM-augmented semantic steering addresses this by using natural-language intent from seed groups, but item-level reasoning triggers one LLM call per item, making cost scale linearly with dataset size. This work introduces group-level semantic prototypes: a single LLM call generates structured group profiles, which are embedded and fused with seed centroids. Intent is then propagated via embedding-space soft assignment, abstention, and alignment-scaled updates before reprojection. On a 5K-document LitCovid corpus, global alignment matches per-item steering while reducing LLM calls by over three orders of magnitude, and the approach extends to multimodal embeddings.","Scalable Semantic Steering of Embedding Projections  \nWei Liu* Virginia Tech  \nEric Krokos†  \nDepartment of Defense  \nKirsten Whitley‡ Department of Defense  \nRebecca Faust § Tulane University  \nChris North¶ Virginia Tech  \narXiv :2607 .03978v 1 [ cs .HC] 4 Jul 2026  \nFigure 1: Prototype-based semantic steering on a text (LitCovid) dataset and an image (Human Actions) dataset. In each pair, the baseline projection (left) is reorganized into a more semantically coherent layout (right) using five user-selected seed examples per group (circled) . Our method propagates analyst intent across the full collection through embedding-space operations driven by a single LLM call.  \nABSTRACT  \nLow-dimensional projections support interactive visual analysis of high-dimensional data embeddings, but their structure often does not align with analyst-defined semantic relationships. Recent LLMaugmented semantic steering methods address this gap by externalizing analyst intent from user-defined groups of seed examples, but they propagate intent through per-item LLM reasoning, causing LLM calls and cost to grow linearly with collection size. We propose a scalable semantic steering method that shifts semantic computation from individual items to user-defined groups. A single LLM call generates structured profiles for all groups, which are embedded and combined with seed centroids to form hybrid semantic prototypes. The method then propagates intent without retraining, using embedding-space soft assignment, abstention, and alignmentscaled updates before reprojection. On a 5K-document LitCovid corpus, our method achieves global alignment comparable to peritem LLM steering while reducing LLM calls by over three orders of magnitude. An image case study shows that the same prototypebased mechanism extends to multimodal embeddings. These results suggest that group-level representations can make semantic steering more practical for larger embedding collections.  \nIndex Terms: Semantic Steering, Semantic Interaction, Embedding Projections, Large Language Models, Semantic Prototypes.  \n1 INTRODUCTION  \nLow-dimensional projections of high-dimensional embeddings are widely used to support visual analysis of text and image collections, enabling analysts to explore semantic structure through spatial organization [10, 13] . However, projection structure is largely determined by the underlying embedding model and dimensionality  \n*[e-mail: wliu3@vt.edu](e-mail: wliu3@vt.edu), ORCID: 0009-0009-6340-8912  \n†e-mail: [ericpkrokos@gmail.com](ericpkrokos@gmail.com), ORCID: 0000-0003-1350-5297 ‡e-mail: [visual.tycho@gmail.com](visual.tycho@gmail.com), ORCID: 0000-0003-1356-326X § e-mail: [rfaust1@tulane.edu](rfaust1@tulane.edu), ORCID: 0000-0002-7640-1287 ¶ e-mail: [north@vt.edu](north@vt.edu), ORCID: 0000-0002-8786-7103  \nreduction technique, which may not align with the semantic relationships that analysts intend to examine for a given task [11, 1] . Semantic interaction (SI) [6] addresses this gap by allowing analysts to reshape projections through spatial actions such as grouping or repositioning items. More recent work augments SI with large language models (LLMs) to externalize analyst intent as naturallanguage semantics rather than as implicit model parameters, propagating it across the dataset through item-level reasoning [12, 16] .  \nWhile effective, this design has a fundamental scalability bottleneck: semantic propagation is performed per item. Given a small set of seed examples, an LLM is invoked for every remaining item to evaluate its relationship to the analyst’s intent and generate itemlevel augmentations. The cost grows linearly with collection size. On a 5,000-item corpus, a single steering interaction issues over 5,000 LLM calls and costs roughly $32 under current pricing [17](Section 4.3) . As datasets grow, this design becomes increasingly impractical for interactive use.  \nWe make a structural observation: in SI, analysts often express intent at t","cbCaimj4okGzbE0Z","https://ap.wps.com/l/cbCaimj4okGzbE0Z","pdf",1257675,3,1,5,"English","en",105,"# Abstract\n# Introduction\n## Problem: item-level LLM reasoning bottleneck\n## Key idea: group-level semantic prototypes\n## Contributions","[{\"question\":\"What scalability bottleneck limits existing LLM-augmented semantic steering methods?\",\"answer\":\"They propagate semantic intent through item-level reasoning, invoking the LLM for every non-seed item. This makes the number of LLM calls—and cost—grow linearly with collection size.\"},{\"question\":\"How does the proposed method reduce LLM calls while still steering projections semantically?\",\"answer\":\"It externalizes intent once at the group level: a single LLM call generates structured semantic profiles for all user-defined groups. These profiles are embedded and combined with seed centroids to form hybrid semantic prototypes, enabling propagation via embedding-space operations.\"},{\"question\":\"What mechanisms are used to propagate analyst intent after forming hybrid prototypes?\",\"answer\":\"The method uses embedding-space soft assignment, abstention, and alignment-scaled updates before reprojection, avoiding retraining while aligning the projection to the intended semantics.\"}]",1784190269,13,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"scalable-semantic-steering-of-embedding-projections","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/scalable-semantic-steering-of-embedding-projections/83762/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What scalability bottleneck limits existing LLM-augmented semantic steering methods?","Question",{"text":75,"@type":76},"They propagate semantic intent through item-level reasoning, invoking the LLM for every non-seed item. This makes the number of LLM calls—and cost—grow linearly with collection size.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed method reduce LLM calls while still steering projections semantically?",{"text":80,"@type":76},"It externalizes intent once at the group level: a single LLM call generates structured semantic profiles for all user-defined groups. These profiles are embedded and combined with seed centroids to form hybrid semantic prototypes, enabling propagation via embedding-space operations.",{"name":82,"@type":73,"acceptedAnswer":83},"What mechanisms are used to propagate analyst intent after forming hybrid prototypes?",{"text":84,"@type":76},"The method uses embedding-space soft assignment, abstention, and alignment-scaled updates before reprojection, avoiding retraining while aligning the projection to the intended semantics.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,109,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":106,"show_sort_weight":107,"slug":108},"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":22,"slug":137},19,"General","general"]